Skip to main content
Category: Digital Preservation

Normalization

Also known as: Data Normalization, Database Normalization
Simply put

Normalization is a process for organizing data so that it is structured consistently and stored efficiently. In database design, it typically involves arranging data into related tables to reduce duplication and avoid inconsistencies. The term is also used more loosely to mean making data entries appear uniform across fields and records.

Formal definition

Normalization commonly refers to a database design process that organizes data into structured table relationships in order to reduce redundancy, eliminate anomalies, and improve data integrity, consistency, and accuracy. The term carries related but distinct meanings across disciplines: in relational database theory it concerns decomposing data into tables to control redundancy and dependency, while in data preparation and analytics it can denote the practice of standardizing entries so they appear uniform across fields and records. In a separate machine learning sense, normalization refers to transforming numerical features so they span a similar scale. It should be noted that normalization in this data-structuring sense is distinct from records management concepts such as classification, retention, and disposition; the evidence provided addresses database and data-processing usage only, and does not establish a recordkeeping-specific meaning.

Why it matters

Normalization matters because the way data is structured directly affects its integrity, consistency, and accuracy. When data is organized to reduce redundancy and avoid duplication, organizations are less likely to encounter conflicting values scattered across multiple locations, which can undermine trust in the information and complicate downstream reporting, analysis, and retrieval. In database design, controlling redundancy and dependency through structured table relationships helps prevent the kinds of anomalies that arise when the same fact is stored inconsistently in several places.

For information governance professionals, it is important to recognize where normalization sits and where it does not. Normalization in the data-structuring sense is concerned with how data is organized and stored efficiently; it is distinct from records management concepts such as classification, retention, and disposition. Consistent, well-structured data can support governance objectives by making information easier to locate and reconcile, but normalization by itself does not establish whether something is an authoritative record, how long it should be kept, or how it should ultimately be disposed of. Treating the two as interchangeable risks conflating data management practices with recordkeeping controls.

A further reason for care is that the term carries related but distinct meanings across disciplines. In relational database theory it refers to decomposing data into tables to control redundancy and dependency; in data preparation and analytics it can mean standardizing entries so they appear uniform across fields and records; and in a separate machine learning context it refers to transforming numerical features onto a similar scale. Because these usages differ, professionals should confirm which sense is intended in any given context rather than assume a single universal meaning.

Who it's relevant to

Data managers and database designers
Those responsible for designing and maintaining databases rely on normalization to organize data into structured table relationships, reduce duplication, and control the anomalies that can undermine data integrity. Understanding the relational sense of the term is central to their work of storing data efficiently and consistently.
Data analysts and data preparation specialists
Professionals preparing data for reporting or analysis encounter normalization in the sense of standardizing entries so they appear uniform across fields and records. This usage is related to, but distinct from, relational database normalization, and clarity about which is meant helps avoid confusion during data preparation.
Information governance and records professionals
Governance and records practitioners benefit from consistent, well-structured data, but should note that normalization in the data-structuring sense is distinct from records management concepts such as classification, retention, and disposition. The evidence addresses database and data-processing usage only and does not establish a recordkeeping-specific meaning, so the term should not be treated as a recordkeeping control.
Machine learning practitioners
In machine learning, normalization refers to transforming numerical features so they span a similar scale. This is a separate meaning from database normalization, and practitioners in this field should be aware that the shared term reflects different processes and objectives.

Inside Normalization

Format standardization
The conversion of records or their metadata into consistent, agreed formats so that content, structure, and context can be reliably interpreted and managed over time. In recordkeeping this often involves migrating content to stable or open formats to support long-term usability, though the specific target formats depend on organizational policy and preservation objectives.
Metadata harmonization
The alignment of metadata values, element names, and encoding conventions across systems or collections so that records can be described, searched, and retrieved consistently. This may include mapping local metadata schemes to a common scheme, but the appropriate scheme typically depends on jurisdiction, sector, and the standards an organization has adopted.
Value and structure regularization
The process of bringing data values into a controlled and consistent form, for example applying controlled vocabularies, consistent date formats, or standardized naming conventions. This supports the integrity and usability of records as evidence, though it should be applied in ways that do not compromise the authenticity of the original record.
Reconciliation with source records
The retention of a defensible link between normalized outputs and the authoritative source record from which they derive. Normalization typically produces a managed derivative, and preserving the relationship to the original supports authenticity, reliability, and the ability to demonstrate integrity over time.

Common questions

Answers to the questions practitioners most commonly ask about Normalization.

Does normalization mean converting records to a single 'best' format that will last forever?
No. Normalization typically involves converting records into a smaller set of preferred or standardized formats considered more sustainable for preservation or management, but this should not be understood as producing a permanent or universally optimal format. Format suitability depends on the record type, the preservation objectives, and the technology environment, and preferred formats may themselves require future migration. Normalization reduces format diversity to a manageable set; it does not guarantee indefinite longevity.
Is normalization the same as simply making a copy of a record in a new format?
Not exactly. While normalization does produce a version in a target format, the distinction that matters in recordkeeping is what happens to the record's evidential properties. Converting a format can affect authenticity, integrity, and usability, and organizations often need to consider whether the normalized version is treated as the authoritative record or as a derivative, and how the relationship to the original is documented. A routine copy does not necessarily address these considerations, so normalization is better understood as a controlled conversion process rather than mere copying.
At what point in the lifecycle is normalization typically applied?
Normalization can occur at different stages depending on organizational policy. It is often applied at or near the point of capture, so that records enter a repository in preferred formats, or later as part of preservation activity when formats are at risk of obsolescence. The timing chosen typically reflects a balance between controlling format proliferation early and preserving the original as received. Where normalization happens later, retaining the original alongside the normalized version is a common consideration.
How should the original record be handled after normalization?
Practice varies by policy and preservation strategy. In many approaches the original is retained alongside the normalized version, particularly where authenticity or evidential value could be questioned, so that the conversion can be verified or reversed if needed. In other cases, where storage or simplicity is prioritized and the normalized format is deemed sufficient, the original may be dispositioned according to the organization's rules. The decision typically depends on risk, sector requirements, and the intended authoritative status of each version.
What should be documented when normalizing records?
Documentation typically supports the ongoing authenticity and reliability of the normalized record. Organizations often record details such as the source and target formats, the tools and settings used, the date and responsible actor, and any changes or losses observed during conversion. Maintaining this information as part of the record's metadata or audit trail helps demonstrate integrity and allows the conversion to be understood or challenged later. The specific requirements depend on organizational policy and any applicable standards or regulatory obligations.
How can an organization decide which target formats to normalize to?
Selection generally weighs factors such as how well a format supports the record's usability and evidential properties, its openness and level of adoption, the availability of tools to render and validate it, and its suitability for the intended retention period. Because no format is guaranteed to remain sustainable indefinitely, organizations often revisit their preferred formats over time and may pair normalization with ongoing format monitoring. The appropriate choices depend on the record types involved, preservation objectives, and the organization's technology environment.

Common misconceptions

Normalization in recordkeeping is the same as database normalization in the data management sense.
The term normalization carries a specific technical meaning in database design, concerned with reducing redundancy in relational structures, that differs from its use in recordkeeping and preservation contexts, where it typically refers to standardizing formats, metadata, or values for consistency and usability. Practitioners should be explicit about which sense they mean, since records management and data management are distinct disciplines that overlap only partly.
Normalizing a record replaces the original, so the source can be discarded.
Normalization typically produces a managed derivative rather than a substitute for the authoritative record. Whether the source can be disposed of depends on organizational policy, the properties that establish authenticity and integrity, and any applicable retention obligations, which vary by jurisdiction and sector. Destroying a source record is a disposition decision that should be governed separately and not assumed to follow automatically from normalization.
Normalization guarantees that a record's evidential value is preserved.
Standardizing format or values can support usability and consistency, but any transformation carries some risk to authenticity, reliability, or integrity if not carefully controlled. Preservation of evidential value depends on how the process is documented, validated, and linked back to the source, rather than being an inherent outcome of normalization itself.

Best practices

Define and document the target formats, metadata schemes, and controlled vocabularies used for normalization, aligning them with standards and preservation objectives appropriate to your organization's jurisdiction and sector.
Retain a defensible link between each normalized output and its authoritative source record so that authenticity, reliability, and integrity can be demonstrated over time.
Treat any decision to dispose of source records after normalization as a separate, policy-governed disposition action rather than an automatic consequence of the process, and check applicable retention obligations first.
Capture metadata describing the normalization process itself, including what was transformed, when, by whom or by what system, and against which rules, to support later verification.
Validate normalized outputs against the source to confirm that content, structure, and context have been preserved and that no unintended loss of evidential value has occurred.
Clarify which meaning of normalization is intended in policies and system documentation, distinguishing recordkeeping format and metadata standardization from database normalization to avoid confusion across teams.