Skip to main content
Category: Digital Preservation

Representation Information

Simply put

Representation Information is the additional information needed to make sense of a stored digital object, which on its own is just a sequence of bits. It provides the meaning and context required to turn those raw bits into something understandable, such as readable text or a viewable image. Without it, a preserved digital file may be impossible to interpret correctly over time.

Formal definition

In digital preservation practice, Representation Information denotes the information that maps a stored data object, or sequence of bits, into more meaningful concepts so that it can be interpreted and rendered by its intended community. It typically accompanies a digital object to supply the structural and semantic knowledge required to reconstruct meaning from the raw bit sequence, supporting the object's continued usability across time. The specific scope and layering of Representation Information depend on the object type and the preservation framework applied; the sources in this packet describe its general purpose rather than a fixed set of components.

Why it matters

A preserved digital object is, at base, only a sequence of bits. Without additional information to explain how those bits should be interpreted, a stored file may become impossible to render or understand correctly as the software, formats, and knowledge originally used to create it fall away over time. Representation Information addresses this risk by supplying the structural and semantic knowledge needed to reconstruct meaning from the raw bit sequence, which is central to keeping digital material usable rather than merely stored.

The consequence of neglecting Representation Information is that a file may survive intact as bits yet lose its interpretability, effectively becoming unreadable to the community that needs it. This distinguishes bit-level preservation, which keeps the sequence of bits from corruption or loss, from meaningful preservation, which keeps the object understandable and renderable. For records held over long periods, this distinction matters because an object that cannot be interpreted may not reliably serve as usable evidence of the activity it documents.

The extent and layering of Representation Information depend on the type of object and the preservation framework applied, so organizations typically need to consider what their intended community will require to make sense of an object over time rather than assuming a fixed, universal set of components.

Who it's relevant to

Digital preservation specialists
Those responsible for keeping digital objects usable over the long term rely on Representation Information to ensure that preserved bits can still be interpreted and rendered, not merely stored. It underpins the distinction between preserving a file at the bit level and preserving its meaning.
Archivists and records managers handling digital holdings
Professionals managing digital records over extended retention or permanent preservation need to consider what information will be required for future users to make sense of an object. Capturing appropriate Representation Information supports the continued usability, and therefore the evidential value, of records held over time.
Repository and system designers
Those building or configuring preservation repositories and systems must account for how Representation Information is associated with stored objects and how it will be maintained, since the appropriate scope and layering depend on the object types held and the preservation framework applied.
Information governance and compliance leads
Where usable digital records are needed to meet accountability, regulatory, or evidential expectations that vary by jurisdiction and sector, understanding Representation Information helps clarify that surviving bits alone may not guarantee that an object remains interpretable when it is later required.

Inside Representation Information

Structure Information
The information that describes the format and organization of a digital object's bit stream, such as data types, field lengths, and file formats, enabling the raw bits to be parsed into meaningful data elements. Without structure information, a preserved bit stream typically cannot be reliably rendered or interpreted.
Semantic Information
The information that conveys the meaning of the parsed data, including vocabularies, definitions, and contextual rules needed to understand what the data elements represent. This layer builds on structure information and is often essential for a designated community to understand a record's content.
Other Representation Information
Additional information sometimes required to fully interpret an object, which may include software, hardware, or standards documentation referenced by the structure and semantic layers. Depending on the object, this can point to further representation information, forming a recursive network of dependencies.
Representation Network
The interlinked set of representation information objects that together support interpretation, since one piece of representation information may itself require further representation information to be understood. The extent of the network typically depends on the knowledge base assumed for the intended audience.
Designated Community Dependency
The relationship between representation information and the assumed knowledge base of the community expected to use the material. What must be made explicit as representation information often depends on what that community can already be presumed to understand.

Common questions

Answers to the questions practitioners most commonly ask about Representation Information.

Is representation information the same thing as file format metadata?
No, though the two are related and often confused. File format metadata typically identifies the format of a digital object, whereas representation information is the broader body of knowledge required to render and interpret that object meaningfully over time. Representation information may include format specifications but generally extends further, encompassing the structural, semantic, and other contextual information needed to turn a stored bit sequence into something a designated community can understand. Knowing a file's format is often a starting point, not the full set of representation information.
Does keeping the original file guarantee that a record remains understandable in the future?
Not necessarily. Preserving the bit sequence of an original file addresses storage and integrity, but understandability depends on whether the representation information needed to interpret those bits remains available and comprehensible to the intended community. Over time, the knowledge, software, and specifications required to render a format can become obsolete or inaccessible. For this reason, preservation approaches typically treat representation information as something to be identified, documented, and maintained alongside the object, rather than assuming the file alone is self-explanatory.
How can an organization begin identifying the representation information its digital records require?
A common starting point is to analyze the digital objects in scope and determine what knowledge a future user would need to interpret them. This often involves documenting the formats present, the structures within them, and the semantics required to make sense of the content. Organizations frequently consider who the intended audience is, since the amount of representation information needed depends on what that audience can already be assumed to understand. The specific approach will depend on organizational policy, the nature of the holdings, and available resources.
How does the designated community affect how much representation information needs to be captured?
The concept of a designated community is central because representation information is defined relative to the knowledge that community is assumed to hold. Where a community can be relied upon to understand certain formats or terminology, less needs to be documented explicitly. Where less can be assumed, more representation information may need to be recorded to bridge the gap. Because assumed knowledge can change over time, organizations often revisit these assumptions as part of ongoing preservation planning.
Where can representation information be stored in relation to the records it describes?
Representation information can be maintained alongside the object, held in shared repositories or registries, or referenced through links, depending on organizational policy and system design. Practices vary, and some organizations may combine approaches. The key consideration is typically ensuring that the representation information remains available, discoverable, and usable for as long as the associated records must remain understandable, rather than mandating any single storage arrangement.
How should representation information be maintained over the long term?
Because the knowledge and technology needed to interpret digital objects can shift over time, representation information is generally treated as something requiring ongoing management rather than a one-time capture. This can involve periodic review of whether the documented information remains sufficient for the intended community, and updating or supplementing it as assumptions and technologies change. The frequency and depth of such review will depend on the significance of the records, applicable policy, and available resources.

Common misconceptions

Representation information is the same as descriptive metadata used to locate or catalogue a record.
Representation information is chiefly concerned with enabling the interpretation and rendering of a preserved object from its bit stream, rather than with discovery or resource description. Descriptive metadata often serves a distinct purpose, and the two should not be treated as interchangeable, though they may coexist within a broader preservation package.
Preserving the file itself is sufficient, so representation information is optional.
A preserved bit stream may be unusable over time without accompanying structure and semantic information to parse and interpret it. Representation information is generally what allows an object to remain understandable to its designated community as technology and knowledge change, so it is typically regarded as integral rather than optional.
Representation information is a single, fixed document attached to an object.
Representation information is often better understood as a network in which one component may depend on further components. What needs to be captured typically varies with the object and with the assumed knowledge base of the intended community, so it is rarely a static, self-contained artifact.

Best practices

Define the designated community and its assumed knowledge base explicitly, since this determines how much structure and semantic information must be made available rather than presumed.
Capture both structure and semantic information for preserved objects, so that a bit stream can be parsed into data and that data can be understood in context.
Map the representation network for significant object types, identifying where one piece of representation information depends on further information, and record those dependencies deliberately.
Review representation information periodically against changes in technology and in the community's knowledge base, updating or extending it where interpretation would otherwise be at risk.
Keep representation information linked to the objects it supports, so that the relationship between an object and the information needed to interpret it is maintained over time.
Distinguish representation information from descriptive or discovery metadata in your documentation and systems, to avoid conflating the interpretation function with cataloguing functions.