Skip to main content
Category: Digital Preservation

Content Data Object

Also known as:
Simply put

A Content Data Object is the actual digital material that an organization sets out to preserve, such as a file or set of data. On its own, it may not be fully understandable, so it is preserved together with additional information that explains how to interpret it. The term comes from the field of digital preservation rather than from everyday computing usage.

Formal definition

Within the Open Archival Information System (OAIS) reference model, a Content Data Object is the Data Object that, together with its associated Representation Information, forms the original target of preservation. It denotes the specific digital object (for example, a bitstream or structured data) whose long-term preservation is the objective, and its interpretability typically depends on the Representation Information bound to it. The term as used here is grounded in the digital preservation and archival context and should be distinguished from more general software or application-development senses of 'data object,' which describe programmatic structures for grouping related data fields.

Why it matters

The concept of a Content Data Object matters because it names, with precision, the specific digital material that a preservation effort is actually trying to keep. In digital preservation practice, it is not enough to store a bitstream or a set of structured data; over time, the technical and contextual knowledge needed to open, render, and correctly interpret that material can be lost. Isolating the Content Data Object as the target of preservation forces an organization to ask a companion question: what additional information is needed to make this object understandable in the future? This framing helps preservation programs avoid the common trap of assuming that a preserved file will remain usable simply because it is stored intact.

The term also serves an important disambiguating function. Because 'data object' has a looser, more general meaning in software and application development, where it typically describes a programmatic structure for grouping related data fields, professionals working across disciplines can talk past one another. Within the Open Archival Information System (OAIS) reference model, the Content Data Object has a narrower and more deliberate meaning tied to long-term preservation. Keeping this distinction clear helps records and preservation teams communicate accurately about what is being preserved and why interpretability, not just storage, is the objective.

Who it's relevant to

Digital preservation specialists
Those responsible for long-term preservation programs work directly with the Content Data Object as the defined target of preservation. Understanding the term in its OAIS sense helps them plan for interpretability by ensuring that Representation Information is captured and bound to the object, rather than assuming that a stored bitstream will remain usable on its own.
Archivists managing digital holdings
Archivists dealing with digital objects benefit from the precise scoping this term provides. It helps distinguish the material being preserved from the supporting information needed to interpret it, supporting more deliberate decisions about what must accompany an object into long-term custody.
Records and information governance professionals
For those coordinating across records management and preservation functions, the term is useful for clear communication. Recognizing that 'Content Data Object' carries a specific digital preservation meaning, distinct from the general software or application-development sense of 'data object,' helps avoid confusion when preservation requirements are being defined.

Inside CDO

Content Component
The substantive information payload that a Content Data Object carries, representing the actual content that is intended to be preserved and made usable over time. In preservation contexts, this is often the bit sequence or set of bit sequences that requires interpretation to become meaningful.
Representation Information
The information typically required to interpret the content component and render it intelligible to a designated community. Depending on the preservation model in use, this may include structural, semantic, and other contextual details, though the exact composition varies by framework and organizational policy.
Bit-Level Encoding
The underlying digital encoding of the content, which generally must be understood in combination with representation information before the object can be interpreted. On its own, the raw bit sequence is not equivalent to a usable record.
Associated Metadata Linkages
The relationships between the content and its descriptive, technical, and contextual metadata. These linkages are often what allow a Content Data Object to be located, understood, and maintained as authentic and usable, though the specific metadata retained depends on jurisdiction, sector, and organizational policy.

Common questions

Answers to the questions practitioners most commonly ask about CDO.

Is a content data object the same thing as a record?
No. A content data object is the digital bitstream or data content that carries the informational payload, whereas a record is the authoritative evidence of an activity, comprising content together with its context and structure. A content data object may form part of a record, but on its own it typically lacks the metadata, contextual linkages, and management controls that give a record its authenticity, reliability, integrity, and usability. Treating the two as identical risks losing the evidential properties that distinguish a record from mere information.
Does capturing a content data object mean the underlying information has been preserved as a record?
Not necessarily. Capturing a content data object secures the bitstream, but preservation as a record generally requires that associated metadata, context, and structure be captured and maintained as well, so that authenticity and usability can be sustained over time. Depending on organizational policy and the applicable preservation approach, a captured content data object without its supporting information may amount to transitory or partial content rather than a fully constituted, preservable record.
How should a content data object be linked to its associated metadata during capture?
In many implementations, the content data object is bound to descriptive, structural, and contextual metadata at the point of capture so that the two remain associated throughout the lifecycle. Approaches vary by system and by organizational policy, and can include packaging content and metadata together, maintaining persistent identifiers, or holding references within a controlled repository. The aim is typically to ensure the association is maintained through classification, retention, and any subsequent transfer or migration.
What happens to a content data object during migration or format conversion?
When a content data object is migrated or converted, the objective is generally to retain its usability and integrity while updating the format or storage environment. Because conversion can alter the bitstream, organizations often document the transformation, retain evidence of the process, and verify that the informational content and any relevant properties are preserved. The controls applied typically depend on organizational policy, the preservation strategy in use, and the evidential value assigned to the content.
How does the disposition of a content data object relate to the disposition of the record it supports?
Disposition decisions generally apply at the record level, and the content data object is acted upon in accordance with the disposition authorized for the record it forms part of. Depending on organizational policy, disposition may involve transfer, permanent preservation, or destruction, and it should not be assumed to mean destruction alone. Where a single content data object supports more than one record, or is shared across contexts, the interaction of applicable retention and disposition rules typically needs to be resolved before action is taken.
How can the integrity of a content data object be verified over time?
Integrity verification commonly relies on techniques such as fixity checks, which detect unintended changes to the bitstream, alongside documented handling and audit information. Establishing integrity at capture and re-verifying it periodically, and around events such as migration or transfer, can help demonstrate that the content data object has not been altered without authorization. The specific methods and their frequency typically depend on organizational policy, system capability, and the evidential requirements applicable to the content.

Common misconceptions

A Content Data Object is the same thing as a record.
The two concepts are related but distinct. A Content Data Object typically refers to the content payload requiring interpretation, whereas a record is content plus the context and structure that give it evidential value. A record generally must exhibit properties such as authenticity, reliability, integrity, and usability, and a Content Data Object alone does not necessarily satisfy these. The relationship depends on the preservation model and organizational policy in use.
The raw bit sequence carries meaning on its own.
In most preservation approaches, a bit sequence is not intelligible without associated representation information that enables interpretation. Content and the information needed to render it usable are typically treated as separate but linked elements, and losing the interpretive information can render the content unusable even when the bits are intact.
Preserving a Content Data Object is equivalent to archiving or long-term retention on its own.
Holding a Content Data Object does not by itself constitute a full preservation, retention, or disposition outcome. Retention, transfer, permanent preservation, and destruction are distinct lifecycle actions, and maintaining usability over time generally requires ongoing management of both the content and its associated representation information and metadata, as governed by applicable policy and jurisdictional requirements.

Best practices

Maintain the content component together with the representation information needed to interpret it, since separating the two often renders the content unusable over time.
Establish and preserve the linkages between the Content Data Object and its associated metadata so the content remains locatable, interpretable, and defensible.
Do not treat a Content Data Object as automatically equivalent to an authoritative record; assess whether the necessary context and structure are present to support authenticity, reliability, integrity, and usability.
Document the assumptions about the designated community or intended users, as the level of representation information required to keep content intelligible depends on who must be able to interpret it.
Apply lifecycle actions such as retention, transfer, or disposition to Content Data Objects according to applicable organizational policy and jurisdictional requirements, rather than assuming preservation is a single fixed action.
Where the specific preservation model, standard, or metadata scheme in use affects how a Content Data Object is defined and managed, record that model explicitly so future custodians understand the interpretive framework applied.