Skip to main content
Category: Systems and Technology

Optical Character Recognition

Also known as: OCR, Optical Character Recognition, text recognition
Simply put

Optical Character Recognition (OCR) is a technology that converts images of text, such as scanned documents or photographs, into machine-readable text. This allows text that was previously locked inside an image to be searched, copied, or edited. OCR is commonly applied to typed, printed, and in some cases handwritten material.

Formal definition

Optical Character Recognition (OCR) is an automated process that analyzes images of text and extracts the characters into a machine-encoded, machine-readable format. It typically operates on scanned images, photographs, or image-based documents such as non-searchable PDFs, converting typed, printed, or handwritten content into text that can be indexed, searched, copied, or edited. In a recordkeeping context, OCR is a processing step applied to digitized or born-digital image content; the accuracy and completeness of its output can vary with source quality, and organizations should note that OCR output is a derived text layer rather than an alteration of the underlying image, which has implications for authenticity, integrity, and how the digitized record and its text representation are captured and managed.

Why it matters

For records professionals, OCR is often the step that turns a large body of digitized images into content that can be searched, indexed, and retrieved. Without a text layer, a scanned document is effectively an image whose contents cannot be located through full-text search, which limits discoverability and can hamper responses to access requests, litigation-related searches, and routine business use. Applying OCR to image-based material therefore has a direct bearing on how efficiently an organization can find and use its records.

OCR also carries implications for authenticity and integrity that professionals should treat carefully. The text produced by OCR is a derived representation rather than a modification of the underlying image; the original digitized image typically remains unchanged. This distinction matters when the digitized image is intended to serve as the authoritative record, because the searchable text layer is a convenience derived from that image and may contain errors. Depending on organizational policy and the sensitivity of the material, it may be important to document how and when OCR was applied and to retain the source image as the reference point for the record's content.

Because OCR accuracy and completeness can vary with the quality of the source, professionals should not assume that a searchable document contains a perfect transcription of its image. Faint, damaged, skewed, or handwritten source material may yield incomplete or incorrect text. Where OCR output is relied upon for search, redaction, or downstream processing, organizations typically need to account for this variability in their quality controls and in any assurances they give about completeness of retrieval.

Who it's relevant to

Records managers and digitization program leads
Those overseeing scanning and digitization projects rely on OCR to make image-based holdings searchable and usable. They typically need to decide when OCR is applied in the workflow, how its output is stored alongside the source image, and how to handle variability in accuracy so that retrieval and use of records are supported without overstating the reliability of the text layer.
Information governance and compliance officers
OCR affects the discoverability of content across a repository, which bears on access requests, searches, and consistent application of policy. Governance staff often need to understand that OCR output is a derived and potentially imperfect representation, so that assurances about search completeness and content retrieval are qualified appropriately.
Archivists and preservation staff
For those responsible for authoritative digitized records, the distinction between the underlying image and the derived text layer is significant. OCR does not alter the source image, but its output introduces a text representation whose creation and quality may need to be documented to support authenticity, integrity, and long-term usability of the record.
Legal, discovery, and access-to-information staff
Staff conducting searches for litigation, freedom of information, or similar processes depend on searchable text to locate relevant material. Because OCR accuracy varies with source quality, these users should be aware that image-based records may not be fully captured in text-based searches, and that additional review may be warranted depending on jurisdiction and the stakes of the matter.

Inside OCR

Text Recognition Engine
The core software component that analyzes scanned or digital images and converts detected shapes into machine-encoded characters. Recognition accuracy typically depends on image quality, font characteristics, language, and the condition of the source material.
Image Pre-processing
Steps applied before recognition, such as deskewing, despeckling, contrast adjustment, and binarization, intended to improve the legibility of the source image. The extent and effectiveness of pre-processing often influences the reliability of the resulting text output.
Output Text Layer
The machine-readable text produced by the process, which may be embedded within an image-based document (for example, a searchable image format) or exported separately. This layer supports indexing and full-text search but is not necessarily a faithful reproduction of the original.
Confidence and Error Handling
Mechanisms that flag uncertain recognition results, often expressed as confidence scores. These support quality review but do not by themselves guarantee correctness, and manual verification is frequently required for records where accuracy is critical.
Language and Character Set Support
The range of scripts, alphabets, and symbol sets a given engine can interpret. Coverage varies by tool and configuration, and recognition of handwriting, degraded text, or unusual layouts is typically less reliable than recognition of clean printed text.

Common questions

Answers to the questions practitioners most commonly ask about OCR.

Does running OCR on a scanned document turn the image into an authoritative record?
Not by itself. OCR produces a machine-readable text layer derived from an image, but the OCR output is an interpretation of the source and may contain recognition errors. Whether the resulting object qualifies as an authoritative record depends on the recordkeeping context, including how authenticity, reliability, integrity, and usability are established and maintained. Typically the scanned image (and its capture metadata) remains the reference point, while the OCR text supports search and access. Organizations should define, through policy and often with reference to standards such as ISO 15489, which representation is treated as the record.
Is OCR the same as digitization or scanning?
No. Scanning or imaging captures a visual representation of a physical or digital source, producing an image file. OCR is a subsequent processing step that attempts to recognize and extract text from that image. Digitization is often used more broadly to describe converting analog material into digital form, which may or may not include OCR. In practice these steps are frequently combined in a single workflow, but they are distinct: an object can be scanned without being OCR-processed, and OCR quality does not affect the fidelity of the underlying image.
What factors affect OCR accuracy in a records workflow?
OCR accuracy typically depends on factors such as the quality and resolution of the source image, the condition of the original (for example, faded, skewed, or damaged material), the typeface or whether the text is handwritten, language and character set, and the presence of tables, stamps, or annotations. Depending on the tool and configuration, results can vary considerably. Many organizations assess accuracy against representative samples before relying on OCR output for downstream processes.
How should OCR errors be handled when the output supports retention or disposition decisions?
Because OCR output can contain errors, organizations that use it to inform classification, retention, or disposition often build in verification steps, such as human review of exceptions, confidence-score thresholds, or quality-assurance sampling. Depending on organizational policy and risk tolerance, it may be appropriate to treat OCR-derived text as an aid rather than a sole basis for decisions that carry legal or compliance consequences. The tolerance for error generally varies with the sensitivity and evidential importance of the records.
Should OCR-derived text be retained, and how does it relate to the original record?
Whether to retain the OCR text layer and how to manage it are matters of organizational policy. In many workflows the OCR text is stored alongside or embedded within the image object to enable search and retrieval, while the image remains the reference representation. Clarifying the relationship between the two, which is the record, which is a derived access aid, helps preserve integrity and avoid confusion during retrieval, transfer, or disposition.
What metadata should be captured when OCR is applied?
Depending on the recordkeeping requirements, it can be useful to record process metadata documenting that OCR was performed, such as the tool or process used, when it was applied, and any quality-control outcomes. Capturing this information supports the usability and, where relevant, the demonstrable integrity of the object over time, and can help those relying on the text later understand its provenance and limitations. Specific metadata requirements generally depend on organizational policy and applicable standards.

Common misconceptions

OCR produces a perfectly accurate transcription of the source document.
OCR output typically contains errors that vary with image quality, font, layout, and source condition. For records where content must be relied upon as evidence, the recognized text should be treated as a derived aid rather than an authoritative transcription, and verification is often necessary depending on organizational policy and risk.
Applying OCR to a document creates the authoritative record.
OCR generally adds a machine-readable text layer or a separate text output; it does not by itself establish which item is the authoritative record. The relationship between the source image, the recognized text, and any authoritative record should be defined by organizational recordkeeping practice, and care should be taken to preserve authenticity, integrity, and usability.
OCR and digitization are the same thing.
Digitization refers to converting a physical item into a digital image, while OCR is a subsequent process that interprets that image to produce machine-readable text. A document can be digitized without being subjected to OCR, and the two serve different purposes within capture and processing workflows.

Best practices

Retain the original source image alongside any OCR-generated text, so that the recognized text can be verified against the source and the authenticity and integrity of the record are not compromised.
Improve image quality through appropriate pre-processing, such as deskewing and contrast adjustment, before recognition, since output reliability often depends heavily on the legibility of the source.
Apply quality review proportionate to risk, using confidence indicators to prioritize manual verification for records where accuracy is critical or where the content may be relied upon as evidence.
Define in organizational policy how OCR text relates to the authoritative record, clarifying that recognized text is typically a derived search and access aid rather than a substitute for the source.
Confirm that the OCR configuration supports the relevant languages, scripts, and character sets for the material being processed, and set expectations accordingly for handwriting or degraded documents.
Document the OCR processing applied to a record as part of its metadata, supporting transparency, usability, and the ability to reproduce or reassess the results over time.