Optical Character Recognition
Optical Character Recognition (OCR) is a technology that converts images of text, such as scanned documents or photographs, into machine-readable text. This allows text that was previously locked inside an image to be searched, copied, or edited. OCR is commonly applied to typed, printed, and in some cases handwritten material.
Optical Character Recognition (OCR) is an automated process that analyzes images of text and extracts the characters into a machine-encoded, machine-readable format. It typically operates on scanned images, photographs, or image-based documents such as non-searchable PDFs, converting typed, printed, or handwritten content into text that can be indexed, searched, copied, or edited. In a recordkeeping context, OCR is a processing step applied to digitized or born-digital image content; the accuracy and completeness of its output can vary with source quality, and organizations should note that OCR output is a derived text layer rather than an alteration of the underlying image, which has implications for authenticity, integrity, and how the digitized record and its text representation are captured and managed.
Why it matters
For records professionals, OCR is often the step that turns a large body of digitized images into content that can be searched, indexed, and retrieved. Without a text layer, a scanned document is effectively an image whose contents cannot be located through full-text search, which limits discoverability and can hamper responses to access requests, litigation-related searches, and routine business use. Applying OCR to image-based material therefore has a direct bearing on how efficiently an organization can find and use its records.
OCR also carries implications for authenticity and integrity that professionals should treat carefully. The text produced by OCR is a derived representation rather than a modification of the underlying image; the original digitized image typically remains unchanged. This distinction matters when the digitized image is intended to serve as the authoritative record, because the searchable text layer is a convenience derived from that image and may contain errors. Depending on organizational policy and the sensitivity of the material, it may be important to document how and when OCR was applied and to retain the source image as the reference point for the record's content.
Because OCR accuracy and completeness can vary with the quality of the source, professionals should not assume that a searchable document contains a perfect transcription of its image. Faint, damaged, skewed, or handwritten source material may yield incomplete or incorrect text. Where OCR output is relied upon for search, redaction, or downstream processing, organizations typically need to account for this variability in their quality controls and in any assurances they give about completeness of retrieval.
Who it's relevant to
Inside OCR
Common questions
Answers to the questions practitioners most commonly ask about OCR.