OCR (optical character recognition)

Definition

OCR converts scans, photos and PDFs into machine-readable text. It is the entry step of document automation and delivers characters only — the meaning, validation and posting logic are added by the layers above it.

Modern OCR handles layout, tables and mixed languages reliably, but quality still depends on the source: resolution, contrast and scan orientation.

Treat OCR output as raw material that must be validated, never as a verified value.

In practice

  • Quality depends on source resolution and layout
  • Output requires validation before it is posted anywhere
  • One building block of intelligent document processing

Related terms

Back to the glossary