Scanned documents

Recognize scanned PDFs and the limits of OCR

A scanned PDF may have no usable text layer. Basic extraction reads existing text; the separate local OCR mode recognizes printed English, Spanish or Simplified Chinese from scans and images when you start it.

Workflow

  1. Test what the PDF contains

    Try selecting a known sentence and searching for a visible word. If both fail, the page may be image-only or have a broken or hidden text layer.

  2. Choose the local OCR mode when needed

    Correct orientation, keep readable resolution, and select the document language explicitly. Existing-text pages are skipped by default; disable this for a scanned page with only a text header. Then download text and the searchable PDF.

  3. Proofread against the scan

    Check names, dates, totals, decimal points, tables, handwriting, page order, and low-contrast areas; keep the scan as the visual record.

Limits to understand

OCR can confuse similar characters, reorder columns, miss handwriting, invent spaces, or omit faint text. A searchable layer is not proof of accuracy, authenticity, accessibility, or legal equivalence. Printed text only. OCR can misread characters and does not reproduce editable layouts, tables or handwriting. OCR pages become images with an invisible text layer; links, forms, annotations and signatures are not retained on those pages. This is not PDF/A certification.

Privacy mode

Your document stays in browser memory. OCR code and the selected language model download from this site; this is not a file upload. Leaving or clearing discards the document and results.

Verify the result

Compare a sample from every layout type, search for expected and deliberately difficult terms, copy figures into plain text, and have a person verify information used for decisions or compliance.

Frequently asked questions

Why does PDF to text return nothing?

Basic extraction only reads an existing text layer. If the page is image-only, choose the separate OCR mode; recognition can still miss faint, blurred or unsupported text.

Does converting a scan to JPG help OCR?

It can provide a compatible input, but unnecessary downscaling or compression can make recognition worse. Preserve a high-quality source.

Editorial review