Scanned documents
Recognize scanned PDFs and the limits of OCR
A scanned PDF may have no usable text layer. Basic extraction reads existing text; the separate local OCR mode recognizes printed English, Spanish or Simplified Chinese from scans and images when you start it.
Workflow
Test what the PDF contains
Try selecting a known sentence and searching for a visible word. If both fail, the page may be image-only or have a broken or hidden text layer.
Choose the local OCR mode when needed
Correct orientation, keep readable resolution, and select the document language explicitly. Existing-text pages are skipped by default; disable this for a scanned page with only a text header. Then download text and the searchable PDF.
Proofread against the scan
Check names, dates, totals, decimal points, tables, handwriting, page order, and low-contrast areas; keep the scan as the visual record.
Limits to understand
OCR can confuse similar characters, reorder columns, miss handwriting, invent spaces, or omit faint text. A searchable layer is not proof of accuracy, authenticity, accessibility, or legal equivalence. Printed text only. OCR can misread characters and does not reproduce editable layouts, tables or handwriting. OCR pages become images with an invisible text layer; links, forms, annotations and signatures are not retained on those pages. This is not PDF/A certification.
Privacy mode
Your document stays in browser memory. OCR code and the selected language model download from this site; this is not a file upload. Leaving or clearing discards the document and results.
Verify the result
Compare a sample from every layout type, search for expected and deliberately difficult terms, copy figures into plain text, and have a person verify information used for decisions or compliance.
Frequently asked questions
Why does PDF to text return nothing?
Basic extraction only reads an existing text layer. If the page is image-only, choose the separate OCR mode; recognition can still miss faint, blurred or unsupported text.
Does converting a scan to JPG help OCR?
It can provide a compatible input, but unnecessary downscaling or compression can make recognition worse. Preserve a high-quality source.