Practical workflows

Make a scan searchable with OCR

Recognize text in a scan locally and download plain text plus a searchable PDF for later retrieval.

Starting parameters

Choose English, Spanish or Simplified Chinese recognition. Limits include 20 MiB input, ten PDF pages and 12 megapixels per source image. Existing-text pages are skipped by default.

Processing steps

  1. Load a scan or the synthetic text sample. Select the language printed in the document, not necessarily the interface language.
  2. Start OCR; the selected model downloads on demand. If a page has only a text header above a scan, turn off skipping existing-text pages and retry.
  3. Inspect the recognized text, then download the text file and searchable PDF. Search for a known phrase in an independent PDF reader.

Check the result

Compare names, numbers and important passages with the image. A confidence estimate is not proof of accuracy.

Limits and losses

OCR pages become raster images with invisible text. Original editable layout, forms and links are not preserved; this is not certified PDF/A archival conversion.

Local processing

Files and outputs stay in this browser; originals are not overwritten. Samples contain invented data. Download results before clearing or leaving. OCR downloads its recognition model, not your document.

Reviewed