Extract text from PDF

Extract existing PDF text, or choose local OCR for printed English, Spanish and Simplified Chinese scans and images.

When this tool is useful

Copying readable text from digitally generated reports, articles, invoices, and notes for search or editing.

How local processing works

Basic extraction: The browser reads each page's text items, joins them in an approximate reading order, separates pages, and creates a UTF-8 plain-text download. Your document stays in browser memory. OCR code and the selected language model download from this site; this is not a file upload. Leaving or clearing discards the document and results. These pages are copied without OCR. Turn this off if a scanned page has a text header but the rest still needs recognition.

How to use it

  1. Choose the source files.
  2. Review the available settings and page order.
  3. Process locally, inspect the result, and download it.

Important limitations

Printed text only. OCR can misread characters and does not reproduce editable layouts, tables or handwriting. OCR pages become images with an invisible text layer; links, forms, annotations and signatures are not retained on those pages. This is not PDF/A certification.

Verify the downloaded result

Reopen the downloaded file in its destination application, check dimensions or page count and visible content, and keep the source until the result has been accepted.

Private by design

Your files stay in this browser. The tool does not upload the source or output to our servers.

Questions and answers

Does this tool upload my files?

No. Processing happens in your browser; closing the page clears the working session.

Will the output always match the source exactly?

Printed text only. OCR can misread characters and does not reproduce editable layouts, tables or handwriting. OCR pages become images with an invisible text layer; links, forms, annotations and signatures are not retained on those pages. This is not PDF/A certification.

What should I check before using the result?

Reopen the downloaded file in its destination application, check dimensions or page count and visible content, and keep the source until the result has been accepted.

Which browsers and devices are supported?

Use a current browser with JavaScript, workers, canvas, WebAssembly where the format requires it, Blob downloads, and enough memory. Reopen important downloads in the destination application.

Concrete input and output example

Extract selectable text from a six-page digital report with the basic mode. For an image-only English, Spanish or Simplified Chinese scan, explicitly choose OCR, download the language model and compare the text plus searchable PDF against the original. Files are not uploaded.

Browser compatibility notes

Reviewed against the implementation and required web APIs on 2026-09-02. Use a current browser where these APIs are available and verify important output on your device.

Chrome / Edge (current desktop)
Requires local PDF parsing, worker/canvas APIs where applicable, document creation, and Blob download.
Firefox (current desktop)
Uses the local PDF workflow; always verify the generated download in its destination viewer.
Safari (current desktop)
Uses the same local workflow; large page sets can have a lower practical memory ceiling.

A tool-specific question

Can it read text from a scanned PDF?

Yes, through the separate on-demand OCR mode for printed English, Spanish or Simplified Chinese. Basic extraction still reads existing text only. OCR is imperfect, and recognized pages become images with an invisible text layer rather than editable original layouts.

Tool guidance reviewed

Related tools