All tools

OCR PDF

Make scanned PDFs searchable.

Drop files or click to browse

PDF · up to 100MB per file

About OCR PDF

A scan is a photograph of a document. To a computer it is pixels: you cannot search it, copy from it, or have a screen reader read it aloud. Optical character recognition looks at those pixels, identifies the characters, and attaches a text layer to the page.

The result looks identical but behaves like a real document — searchable with Ctrl+F, selectable, and usable by assistive technology.

How to use OCR PDF

  1. Step 1

    Upload the scanned PDF

    Add a document made of page images.

  2. Step 2

    Recognise the text

    Characters are detected and a text layer is written behind the image.

  3. Step 3

    Download

    Save the searchable PDF.

When it helps

  • Making an archive of scanned invoices searchable by reference number.
  • Copying a quotation out of a scanned book page.
  • Meeting accessibility requirements for published documents.
  • Preparing scans so PDF to Word or Extract Text can work on them.

Good to know

  • Accuracy depends on the scan. Straight, well-lit pages at 300 DPI or better give the best results.
  • Handwriting, decorative fonts, faint carbon copies and heavy skew all reduce accuracy.
  • Always proofread OCR output before relying on it for numbers or legal text.

OCR PDF: frequently asked questions

Working with AI on documents, sensibly

Extraction comes before intelligence

Every AI document feature begins with the same unglamorous step: turning the file into text. A digital PDF exposes its characters directly; a scan must be read by OCR first. Whatever reaches the model is only as good as that extraction, which is why a crisp export and a crooked photograph of the same page produce noticeably different answers.

If a result looks wrong, check the extracted text before blaming the model. Missing columns, merged words and dropped diacritics almost always trace back to the source rather than the analysis.

Context windows and long documents

A language model can only consider a limited amount of text at once. Long documents are therefore handled in sections, with the results combined — a reliable approach for summaries and question answering, and a weaker one for questions that require holding the entire document in mind at the same time, such as counting every occurrence of a term across four hundred pages.

Ask focused questions about specific sections rather than sweeping questions about the whole file, and you will get markedly more dependable answers.

Verify before you rely

Models produce fluent text whether or not they are certain, so a confident summary can still misstate a figure, a date or a negation — the last being especially costly in contracts, where dropping a single 'not' inverts the meaning. Treat output as a well-informed first draft.

Spot-check anything consequential against the source page, and never paste material you are not permitted to share into any AI tool, including this one. Where a document is confidential, prefer the browser-only tools that never transmit it anywhere.

Common problems and fixes

The tool says it found no text
The document is image-only. Run OCR to create a text layer, then repeat the request.
The summary misses an important section
Summaries prioritise recurring themes. Ask directly about the section you care about instead of requesting a general overview.
Numbers or names appear slightly wrong
OCR routinely confuses 0 with O and 1 with l, and models can carry the error forward. Verify figures against the original page before using them.

Terms worth knowing

OCR
Optical character recognition — converting pictures of text into machine-readable characters.
Context window
The maximum amount of text a language model can consider in a single request.
Hallucination
A confident but incorrect statement produced by a language model that is not supported by the source.

Related PDF tools

Continue your workflow with other free ai-powered PDF tools or browse all PDF tools.

Popular guides