Make your scans searchable

Straighten pages, make them easier to read, and recognise the printed text — so a scanned document can be searched and copied from like any other. Everything happens on this device.

HomePDF Tools › Searchable Scan
Your scan stays on this device
What that means

Your files are read, cleaned and recognised by code running in this browser tab. The recognition engine and its language packs are served from this address, not from anyone else's network, and are fetched only when you choose a language. There is no upload, no server and no analytics here. Nothing is stored unless you download it.

  1. Add pages
  2. Clean up
  3. Recognise
  4. Review & download

Add your pages

The sample is an invented receipt, with a known transcription to compare against.

What this tool does, and what it does not

How your scan is treated
  • The text is real. It is written into the PDF as invisible text positioned over the words it came from, so a reader can search it, select it and copy it. It is not a separate text file shown next to an unchanged image.
  • Your original file is never modified. A page is the file you gave plus a list of changes, and that list can always be set back to nothing.
  • A page that already has text keeps it. Adding recognised text on top would make every search find the page twice and copying produce nonsense. Such pages are kept as they are unless you deliberately rebuild them.
  • Untouched scans keep their own pixels. If you have not cleaned a page, the original page is kept byte for byte and the text is appended to the file, so the scan looks exactly as it did.
  • The output is checked before you get it. The finished PDF is reopened and its text read back page by page. If a page does not verify, it says so instead of showing a tick.
Limits, honestly
  • Printed text only. Handwriting, equations, decorative lettering, vertical text and very low-resolution photographs will mostly fail. Nothing here invents missing words or tidies up the recognition with a language model.
  • Recognition makes mistakes. “Text layer verified” means the words are present and searchable in the file. It does not mean they match your scan. Check anything that matters.
  • Six languages: English, Spanish, French, German, Italian and Portuguese, all printed. The text layer uses the fonts built into every PDF reader, which cover those alphabets; a language outside them would need a font embedded in the file, which this does not yet do.
  • Memory, not file size, is the limit. An A4 page at 300 dpi is about 9 megapixels and one copy of it is 35MB in memory. Pages are worked on within a budget and scaled down if they exceed it; the page says when that happened.
  • Rebuilt pages are flattened. Once you crop, straighten or clean a page it is rebuilt from what you can see, which does not carry over form fields, links, attachments, layers or tags. Signatures on a rebuilt page will not survive, and nothing here validates or creates them.
  • This is not redaction. Cropping moves the edge of what is drawn. Do not use it to remove something confidential.
  • Nothing is saved between visits. Download before you close the tab.

Already have a PDF with text in it? PDF Compare finds what changed between two versions, and PDF Form Builder adds fillable fields.