OCR a Scanned PDF Online
Upload a scanned PDF, run local OCR, then search within your session or export a copy with a searchable text layer.
Drag and drop a PDF file here
or
Files are processed in your browser and are not uploaded to our servers.
When to use OCR PDF
Use this when a PDF is a scan or photo of a document and its text can't be selected or searched. Running OCR first is also the prerequisite for using Search PDF on that same file.
How to OCR a Scanned PDF
- Select a scanned or image-based PDF.
- Choose the document language, if more than one is available.
- Run OCR — it processes locally in your browser.
- Review or search the recognized text, then export a searchable copy.
What this tool changes
- Adds a hidden, searchable text layer behind the existing scanned page images (when exporting a searchable copy).
What it preserves
- The visible page images are unchanged.
Good to know
- OCR accuracy is reduced for low-resolution scans, handwriting, rotated pages, mixed languages, complex tables, and decorative fonts.
- Initial language support is English; additional languages are added only once verified reliable.
- OCR runs entirely on your device, so processing time depends on your hardware.
What recognition adds, and what it leaves alone
Recognition does not replace your scanned image with text. The photograph stays exactly where it is, pixel for pixel. What gets added is a second layer of real characters, positioned over the words in the picture and rendered completely invisibly.
So the finished document holds two things on every page: the picture you see and an invisible transcript you do not. Search matches the transcript; your eyes read the picture. Because the visible page is untouched, recognition cannot make your document look worse — a misread word is a wrong entry in an index, not a wrong word on the page.
It also explains the quirk people notice first: selecting text on a recognised scan often highlights slightly off from the printed words, because the invisible layer is positioned from recognition bounding boxes and those are close rather than perfect.
Expect very good, not perfect
On a clean, straight, 300 DPI scan of ordinary printed text, accuracy is high enough that search works reliably and errors rarely surface. It is not transcription, and treating it as though it were is how people get burned.
What degrades it, roughly in order: low resolution, skew, poor contrast, unusual or very condensed typefaces, complex multi-column layouts, and scanner noise. Handwriting is not attempted — print only. Classic confusions are 0 for O, 1 for l, and rn read as m.
Never rely on unchecked recognition for numbers that matter — account numbers, dosages, reference codes, monetary amounts. Read those off the picture with your own eyes.
Prepare the pages, and mind the order
Almost all of the quality is decided before recognition starts. Scan at 300 DPI in greyscale; below that, character shapes lose the detail that separates similar letters. Straighten the pages first, because a sideways page produces near-total nonsense. Crop scanner edge shadows, which are noise the engine will try to read as characters.
If you are also compressing, compress first and recognise second. Compression re-renders pages as images and would discard a text layer you had just created.
Recognition runs over every page of the document, not only the picture ones — on a mixed file, pages that already carry real text receive a second layer over the top. Usually harmless, occasionally the cause of duplicate matches. It also runs on your own hardware, so a long document takes real time; leave the tab open and work in batches of thirty or forty pages.
The text layer is invisible on purpose
Recognition does not replace your page or redraw it. It reads the image, works out where each word sits, and writes those words back onto the page as text drawn at zero opacity — positioned over the picture of the same words, and invisible.
That is why an OCR’d scan looks exactly as it did before while suddenly answering to search, selection, and copy. Your page is untouched; what you gain is a layer nobody sees. It also explains a small oddity worth expecting: selecting text on a recognised scan sometimes highlights a slightly different rectangle from the printed word, because the invisible text is placed to match the recognised box rather than the ink.
The text is written in Helvetica, which is referenced by name rather than embedded, so the layer costs very little and renders nowhere. What it costs in accuracy depends entirely on the scan: sharp, straight, well-lit pages at a sensible resolution recognise very well, and a crooked or low-resolution one does not. Straightening and cropping before recognition, not after, is the difference that matters most.
Frequently asked questions
- Is OCR perfectly accurate?
- No OCR engine is perfect. Recognition quality depends heavily on scan resolution, font, and layout. Always review recognized text for anything important.
- Where does OCR processing happen?
- Locally, in a background worker in your browser. Page images are never uploaded.
- Can I cancel OCR partway through?
- Yes — cancelling stops the worker and keeps whatever pages were already recognized.
- Will my scan look different after recognition?
- No. The recognised words are written onto the page at zero opacity, over the image of the same words. The page renders exactly as before; only search, selection, and copy change.
- Should I rotate or crop before or after running OCR?
- Before. Recognition on a sideways page produces near-nonsense, and cropping afterwards can clip the region the text layer was positioned against. Straighten, crop, then recognise.