OCR PDF

Make a scanned PDF searchable and selectable. The OCR engine and language model run entirely on your device.

Drop a scanned PDF here

or click to browse

Select file
PDF
Your file data uploaded

0 bytes

Processed on this device
0 bytes
Files opened here
0
Requests carrying a file
0

Something not working as it should? Report an issue with OCR PDF

About OCR PDF

A scanned PDF is a stack of photographs of paper. It looks like a document, but searching it finds nothing and selecting text is impossible, because as far as any software is concerned there is no text in the file at all.

OCR fixes this by recognising the characters in each page image and adding an invisible text layer positioned over them. The page looks exactly the same, but the document becomes searchable, selectable and copyable. That is why an OCR'd file opens identically in every reader while suddenly responding to a search.

Recognition runs on your device using Tesseract compiled to WebAssembly, with support for more than a hundred languages. This is the part that usually forces a compromise: OCR is computationally heavy, so most online services upload your documents to their servers. Old bank statements, medical records and years of personal correspondence are exactly the material people least want to hand over, and are exactly what tends to need OCR.

Accuracy depends mostly on the scan. Clean, straight pages at three hundred DPI typically reach the high nineties. Skewed, low resolution or heavily marked pages do worse, and handwriting is not reliably recognised by this engine. Pages process in parallel using Web Workers, so a fifty page document usually takes a few minutes.

How to ocr pdf

  1. Load the scanned PDF

    Drop the file in. Nothing is transmitted.

  2. Choose the language

    Select the language of the document so the right recognition model is used. This matters more than most settings for accuracy.

  3. Run OCR

    Pages are processed in parallel. Progress is shown as it goes.

  4. Download

    You get a standard PDF that looks unchanged but is now fully searchable.

Frequently asked questions

Does OCR change how my document looks?

No. The original page images are untouched and the recognised text is added as an invisible layer behind them. Visually the file is identical.

How accurate is it?

Clean printed scans at three hundred DPI usually reach the high nineties. Poor scans, unusual fonts and skewed pages reduce that. Handwriting is not reliably recognised.

Which languages are supported?

Over a hundred, including English, Spanish, French, German, Portuguese, Russian, Arabic, Hindi, Chinese, Japanese and Korean. Choosing the correct language noticeably improves accuracy.

How long does it take?

Roughly two to five seconds per page on a modern device, so a fifty page document takes a few minutes. It runs on your processor rather than a server, so speed depends on your hardware.

Is my scanned document uploaded?

No. The recognition engine and language data are downloaded to your browser and the processing happens there. You can watch the Network tab during a run and see that no page ever leaves.