🔍 PDF300
Read the text out of a scanned PDF
A scan looks like a document but it is a picture of one — you cannot select a word or search for a name. OCR looks at the shapes and works out the letters.
Open the toolRuns on our server — the file is uploaded, processed and sent back.
How it works
- Add the scanned PDFThe file is sent to the server, which is where the reading happens.
- Let it readClear, straight scans come out best. A crooked or blurry photo comes out worse.
- DownloadTake the text, or the PDF with the text laid behind the picture.
Questions
Yes — this one runs on a server, so the file is sent over the network and processed there. If you would rather nothing left your device, use the browser tools on the main page instead.
Neither. No account, no watermark on the output. The practical limit is your device memory — large scanned files use more of it.
Guide: doing it well
When you need OCR
A scanned PDF is a picture of a document. It looks like text, but your computer only sees pixels: you cannot search it, copy a sentence out of it, or have it read aloud. OCR (optical character recognition) looks at the shapes of letters and works out what they say. It is what you need when you want to quote from an old scanned contract, search a pile of scanned invoices, or turn a printed page back into editable text.
Two kinds of output
Plain text. The recognised text comes back as a .txt file or can be copied to the clipboard. Use this when you want to edit or reuse the words.
Searchable PDF. The scan stays exactly as it looks, and an invisible layer of recognised text is placed behind each page. You can then search it and select text in any PDF reader. Use this when the document itself matters — signatures, stamps, layout — but you also want it to be searchable.
Step by step
- Choose the scanned PDF in the OCR tool.
- Choose the language.Options are English, Korean, Korean + English, Japanese and Simplified Chinese. Choosing the right language matters a lot for accuracy.
- Run it and check the result.Skim the text for obvious mistakes, especially numbers.
Getting accurate results
OCR is only as good as the scan. A flat, straight page scanned at 300 dpi in black and white or greyscale gives the best results. Phone photos work if the page is lit evenly and fills the frame. Accuracy drops with handwriting, very small print, coloured backgrounds, tables with many lines, and pages that are skewed. Numbers deserve a second look: 0 and O, 1 and l, 5 and S are the usual confusions.
If the PDF already has a text layer (you can select words in it), you do not need OCR — use PDF to Text, which runs in your browser and is exact.
Privacy
OCR runs on our server, not in your browser, because the recognition engine (Tesseract, via OCRmyPDF for searchable PDFs) is too heavy to run in a web page. The file is uploaded, processed in a temporary folder, and that folder is deleted when the response is sent. We do not keep the file or the recognised text. For sensitive documents, consider redacting anything you do not need first.
Common problems
The text is gibberish. The wrong language was selected, or the scan is too low-resolution. Try again with the correct language.
File too large. OCR accepts files up to 30 MB. Split long scans into parts with Split.
Server unavailable. The server tools depend on one machine. If it is down, the browser tools still work; try OCR again later.
More questions
The limit is 30 MB per file. Very long documents take longer; split them if the request times out.
The searchable PDF keeps the scan as it was and adds hidden text behind it.
Poorly. The engine is trained on printed text.
Server tools are free for 5 uses per day. A $0.99 day pass removes the limit for 24 hours.
No. It is processed in a temporary folder that is removed after the result is sent back.