OCR PDF
Extract text from scanned or image-based PDFs
Make scanned PDFs searchable with OCR
A scanned document is really just a picture of text — you cannot search it, select it or copy from it. OCR (optical character recognition) reads the characters in the image and turns them into real, selectable text, so a scanned contract, receipt or book page becomes searchable and ready to reuse.
This unlocks old paper archives, lets you find a clause in a scanned agreement instantly, and is the essential first step before converting a scan to Word or Excel.
What decides the accuracy
Upload your scanned PDF and each page is analysed to recognise the text it contains, so you can search and extract it. Recognition quality is determined almost entirely by the scan rather than by the software.
Resolution is the first factor: 300 DPI is the practical minimum for reliable results on normal body text, and accuracy falls away sharply below it as character shapes blur together. Straightness is the second — even a couple of degrees of skew measurably hurts recognition, because the process assumes roughly horizontal baselines. Contrast is the third: crisp black on white is ideal, while grey photocopies, yellowed paper and shadowing from a curved book spine all degrade results.
Typeface matters too. Ordinary serif and sans-serif text around 10 to 12 points is what these systems handle best. Decorative faces, very small print and handwriting are progressively harder, and handwriting in particular is a fundamentally different problem that general-purpose recognition should not be expected to solve.
Errors that a spellchecker will not catch
The failure mode worth understanding is that recognition errors are frequently valid words. Characters that look alike get confused: 'rn' reads as 'm', the digit 1 blurs with lowercase l and capital I, 0 and O swap, 'cl' becomes 'd'. So 'modern' becomes 'modem' and '1,000' becomes 'l,OOO'.
None of these are flagged by proofing tools, because they are real words and plausible strings. In a scanned invoice, a misread digit is both entirely possible and invisible to automated checking.
The practical conclusion is that OCR output is excellent for making a document findable and unreliable as an authoritative transcript. Use it to locate the clause; read the original scan to rely on what it says. For legal, financial and medical documents, the scan remains the record.
Tables, columns and the limits of recognition
Recognition reads the page as lines of text and has limited understanding of layout. Tables often emerge with the relationship between cells lost, and multi-column pages can interleave text from adjacent columns into nonsense.
If you need tabular data from a scan and the figures matter, extracting manually is frequently faster than correcting an automated attempt, and considerably safer.
Before running OCR at all, check whether you need it. If your PDF came from an export rather than a scanner it already contains real text — try selecting a line, and if individual words highlight, the text is there. The same test confirms whether OCR has worked afterwards.
Recognition on your own hardware
Text recognition is performed on your own device, so scanned documents are never uploaded. This is the operation where that matters most, because OCR is overwhelmingly applied to paper records being digitised — medical letters, historical correspondence, financial statements, legal files — and those are precisely the documents that should not be handed to an unknown server.
The honest trade is speed. Running recognition locally is slower than a server would be, particularly across a long document or on an older phone, because the work is done by your own processor. What you get for the wait is that nothing left the machine.
Frequently Asked Questions
Is the OCR PDF tool free?▾
Yes. Running text recognition on your scanned PDF is completely free with no signup.
What does OCR do to my PDF?▾
OCR (Optical Character Recognition) reads the text in scanned images and embeds it into the PDF so you can search, copy, and select it.
What text can the OCR read?▾
The tool recognizes English-language text in scanned PDFs and embeds it so you can search, select, and copy it.
Is my scanned document kept private?▾
Yes. OCR processing happens in your browser — your document is never uploaded to or stored on our servers.