What scanning produced, and what recognition adds
A scanner or a phone camera produces an image of each page. The PDF wraps the images, so it looks like the paper but holds no text. Text recognition reads the images and creates characters that are stored in an invisible layer behind the picture. The OCRmyPDF documentation describes this: it adds text layers to images in PDFs, making scanned image PDFs searchable, and can produce a minimally altered PDF.
A minimally altered PDF is not the same as an untouched one. The same documentation says that its image-processing options, such as deskew, integrate the text layer into the processed image, that it generates PDF/A-2b output by default, and, in its limitations section, that when Ghostscript is used it may transcode some images, potentially lossily. So the aim is a hidden text layer behind the page pictures, but the pictures in the new file can differ from the originals. That is one reason the originals are kept untouched and the searchable files are separate copies.
What decides how good the text is
Recognition quality depends on the input. The Tesseract documentation says it works best on images with a resolution of at least 300 dpi, that line segmentation falls significantly when a page is too skewed, and that noise and dark borders from the scanner can make it read extra characters. OCRmyPDF adds that its output quality depends on the quality of the input, that its accuracy may not match commercial solutions and that it cannot recognise handwriting.
In practice a faint photocopy, a crooked page, a small font or a stained page will recognise worse than a clean, straight, high-resolution scan. If you can still rescan, the cheapest improvement is usually a better scan, not a smarter program.
- Resolution: very low resolution scans lose small characters.
- Skew: a crooked page upsets the program's idea of where the lines are.
- Noise and borders: specks and scanner edges can turn into false characters.
- Handwriting: not recognised reliably, so a handwritten page is outside what this can honestly promise.
Test it with words you already know
Do not ask for, or accept, a single accuracy percentage for a mixed folder. Instead pick 15 to 20 words or numbers that you know appear on particular pages: a surname, a reference code, a place name, a hyphenated word, a figure. Vary them on purpose: some in headings, some in small print, some on the worst scans. After processing, search for each one and record whether it was found on the page where you know it appears.
Report the result as counts: found, not found. Also list pages where the recognised text is empty or very sparse. This tells you what a search can and cannot be trusted to find, which is the thing you actually need to know.
Archival format is not the same as searchable
People often ask for PDF/A to make an archive searchable. PDF/A is a format for long-term preservation: the Library of Congress description says it aims to preserve a document's static visual appearance over time. It also says that level B files made from scanned page images do not necessarily support indexing of the document text. A PDF/A file of scans can therefore be a perfect archive of pictures and still be unsearchable.
Ask for an archival format if your archive policy needs one, but still run the search test, whatever the format.
Where the paid job fits, and where it does not
Making up to 500 already-scanned pages with no personal or confidential content searchable is a one-off job from £195 (an untested proposal). It is accepted by a count of every file with a searchable copy or a stated reason, a search test on your terms held to a pass floor agreed in writing before processing (every term on a page classed clean in advance must be found, and every miss is shown on the page image), a list of poor pages, matching page counts and originals that are unchanged by checksum. Payment follows your sign-off. No accuracy is promised for the batch as a whole.
It does not scan paper, read handwriting, correct the recognised text or handle documents containing personal or confidential information. If the real goal is a spreadsheet of values from the scanned pages, that is a different job with different checks. The first enquiry needs a page count, the language and a few words you would search for, never the documents. Work starts only after a secure way to hand the files over has been agreed in writing; no upload portal exists yet.
Sources and limits
- OCRmyPDF introduction Checked 2026-10-11.
- It adds text layers to images in PDFs, making scanned image PDFs searchable, and can produce a minimally altered PDF.
- Its image-processing options, such as deskew, integrate the OCR layer into the processed image; it defaults to PDF/A-2b output; its limitations section says that when Ghostscript is used it may transcode grayscale and colour images, potentially lossily.
- It says its accuracy may not match commercial solutions, says output quality depends on input quality and says it cannot recognise handwriting.
- Tesseract documentation: improving the quality of the output Checked 2026-10-11.
- Tesseract works best on images with a resolution of at least 300 dpi.
- Line segmentation quality falls significantly when a page is too skewed; noise and dark scanner borders can also reduce accuracy.
- Library of Congress: PDF/A, PDF for Long-term Preservation Checked 2026-10-11.
- PDF/A is intended for long-term preservation of page-oriented documents and preserves their static visual appearance.
- Level B files made from scanned page images do not necessarily support indexing of the document text.