Job pdf-scanned-batch-to-searchable-text · revised 11 October 2026
Make a batch of scanned PDFs searchable and test it with terms you choose
Image-only PDFs in one folder get an added text layer, with originals kept. A search test of terms you choose is held to a floor agreed in writing beforehand, and poor pages are listed.
You might be seeing
- A search for a word you can see on the page returns no result
- You cannot select or copy any text from the PDF
No passwords, keys, card details or admin invites needed to start.
What usually happened
Scanning a paper document produces page images. Without a text layer, a PDF cannot be searched, selected or indexed. Text recognition can add one, but its quality depends on the resolution, skew, noise and language of the scan, and it can misread characters. A folder can look searchable while the words people look for are missing from the layer.
Who it’s for: Someone responsible for an office archive of printed business documents that has already been scanned to PDF but cannot be searched, and who needs to find documents by their words.
Usually starts when: People open scan after scan looking for one name or phrase, because searching the folder finds nothing inside the files.
The result: You receive a searchable copy of each PDF in the folder, with the original untouched, and a report on the search terms you chose: which pages they were found on, which expected pages missed them, with the page image shown for every miss, and which pages had poor recognition. The report is held to a pass floor agreed in writing before processing, so you know what a search can and cannot find.
Check whether this job fits
Answer these without sending any document. Nothing is submitted unless you choose to contact us.
Checks you can run yourself
Check whether the PDFs hold any text
Open one PDF, press the search shortcut and type a word you can see on the page. Then try to select a line of text with the mouse.
Look for: No result and no selectable text mean there is no text layer. A result on some pages only means the batch is mixed.
What you get
- A searchable copy of each PDF in the batch, with the original kept
- A scan assessment that classes each page as clean or other, and the pass floor agreed from it
- A search-test report: each term, the page you expected, whether it was found, and the page image for every term that was not
- A list of pages with no or poor recognised text, with a reason where we can give one
Included
- Up to 500 pages of already-scanned PDFs in one or two named languages, with no personal or confidential content; we do not scan paper
- Assess the scans for resolution, skew and blank pages before processing, class each page as clean (printed text, upright within an agreed angle, at least 300 dpi, not faded) or other, and tell you what is likely to recognise poorly
- Agree a pass floor in writing after that assessment and before processing
- Add a text layer to copies, keeping the originals, in an archival PDF format if you ask for it
- Run a search test with 15 to 20 terms you know appear on known pages, and list pages where recognition was poor
Not included
- Scanning paper, or rescanning poor pages
- Recognising handwriting
- A promised recognition accuracy for the batch as a whole, or retyping and correcting the recognised text
- Redaction, legal admissibility, or advice on records and retention
- Extracting data into a spreadsheet
- Documents holding personal, confidential or regulated information
How we know it’s done
Agreed with you before work starts. Each check produces evidence you keep.
Every input PDF has a searchable output copy, or is listed with the reason it has none, and the page count of each output equals that of its original.
Evidence: The file list with page counts and exceptions.
For each of your 15 to 20 agreed search terms, the report states whether the term was found on the page where you know it appears and shows the page image for every term not found. The result meets the pass floor agreed in writing before processing: every term whose known page the assessment classed as clean is found on that page, and at least the number of terms agreed in writing for the other pages is found.
Evidence: The scan assessment, the agreed floor, and the search-test report with term, expected page, found or not found and the image for each miss.
Pages whose recognised text is empty or below the agreed threshold are listed with their file and page number.
Evidence: The list of poor pages and the agreed threshold.
The original files are identical before and after the job.
Evidence: Checksums of the originals taken before and after.
Sign-off. You run some of your own searches in the copies, read the report and check the page images for any misses, then sign off in writing. Payment follows sign-off.
If it fails. If the agreed checks do not pass, including the pass floor, you do not pay for this fixed scope. If the scans prove too poor or the documents turn out to be handwritten or sensitive, we explain why, hand back what we have and agree whether to stop. No surprise work.
When it fits, and when we stop
It fits when
- The files are PDFs of scanned pages, typed or printed text rather than handwriting, in one or two named languages
- You can supply 15 to 20 search terms and the pages where they are known to appear
- The documents hold no personal, confidential or regulated information, and you may share them with a contractor
We stop and tell you if
- The pages are mainly handwriting, or the scans are too poor to read
- The documents contain personal, confidential or regulated information
- You want a guaranteed recognition accuracy
- The scan assessment shows that no honest pass floor can be set; we say so and nothing is charged
What could go wrong
The originals are never changed. The searchable copies are separate files; delete them if you do not want them.
Scroll the table sideways to read it all.
| Risk | How we handle it |
|---|---|
| A word is misread, so a search for it misses the page. | The report says so plainly and shows the page image for each miss. The search test shows which of your known terms are and are not found against a floor agreed beforehand, and poor pages are listed. |
| The text layer looks fine but is empty on some pages. | Pages with no recognised text are listed, never left looking searchable. |
| Sensitive documents are exposed. | Documents with personal or confidential content are excluded outright. You confirm in writing that there are none, no documents go in the first enquiry, and a secure route and deletion plan are agreed in writing before any file moves. If such content appears we stop and agree with you how the files are returned or deleted. |
A reviewer who did not process the batch re-runs the search test with other terms you supply and spot-checks pages with poor recognition. You decide whether the result is good enough for your purpose.
How we deliver
We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.
- Agree the languages, output format and the search-test terms in writing
- Receive the PDFs by the agreed secure route, inspect the scans for resolution, skew, noise and blank pages, class each page as clean or other and report likely problems
- Agree the pass floor with you in writing before any processing
- Add a text layer to copies of the PDFs, never to the originals
- Run the search test, show the page image for every miss and list pages with no or poor recognised text
- Verify that page counts match and that the originals are unchanged; have an independent reviewer repeat the search test on different terms
- Hand over the copies and the report
This is a one-off job, not a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after the agreed files, the secure handling route and the acceptance checks are written down. An enquiry creates no charge or booking. We give no legal, medical, financial or tax advice.
Need to keep it working?
If new scans arrive regularly, making each batch searchable can be agreed as a recurring service. That is a separate quoted agreement.
Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.
Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.
What you can check
This is a new service. We have not delivered this job for a client yet.
Other ways to get this done
- Many scanners and PDF programs have a built-in recognise-text option. For a short document, running it yourself and testing the search may be enough.
- Scan quality matters most. The Tesseract documentation says it works best on images of at least 300 dpi, so rescanning a few poor pages may help more than processing them again. tesseract-ocr.github.io
Questions
How accurate will the recognised text be?
We do not promise an accuracy. We test the terms you choose, show every miss with its page image, and list pages that recognised poorly. Before processing we agree in writing a pass floor, and you do not pay if the result falls short of it.
What if the scans are too poor for any floor to be set?
We tell you after the scan assessment, before any processing, and nothing is charged.
Can you read handwriting?
No. Handwritten pages are outside this job and will be listed.
Do you scan the paper?
No. We work on PDFs you have already scanned.
Will it be in PDF/A?
If you ask for it, yes. A level B PDF/A file made from page images does not necessarily support indexing of its text, which is why we run a search test whichever format you choose.
Send an enquiry
Send us
- The number of pages, the language, the era or type of the documents and the resolution of the scans if known
- Whether you can select any text in the PDFs now
- Five or six words you would search for; no documents in the first enquiry
Later, once you agree
- The PDFs through the secure handover route we agree in writing before any file moves, with your confirmation that they hold no personal or confidential information and that you may share them
- A list of 15 to 20 search terms and a page where each is known to appear
- Your agreement to the pass floor, once you have seen the scan assessment, and a named person to review the report and sign off
You keep ownership of every file we work on and of what we return. No upload portal exists yet, and the first enquiry never includes documents: send a description, counts and an invented or redacted sample only. Work can start only after we have agreed with you, in writing, a secure way to hand the files over, who may see them, how long we keep them and how we delete them. If we cannot offer a way you are comfortable with, we decline and nothing is charged. Files may be handled by automated software, including services run by other companies, and we tell you which in writing before any file moves. We work on copies, never on your only original, and we do not send, publish or contact anyone for you.
Email fallback: open your mail app
If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “pdf-scanned-batch-to-searchable-text” as the subject.