Two kinds of PDF that look the same
A PDF exported from a word processor, accounting package or spreadsheet usually stores each character as text placed at a position on the page. A PDF made by scanning paper or photographing it usually stores one picture per page. Both look the same when you open them. Only the first can be searched, selected and read by a program as text. Some files mix the two, such as a digital cover page followed by scanned appendices.
A PDF does not store a table as a table. It stores pieces of text at positions, sometimes with ruled lines drawn around them. Tools that extract tables have to infer rows and columns from those positions or lines. The pdfplumber project describes exactly this: it looks for explicit or implied lines, finds where they cross, and builds cells from the intersections.
- Text-based PDF: you can select a line, search for a word, and paste readable text elsewhere.
- Scanned PDF: selecting drags a box around the whole page, and searching finds nothing.
- Mixed PDF: test the first, middle and last pages separately and write down which page numbers are images.
A two-minute test that sends nothing anywhere
Open the file in your usual PDF reader. Search for a word you can plainly see on the page. Then drag across a line of text, copy it and paste it into a plain text editor. Try this on at least three pages, including the last one.
Do this locally. Free online converters ask you to upload the file, which means sharing the document with a third party. If the document is not yours to share, that is a problem before any extraction has started.
- Readable pasted text with numbers as digits means a text layer exists.
- A blank paste, or characters that are nonsense, means there is no usable text layer on that page.
- Note the page numbers that fail. A mixed file needs two methods, and the failed pages are a separate job.
What each result means for the method
With a text layer, a table can usually be read from the positions of its text. Microsoft documents a PDF connector for Excel that returns the tables it finds, with options for page ranges and for whether similar tables on consecutive pages are combined. The same page says rows that spread over several lines may not be identified properly and may need cleaning afterwards. The pdftotext manual page describes a layout option that tries to keep the original physical arrangement of text.
Without a text layer there is nothing to extract yet. A text-recognition step must first create text from the page images, and its quality depends on the scan. Tesseract's documentation says it works best on images of at least 300 dpi and that heavy skew, noise and dark scanner borders reduce quality. That is why a scanned document cannot honestly be given a promised accuracy before anyone has tried it.
These tools are named only because their documentation explains how PDFs behave. What a buyer receives is a checked spreadsheet, not a tool.
Look-alike problems that are not a missing text layer
Sometimes text can be selected but pastes as nonsense. The pdftotext manual page says text in PDFs with badly corrupted font encodings cannot be extracted short of text recognition. Sometimes the text is real but pastes in a scrambled order, because reading order differs from the visual layout. And sometimes a digital PDF contains a table that was inserted as a picture, so the page has text but the table does not.
Password protection and restricted copying are separate problems. If a file refuses to open or copy, ask the person who issued it before trying to get around it.
- Nonsense characters: the font encoding is the problem; treat the page like a scan.
- Scrambled order: the text exists but the layout needs rules; extraction can still work.
- A table pasted as a picture: no table text exists for that table, even though the page has other text.
Where the paid jobs fit, and what to send first
One text-based price list or table of up to 60 pages in a single layout can be turned into a spreadsheet by the one-off extraction job, from £295 after we have seen a redacted sample page and agreed the columns. It is accepted by page-by-page row counts, a sample compared field by field and an exceptions list, not by a promised accuracy. A scanned batch of up to 500 pages can be made searchable by a separate job from £195, accepted by a search test of terms you choose. Both prices are untested proposals, and payment follows the agreed checks and your sign-off.
For a first enquiry, send the page count, whether the text could be selected, how many different layouts the document has and your wanted column headings. Do not send the document, and never one with personal or confidential data. Secure handling is agreed after scoping.
- This guide does not establish that anyone has requested or paid for these jobs; they are new offers.
- Neither job is legal, financial or tax advice, and neither promises that every row or word is read correctly.
Sources and limits
- pdfplumber README: table extraction and its limits Checked 2026-10-11.
- The project says it works best on machine-generated rather than scanned PDFs.
- Its table finder looks for explicit or implied lines, finds their intersections and builds cells from them; line detection can use lines, text or explicit strategies.
- It offers no text recognition and lists weak support for tables from OCR'd documents as a missing feature.
- Power Query PDF connector (Microsoft Learn) Checked 2026-10-11.
- Pdf.Tables returns any tables found in a PDF; options include a start and end page, combining similar tables on consecutive pages and enforcing border lines as cell boundaries.
- Where multi-line rows are not identified properly, the page says the data may need cleaning with further steps in Power Query.
- To import several PDF files at once, it points to a multi-file connector such as the Folder connector.
- pdftotext manual page (Poppler utilities) Checked 2026-10-11.
- The -layout option tries to maintain the original physical layout of the text; -f and -l set the first and last page.
- Its bugs section says text in PDFs with badly corrupted font encodings cannot be extracted short of text recognition.
- Tesseract documentation: improving the quality of the output Checked 2026-10-11.
- Tesseract works best on images with a resolution of at least 300 dpi.
- Line segmentation quality falls significantly when a page is too skewed; noise and dark scanner borders can also reduce accuracy.