Synthetic Industry

Collection · updated 2026-10-11

From a PDF to a checked spreadsheet: the order of questions that saves the most effort

Seven steps from "does it have text?" to "who signs off the exceptions?", linked to the guides, the worked example and the checked jobs, including what to do when the same document arrives every month.

1. Does the PDF have text?

Search for a word you can see and try to select a line. If neither works, the pages are pictures, and everything below waits for a text layer. The text-layer guide gives the test. A scanned batch can be made searchable by a separate job, from £195 for up to 500 pages, with no accuracy promised.

  • Next if there is text: go to step 2. If not: the scanned-PDF guide, then return here.

2. Is it a table or a filled form?

A table or list is read from positions of text or ruled lines. A returned fillable form is read from its form fields. The two have different failure modes and different checks. Microsoft's PDF connector and the pdfplumber project describe the first, and the pypdf documentation describes the second.

3. One layout or several?

One set of rules covers one layout. A list whose sections differ, or forms from several template versions, needs rules for each, and the price changes. Note the page ranges or versions before asking anyone for a quote.

4. Decide the columns and the rules for unusual entries

Write the output columns, then say what happens to "price on request", wrapped lines, checkboxes, blanks and repeated headings. These decisions are the owner's. An extraction without them will make them silently.

5. Reconcile, then sample

Count rows per page or files against rows plus exceptions, then compare a sample you chose first, field by field. The reconciliation guide explains the method and the worked example shows a total that matched while two pages did not. Report counts, never a percentage accuracy.

6. List what could not be read, and decide it

Rows or files that could not be read with confidence go on an exceptions list with their page and reason, not into the data. Agree beforehand how many exceptions would make you stop and re-scope. A named person decides each one.

7. If it arrives every month

Write the rules down once and reuse them. The one-off jobs are from £295 for one list of up to 60 pages and from £245 for up to 200 forms. A standing monthly service from £395 a month extracts each batch by the same rules, reconciles it and reports any layout change or image-only page instead of guessing. Prices are untested proposals and payment follows agreed checks and your sign-off. No job gives legal, medical, financial or tax advice, and the first enquiry never includes confidential or personal documents.

Sources and limits

  • Power Query PDF connector (Microsoft Learn) Checked 2026-10-11.
    • Pdf.Tables returns tables found in a PDF; where multi-line rows are not identified properly the data may need cleaning, and similar tables on consecutive pages are combined by default.
  • pdfplumber README Checked 2026-10-11.
    • Table extraction works best on machine-generated PDFs and the project offers no text recognition.
  • pypdf documentation: interactions with PDF forms Checked 2026-10-11.
    • Form field values are read by field name; flattening turns field contents into regular page content; an XFA entry can override the page content.
  • OCRmyPDF introduction Checked 2026-10-11.
    • It adds text layers to scanned image PDFs, depends on input quality and cannot recognise handwriting.