Synthetic Industry

Job csv-delimiter-quoting-embedded-newlines-import · revised 11 October 2026

Stop supplier CSV rows splitting at commas, quotes and line breaks

A named supplier CSV with commas, quotes and line breaks in descriptions imports at the right column count on every row. Bad rows are listed by line, and a file with an unclosed quote is rejected.

You might be seeing

  • A product's price or stock figure appears in the wrong field for some rows only
  • The importer's row count differs from the number of products the supplier says the file holds

No passwords, keys, card details or admin invites needed to start.

What usually happened

The importer splits records at every comma or line break instead of reading quoted fields, or assumes a comma where the file uses a semicolon or tab. Rows with punctuation inside a description then push their values into the wrong columns, or one product is read as two, while ordinary rows import normally and hide the fault.

Who it’s for: Operations or purchasing manager at a distributor whose supplier CSV imports correctly for most products but scrambles a few.

Usually starts when: A price or stock figure lands in the description, or the importer counts more or fewer products than the supplier states, after a description contains a comma, a quote mark or a line break, or after the supplier switches to semicolons.

The result: A synthetic file built from your redacted sample, with the agreed delimiter and quoting cases, imports with every row at the agreed column count and every field equal to the expected value. A row with the wrong number of fields is rejected and listed with its record and line number. A file with an unclosed quote is rejected as a whole, naming the line where the unreadable record starts, so nothing is stored shifted.

Check whether this job fits

Answer these without sending files, code or logins. Nothing is submitted unless you choose to contact us.

Do only some rows come out wrong, usually ones with punctuation in a description?
Does the importer report a different number of products from the number the supplier says are in the file?
In a plain text editor, are fields that contain the delimiter wrapped in double quotes?

Answer the questions to see whether this job fits.

Nothing is sent anywhere until you choose to email us.

Send an enquiry about this outcome

Checks you can run yourself

  1. Count the fields on each line

    Use a spreadsheet's text-import preview with the delimiter set explicitly, or ask your developer to print the number of fields on each line of a copy of the file.

    Look for: Lines whose count differs from the header are the ones to send. A description that continues after a line break shows up as a line with fewer fields than the header.

What you get

  • A change to the importer with its regression test
  • The synthetic file, the expected field values and the reject list produced from it
  • A one-page note on the delimiter, quoting and reject rules now in force
  • Steps to revert the change

Included

  • One existing importer path that reads one named supplier layout with a stated delimiter
  • Reproduce the fault on a synthetic file and record which rows shift and why
  • Read the file with a proper CSV reader using the agreed delimiter, quote character and doubled-quote escape, including line breaks inside quoted fields
  • Check the header and the field count of every row against the agreed layout, so a file in a different delimiter or layout is rejected instead of imported as one wide column
  • Read the whole file before storing any row, so a file with an unclosed quote is rejected as a whole, stores nothing and names the line where the unreadable record starts
  • Reject a row with the wrong number of fields, with its record number, line number, original text and reason, and prove that rows read equal rows accepted plus rows rejected
  • Add a regression test with the synthetic file

Not included

  • Adapting the importer to added, renamed or reordered columns, which are header changes: a file whose header differs from the agreed layout is rejected, not adapted
  • Repairing rows already stored in the wrong fields
  • Character encoding faults
  • Detecting the layout of arbitrary files from unknown suppliers
  • Changing how the supplier exports the file

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. A synthetic file whose description fields contain a comma, a doubled quote mark and a line break imports with every row at the agreed column count and every stored field equal to the expected value.

    Evidence: Expected and stored values side by side, and the column count per record.

  2. A file in the agreed non-comma delimiter imports to the same expected values, while a file in a different delimiter, whose header or field count does not match the agreed layout, is rejected with a clear message instead of importing as one wide column.

    Evidence: Both test runs' output and the rejection message.

  3. In a closed-quote test file, a row with too few fields and a row with too many fields are each rejected to a list giving its record number, line number, original text and reason, and rows read equal rows accepted plus rows rejected.

    Evidence: The reject list and the three counts for the test file.

  4. A test file with one unclosed quote followed by several well-formed rows is rejected as a whole: no row from it is stored, including the rows after the bad one, and the message names the line where the unreadable record starts.

    Evidence: The stored data before and after the attempt, which match, and the rejection message with its line number.

  5. The importer's existing tests still pass with the change applied.

    Evidence: Full test run output attached to the change.

Sign-off. You inspect the expected and stored values and the reject list, and sign off in writing. Payment follows sign-off; applying the change live stays with your maintainer.

If it fails. If the agreed synthetic file does not import as specified, you do not pay for this fixed scope. We hand over what we found and agree whether to stop or re-quote; no surprise work.

When it fits, and when we stop

It fits when

  • The importer is code you can share through an agreed route and it runs locally on synthetic data
  • You can supply a few invented or redacted rows that reproduce the shift, and you know or can ask which delimiter the supplier uses
  • The supplier file is delimited text with one stable delimiter

We stop and tell you if

  • Fields that contain the delimiter are not quoted at all in the supplier's file, so no reader can tell values apart reliably: we show the lines and stop
  • The only reproduction needs production data or credentials
  • The file is fixed-width, spreadsheet or another format

What could go wrong

Before you merge, closing the change leaves the importer as it was. After merge, your maintainer can revert the named commit. The change affects how files are read; reverting does not alter data stored earlier.

Scroll the table sideways to read it all.

RiskHow we handle it
A guessed delimiter makes the sample pass but breaks files from another supplier.The delimiter is agreed per named layout. The regression test keeps a file from another layout where your existing tests have one.
Malformed rows are quietly dropped and the product count looks plausible.Acceptance requires the reject list and the equation rows read = accepted + rejected.
An unclosed quote makes a reader swallow every following line into one field, so one fault hides all the good rows after it, or stores them shifted.The whole file is read before anything is stored, a file with an unclosed quote is rejected as a whole, and an acceptance test places good rows after the bad one and checks that none is stored.
Real product price lists are sent as the sample.The first enquiry asks for invented or redacted rows only.

An independent reviewer checks that quoted line breaks and doubled quotes are read as data, that bad rows are rejected rather than repaired by guesswork, that an unclosed quote rejects the whole file and stores nothing from it, and that the counts reconcile. Your maintainer reviews and applies the change.

How we deliver

We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.

  • Agree the importer path, delimiter, sample rows and expected values in writing
  • Reproduce the shift on a synthetic file and record which rows move and why
  • Switch the reading step to the agreed delimiter and quote rules without changing how fields are used afterwards
  • Add the layout check, the whole-file rule for an unclosed quote, the reject list and the rows-read check, then run the agreed file and the existing tests
  • Independent review of the change and the evidence
  • Hand over the change, the evidence and revert steps; your maintainer reviews and applies it

This is a one-off job, not emergency cover or a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after agreed inputs, secure access and any needed permissions are in place. Hosting, supplier and platform charges are excluded unless the written quote includes them. An enquiry creates no charge or booking.

Need to keep it working?

Discuss a standing check that runs this file test on every new supplier file.

Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.

Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.

What you can check

This is a new service. We have not delivered this job for a client yet.

Other ways to get this done

  • Open the file in a spreadsheet's text-import dialog and set the delimiter and quote character yourself to see whether the rows then line up. If they do, the importer is the one splitting wrongly.
  • Python's csv documentation explains reading files with newlines inside quoted fields and how line numbers differ from record numbers. docs.python.org

Questions

My supplier switched to semicolons. Is that this job?

Yes, if the importer should read the agreed semicolon layout. If the supplier may switch again without notice, a feed change report or a standing check is a better fit.

Will rows that are malformed be fixed automatically?

No. They are rejected and listed with their line number so someone can ask the supplier or decide. Guessing a fix can store wrong product data.

Does it cover a new column added by the supplier?

No. A file whose header does not match the agreed layout is rejected by the layout check, not adapted. Handling a changed header is a separate problem; see the guide on header changes and the feed version report.

Send an enquiry

Send us

  • Five to ten invented or redacted rows as plain text, including a row that goes wrong
  • Which delimiter the supplier states, and the row count it states, if any
  • The importer's name and language. No real price lists, customer data, credentials or code in the first enquiry

Later, once you agree

  • The importer code through an agreed company-controlled route, with local run instructions
  • Redacted sample files and expected field values, agreed in writing
  • A named person who can accept the change

You keep the importer, your database and the supplier relationship. We work on a copy of the importer with synthetic or redacted files and return a change your maintainer reviews and applies. We do not hold supplier credentials or touch production data.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “csv-delimiter-quoting-embedded-newlines-import” as the subject.