Synthetic Industry

Job csv-encoding-bom-garbled-characters-import · revised 11 October 2026

Make a supplier CSV import keep accented and special characters intact

A named supplier CSV with accents, symbols and a leading byte-order mark imports with every agreed test value unchanged and its first column recognised; before and after results attached.

You might be seeing

  • Accented letters, currency symbols or trademark signs turn into strange character pairs after import
  • The importer reports the first column as missing, or skips the first product, although the header looks correct in a spreadsheet

No passwords, keys, card details or admin invites needed to start.

What usually happened

A CSV file is decoded with a different character encoding from the one used to save it, or a leading byte-order mark stays in the first header name as an invisible character. Values then change on the way in and the damage is stored in your product data, or the first column can no longer be matched by name.

Who it’s for: Operations or e-commerce manager at a distributor or online shop whose product import shows garbled names, or cannot find its first column.

Usually starts when: Product names show character pairs such as é or the replacement sign � after an import, or the importer says the SKU column is missing although the header looks right.

The result: A named test file with agreed non-ASCII values and a leading byte-order mark imports through your existing importer with each value identical to the source and the first column recognised. You receive the repair as a reviewable change with before and after results.

Check whether this job fits

Answer these without sending files, code or logins. Nothing is submitted unless you choose to contact us.

What does the import show after the file is read?
When you open a copy of the original supplier file in a plain text editor, do the characters look right there?
Can you make a few redacted or invented rows that still show the fault?

Answer the questions to see whether this job fits.

Nothing is sent anywhere until you choose to email us.

Send an enquiry about this outcome

Checks you can run yourself

  1. Look at the first header name

    Open a copy of the file in a plain text editor that can show invisible characters, or ask your developer to print the first header name with its character codes.

    Look for: An invisible character before the first header name suggests a byte-order mark. Character pairs such as é in place of é suggest the file was read in a different encoding from the one used to save it.

What you get

  • A change to the importer with its regression test
  • A before and after table of expected and stored values for the agreed test rows, shown as code points
  • A short note on which encodings the importer now accepts and what it does with undecodable bytes
  • Steps to revert the change

Included

  • One existing importer path that reads one named supplier CSV layout
  • Reproduce the fault on a synthetic file built from your redacted sample, recording the exact bytes of the header and of up to ten affected values
  • Make the importer decode the file with an explicit, agreed encoding named in the code rather than a language or system default, handle a leading byte-order mark, and reject bytes that are not valid in that encoding instead of storing substitutes
  • Add a regression test that uses the synthetic file

Not included

  • Repairing text already stored garbled in your database; that is a separate reconciliation job
  • Recovering the original characters from a file that an earlier spreadsheet save already damaged
  • Guessing the encoding of arbitrary files from unknown suppliers
  • Changing database character sets or collations
  • Faults caused by delimiters, quotes or line breaks inside fields, which are a different job

How we know it’s done

Agreed with you before work starts. Each check produces evidence you keep.

  1. A synthetic file that starts with a byte-order mark and contains the agreed non-ASCII values imports with every stored value equal to the expected value and the first header name equal to the agreed column name.

    Evidence: Before and after table of expected and stored values as code points, and the header name printed with its character codes.

  2. The same file saved without a byte-order mark imports to identical values.

    Evidence: The second run's table and the regression test output.

  3. A file containing bytes that are not valid in the agreed encoding is rejected with a message that names the line, and no rows from it are stored.

    Evidence: Rejection message and an unchanged row count in the test store.

  4. The importer's existing tests still pass with the change applied.

    Evidence: Full test run output attached to the change.

Sign-off. You inspect the before and after table and the test results, and sign off in writing. Payment follows sign-off; applying the change to your live importer stays with your maintainer.

If it fails. If the agreed test file does not import as specified, you do not pay for this fixed scope. We hand over what we found and agree whether to stop or re-quote; no surprise work.

When it fits, and when we stop

It fits when

  • The importer is code you can share through an agreed route and it runs locally on synthetic data
  • You can supply a redacted or invented sample that reproduces the fault, and you know which program saves the supplier file or can ask the supplier
  • The affected values are product data, not personal data

We stop and tell you if

  • The only reproduction needs production data, a live supplier download or credentials
  • The original characters are already lost in the file you hold, so no decoding can restore them: we explain what to ask the supplier for and stop
  • The importer is a closed product whose code you cannot change: we list the settings to try instead

What could go wrong

Before you merge, closing the change leaves your importer as it was. After merge, your maintainer can revert the named commit. The change affects how files are read; it does not alter data already stored, so reverting restores the earlier reading behaviour only.

Scroll the table sideways to read it all.

RiskHow we handle it
A guessed encoding makes the test pass but changes values for another supplier's files.The encoding rule is agreed per named layout, not guessed globally, and the regression test includes a second layout from your existing tests where one exists.
Invalid bytes are replaced with a substitute character and stored without anyone noticing.Acceptance requires that undecodable input is rejected with a message naming the line, and that no rows from that file are stored.
A real price list or customer detail is sent in the sample.The first enquiry asks only for invented or redacted rows. We stop and ask for a replacement if real data arrives.

An independent reviewer checks the byte-level evidence, that undecodable input is rejected rather than silently replaced, and that the change touches only the agreed importer path. Your maintainer reviews and applies the change under your own rules.

How we deliver

We arrange the work and independent review, then show you the result against the agreed checks. You keep authority over your systems.

  • Agree the importer path, the sample rows, the expected values and the test route in writing
  • Reproduce the fault on a synthetic file and record the header and affected values at byte level before changing anything
  • Make the smallest change that decodes the agreed layout correctly, including a leading byte-order mark, and rejects undecodable bytes
  • Run the agreed file, a file without a byte-order mark and an invalid-byte file, and compare with the expected values
  • Independent review of the change and the evidence, with a separate look at anything that writes to storage
  • Hand over the change, the before and after table and revert steps; your maintainer reviews and applies it

This is a one-off job, not emergency cover or a subscription. We confirm eligibility, the total price, a start window and a delivery date before you accept. Work starts only after agreed inputs, secure access and any needed permissions are in place. Hosting, supplier and platform charges are excluded unless the written quote includes them. An enquiry creates no charge or booking.

Need to keep it working?

Discuss a standing check that tests every new file from this supplier for encoding changes.

Ongoing work is separately scoped and quoted: no monitoring, response-time guarantee or automatic subscription is included in this job.

Explore an ongoing engineering lane, or mention the responsibility you need in your enquiry.

What you can check

This is a new service. We have not delivered this job for a client yet.

Other ways to get this done

  • Look for an encoding or file character set setting in your importer and try it on a copy of the file. Python's documentation describes a codec that skips a leading byte-order mark, which is the usual fix when the first header name will not match. docs.python.org
  • Ask the supplier to re-export the feed as UTF-8 and to say whether the file starts with a byte-order mark.

Questions

Can you repair names that are already garbled in my database?

Not in this job. If the original file is still available, reconciliation is a separate agreed scope. If only the garbled copy remains, some characters cannot be recovered.

Which encoding will the importer use?

The one agreed for your named supplier layout, checked against your sample. We do not make the importer guess for unknown files.

Why not rely on the default encoding of the importer's language?

Defaults differ by language, version and platform. Python's csv documentation, for example, says a file opened for reading uses the system default encoding in Python 3.14 and UTF-8 in Python 3.15. So the importer names the agreed encoding explicitly rather than relying on a default.

Does it need my live supplier feed?

No. We work from a few invented or redacted rows. A live download is not needed and should not be sent.

Send an enquiry

Send us

  • Three to ten invented or redacted rows that show the fault, as plain text, including the header line
  • The importer's name and language, and which program the supplier says exports the file
  • What you see: the wrong characters or the missing-column message. No real price lists, customer data, credentials or code in the first enquiry

Later, once you agree

  • The importer code through an agreed company-controlled route, with instructions to run it locally
  • Redacted sample files and the expected values, agreed in writing
  • A named person who can accept the change

You keep the importer, your database and the supplier relationship. We work on a copy of the importer code with synthetic or redacted sample files and return a change for your developer or maintainer to review and apply. We do not hold supplier credentials or touch production data.

A public HTTPS link only, without login details, query strings or fragments. No code or logs.

Sending emails your enquiry and contact address to our team through our mail provider (Resend). It is not kept in a website database. Do not send passwords, keys, recovery links, confidential code or customer records. Your contact email is unverified; nothing is ordered, charged or reserved. Privacy notice.

Email fallback: open your mail app

If website submission is unavailable, review and send the fallback email yourself. An email fallback is not a website receipt. Or write to hello@syntheticindustry.ai with “csv-encoding-bom-garbled-characters-import” as the subject.