Synthetic Industry

Troubleshooting guide · updated 2026-10-11

Supplier CSV shows é instead of é, or the first column goes missing: check the encoding and byte-order mark

Why accented letters turn into character pairs and why an invisible first character breaks the first header name, how to tell which one you have, and what a safe fix looks like.

What is happening to the bytes

A CSV file is only bytes. Whatever reads it must decide which encoding turns those bytes into letters. In UTF-8 the letter é is stored as two bytes, C3 and A9. A reader that assumes a single-byte Windows code page sees those same two bytes as two separate characters and shows é. The euro sign, three bytes in UTF-8, comes out as three characters. That is the whole fault: the file may be perfectly good, and the reader has the wrong idea of how it was written.

There is no rule that settles this for every supplier. RFC 4180, which only records common CSV practice, says usage is commonly US-ASCII and lets a sender name another character set. The W3C tabular metadata vocabulary defaults to UTF-8. The default in the importer's own language can differ, and it can change: the Python 3.14 csv documentation says a file opened for reading is decoded with the system default encoding (see locale.getencoding()), which depends on the machine's settings, while the Python 3.15 documentation says UTF-8. The same importer can therefore read the same file differently on two machines or after an upgrade. So each supplier layout needs an encoding written down and named in the importer's code, not guessed and not left to a default.

  • Character pairs beginning with à or â usually mean UTF-8 bytes were read as a single-byte code page.
  • A diamond with a question mark, or a plain question mark, usually means substitution already happened and the original characters may be gone.

The invisible character at the start of the file

Some exports begin with a byte-order mark: the three bytes EF BB BF. In UTF-8 it is only a signature, and Python's documentation says its use is discouraged. The practical effect is in the first header name. The plain utf-8 codec does not strip it, so the first header becomes an invisible character followed by sku, and a lookup for sku fails or skips the first column. The utf-8-sig codec skips the three bytes when they are present and decodes the same file when they are not, which is why it is the usual reader for files from a supplier you do not control.

The symptom is distinctive: every value looks right, but the first column is reported missing, or the first product is lost, and the header looks correct in any spreadsheet because the character is invisible.

Other diagnoses to rule out

Look at the original file before changing any code. If the characters are already wrong in the file you received, no reader can restore them; the supplier or an earlier spreadsheet save damaged them, and the correct fix is a fresh export. If the file holds replacement characters, check whether your importer decodes with errors set to replace. Python's documentation says that handler substitutes U+FFFD, and ignore skips bad bytes without notice, so either one stores damage silently.

Two other causes look similar. A database column or a web page that is not set to a matching character set shows the same pairs even when the import was correct. And rows that shift into the wrong fields are a delimiter or quote problem, not an encoding one.

  • Compare the same product in the original file, the import result and the store page; the step where it first goes wrong is the cause.
  • A fault in only one supplier's files points to that supplier's export settings.

A safe first investigation

Work on a copy of three to ten rows with invented prices, never the only copy of a real price list, and do not open and re-save it in a spreadsheet, which can change the encoding. Look at the first three bytes of the file. The sequence EF BB BF means a byte-order mark. Decode a line containing é, the euro sign and a dash first as UTF-8, then as the Windows code page, and see which reading gives the supplier's real text. The table below, built from invented characters, shows what a wrong reading looks like.

  • Find the line in the importer that opens the file and check whether it names an encoding. If it names none, the language's default decides, and that default depends on the version and the machine.
  • Ask the supplier which program writes the file and with which encoding.
  • Write the answer next to the layout name so the next person does not guess.

If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.

character | UTF-8 bytes | wrongly read as Windows-1252
é         | C3 A9       | é
ü         | C3 BC       | ü
€         | E2 82 AC    | €
™         | E2 84 A2    | â„¢
(synthetic table; computed, not from a supplier file)

What fixes it, and what does not fit

The fix is a decoding rule agreed per supplier layout and passed to the file-opening call explicitly rather than taken from a default: read as UTF-8 with the byte-order-mark handling, or another named encoding where the supplier says so, and reject bytes that are not valid instead of replacing them. A regression test uses a few invented rows with accents, a currency symbol and a leading byte-order mark.

Not a fit: making the importer guess encodings for unknown files, repairing text already stored garbled in your database, recovering characters that an earlier save destroyed, or changing database collations. Those need their own scope and sometimes the supplier's help.

How the paid job is accepted

The fixed job csv-encoding-bom-garbled-characters-import is £195 for one importer path and one named supplier layout. A synthetic file with a leading byte-order mark and the agreed non-ASCII values must import with every stored value equal to the expected one and the first header name equal to the agreed name. The same file without a mark must give identical values, and invalid bytes must be rejected with a message naming the line. Prices are untested proposals, and payment follows the agreed checks and your sign-off. Nothing is booked or charged by an enquiry. Send invented rows and the importer's name first, never real price lists, credentials or code.

Sources and limits

  • Python codecs documentation Checked 2026-10-11.
    • A UTF-8 variant called utf-8-sig skips a leading byte-order mark when decoding; the plain utf-8 codec keeps U+FEFF as an ordinary character.
    • The errors argument selects what happens to bad bytes: strict raises an error, ignore skips them, replace substitutes U+FFFD when decoding.
  • Python 3.14 csv documentation Checked 2026-10-11.
    • Because open() is used to read a CSV file, the file is by default decoded using the system default encoding (see locale.getencoding()); another encoding is chosen by passing the encoding argument of open().
  • Python 3.15 csv documentation Checked 2026-10-11.
    • Because open() is used to read a CSV file, the file is by default decoded using UTF-8, so the documented default differs from Python 3.14.
  • W3C Metadata Vocabulary for Tabular Data Checked 2026-10-11.
    • Its dialect description lists utf-8 as the default encoding.
  • RFC 4180: common format for CSV files Checked 2026-10-11.
    • It is informational, says common CSV usage is US-ASCII and allows other charsets through an optional parameter; it does not set an Internet standard.