A broken file cannot be read around
The XML specification is blunt. A document is well-formed only if it follows the grammar and the well-formedness constraints, and breaking one is a fatal error. The processor must report it and must not carry on passing the document's data to the application as normal. That is why a supplier feed with one bad item usually fails as a whole, rather than importing the other 9,999 items. An importer that appears to skip the bad item is not following the specification, and the item it skips may be the one that mattered.
This changes what you ask for. When the file is not well-formed, the useful result is a precise message: the line and column of the first error, and what rule it breaks, so the supplier can correct the file. The right behaviour for your importer is to reject the whole file before it changes any product, and keep the previous data.
The usual first errors in a product feed
Four causes are worth checking first. An ampersand in a product name such as Nuts & Bolts that was not escaped as & is the classic case; the specification allows literal & and < only as markup or inside CDATA sections, comments and processing instructions. A control character, for example one pasted from another system, falls outside the legal character range. A closing tag whose letter case differs from the opening tag does not match, because no case folding is applied. And text or a second root element after the final closing tag breaks the single-root rule.
Encoding is a fifth, and it fails in two different ways. The specification makes it a fatal error if the bytes contain sequences that are not legal in the encoding the parser has determined, so Latin-1 text with an accented letter, declared as UTF-8, stops at that byte. A wrong declaration is harder to see. If the bytes happen to be legal in the declared encoding, for example UTF-8 text declared as ISO-8859-1, the parser reports no error and the accented letters come out as character pairs. In a computed run with Python's built-in parser, the word Café in UTF-8 bytes under an ISO-8859-1 declaration parsed cleanly and returned Café. That is a garbled-text fault rather than a parse error, and the guide on encodings and the byte-order mark explains how to spot it.
- Look at the exact line and column in the message before opening the whole file.
- A message that mentions an invalid token usually points at an unescaped character within a few characters of the position.
When the file is well-formed but the data is wrong
The opposite failure is quieter. The file parses, but fields are empty or contain stray markup. That points to the importer: matching tags as plain text rather than with a parser means a prefixed tag, a namespace declared on the root, or a CDATA section is not read the way the supplier meant. Escaped characters such as & need decoding into the real character, and text inside CDATA is data even if it looks like markup. These are faults in how the importer reads the file, and the supplier cannot fix them.
A safe first investigation
Open a copy of the file in a web browser or run it through any XML parser. A well-formed file displays as a tree or parses without error. A fault shows the line and column. Open that line in a text editor and look for the four causes above. Keep your own importer out of this step: you want the parser's verdict first. If the file is well-formed and your import still loses fields, compare one item field by field with what the importer stored.
If the feed comes from anywhere other than a named supplier, treat it as untrusted. Python's XML documentation warns that attacker-controlled XML can exhaust memory or CPU through entity expansion and compressed bombs, and recommends checking which Expat version is in use. A parser should refuse documents that declare entities or external resources unless the layout needs them.
What fixes it, and what does not fit
Fix the supplier's file when it is not well-formed, and send them the message verbatim. Fix your importer when a well-formed file loses fields: use a real parser, configure its namespaces, and map the agreed fields. A regression test uses a small invented sample with a namespaced price, a CDATA description and an escaped ampersand, and a broken copy that must be rejected as a whole.
Not a fit: repairing the supplier's XML with text substitution, which can store wrong products, designing the supplier's schema, tuning very large files, or the outgoing file you send to a shopping platform.
How the paid job is accepted
The fixed job xml-supplier-feed-parse-errors-import is £295 for one importer path and one named supplier layout. The well-formed sample must import every item with the agreed namespaced, CDATA and escaped values; a copy with one unescaped ampersand must be rejected as a whole with its line and column and leave stored products unchanged; and a file declaring an entity must be refused or left unexpanded. Prices are untested proposals, and payment follows the agreed checks and your sign-off. Nothing is booked or charged by an enquiry.
Sources and limits
- W3C Extensible Markup Language (XML) 1.0, Fifth Edition Checked 2026-10-11.
- A well-formedness violation is a fatal error: the processor must report it and must not continue passing character data and structure information to the application as normal.
- Literal < and & may appear only as markup, or inside comments, processing instructions and CDATA sections; end-tag names must match start-tag names exactly, with no case folding.
- The legal Char range excludes most control characters, such as U+0000 to U+0008, and a document has exactly one root element.
- It is a fatal error if an entity is determined to be in a certain encoding but contains byte sequences that are not legal in that encoding; UTF-8 is the encoding assumed when an entity carries no declaration and no byte-order mark.
- Python xml documentation, XML security section Checked 2026-10-11.
- Attacker-controlled XML can be used for denial of service, local file access and network connections, and the page names billion laughs, quadratic blowup and decompression bombs.
- Python's built-in parsers use libexpat, and by default Expat does not access local files or make network connections.