Synthetic Industry

Free tool · updated 2026-10-11

CSV header diff: find missing, added, renamed and duplicate columns between two files

Paste two CSV header rows, or read the header of two small files in your browser, and see missing, added and likely renamed columns, duplicate and empty headers and the delimiter. Nothing is uploaded.

What it compares

Paste the header row of the file you expect (A) and the header row of the new or changed file (B). The tool reports the columns missing from B, the columns added in B, columns that look renamed, columns whose position moved, duplicate names within each header, empty header cells and the delimiter it found in each. You can also choose a CSV file for A or B: the browser reads the first part of the file, keeps only its header row and fills the box. The end of the header is found with the same quoting rules the comparison uses, so a stray double quote inside a name does not pull the rows below it into the box. If the header's opening quote is never closed in the part read, only the first line is kept and the page says so. The rest of the file is never placed on the page, and nothing is uploaded.

Do not paste confidential data. A header row rarely contains any, but a pasted row can include data by mistake, and the tool shows what you give it on your own screen.

  • Only header names are compared. Data rows, types and values are out of scope.
  • Leading blank lines are skipped and only the first record is read; a quoted header cell may span lines.
  • Files are decoded as UTF-8, or UTF-16 when a byte order mark says so; other encodings can show replacement characters.
  • Work limits keep the page responsive: up to 1,000 columns and 1,000,000 characters per side are accepted, and the similarity step has its own bounds, described under how columns are matched.

How columns are matched

Matching runs in three passes and uses each column once. First, identical text. Second, the same name after normalising: Unicode compatibility forms are applied, hidden characters (byte order mark, zero-width characters, soft hyphen) are removed, and case, spaces and punctuation are ignored, so 'Unit Price', 'unit_price' and 'UNIT-PRICE' are one name. Accents are kept: 'Cafe' and 'Café' are different names, though similar enough to be offered as a guess.

Third, for columns still unmatched, a similarity score. The score is the larger of two numbers: one minus the edit distance divided by the longer normalised name, and the overlap of the two word sets (words are split at spaces, punctuation and lower-to-upper camel-case boundaries). A pair is offered as a likely rename when the score is at least the threshold, which you can change from 0.40 to 1.00 (default 0.60). Pairs are assigned from the highest score down. A rename candidate is a guess to confirm, not a fact. To keep the page responsive, a name longer than 200 letters and digits (after normalising) is never scored for similarity, and scoring is skipped for all remaining columns if they would need more than 250,000 column pairs or more than 100 million character comparisons. When either bound applies the page says so in its warnings; identical and same-after-normalising matches are always checked.

  • 'Order Qty' to 'Order Quantity' scores 0.62 and is offered at the default threshold but not at 0.70.
  • 'Qty' to 'Quantity' scores 0.38 and is not offered: very short abbreviations are missed, so check the missing and added lists as well.
  • Repeated names are paired in the order they appear, so a second 'Price' with no partner is reported as missing or added.
  • A name longer than 200 letters and digits can still match by identical text or the same name after normalising, but no rename guess is offered for it.

Delimiter, quoting and the problems it flags

Each of comma, semicolon, tab and pipe is tried on the first record, honouring double quotes, and the one giving the most cells (at least two) wins; ties go to the earlier one. You can override either side. RFC 4180, the common description of CSV, defines only the comma, and says fields containing commas, double quotes or line breaks should be quoted with a quote inside a field doubled; semicolon, tab and pipe are common variants it does not define. If one file uses a semicolon and the other a comma, a parser set to one delimiter will read the other as a single column, so the tool warns about that.

It also flags duplicate names inside one header (exact or only after normalising), empty header cells including a trailing delimiter, cells with no letters or digits, leading or trailing spaces, hidden characters, a byte order mark on only one side, and stray or unterminated quotes. A positional import that maps by column number needs the position list too: the tool says whether the shared columns kept their relative order and which ones moved.

  • A trailing delimiter makes an empty last header cell.
  • Padded names such as ' Price' break exact-name matching in some importers even though they look the same on screen.
  • A byte order mark can be glued onto the first column name by a parser that does not strip it.

What it cannot tell you

It cannot tell you whether the data under a column means the same thing, whether units, currencies or date formats changed, or whether rows are the right shape: header names only. A renamed-looking pair may be two different fields and a genuinely renamed pair may not look similar. A header that was never a header (a file with no header row) will be compared as if it were one.

For why a changed supplier layout needs a declared header contract and row-shape checks, and for a checklist of import checks, see the supplier-feed header drift guide and the supplier-feed import checks collection linked on this page. If you receive a new version of a supplier feed before each import, the first related outcome is a change report comparing two named versions of one feed by product key, with price and stock moves over thresholds you set; it does not decide whether to import the new file. If an import breaks on quotes, line breaks, a byte order mark or accents, the two CSV outcomes listed with this tool each cover one named file. None of them is needed to use this tool.

  • No file leaves your browser, and no input is put in a URL or stored.
  • The result is a comparison of two strings of text, not a validation of your CSV.

Use the tool

Sources and limits

  • Header comparison implementation Checked 2026-10-11.
    • Columns are matched in three passes: identical text, then the same name after normalising case, spacing, punctuation, hidden characters and Unicode compatibility forms, then a similarity score against an editable threshold (default 0.60).
    • Only the first record of each input is read, delimiters tried are comma, semicolon, tab and pipe, and quoting follows RFC 4180 conventions.
    • The tool makes no network request and stores nothing; chosen files are read with FileReader in the browser and only the header row is placed in the page.
  • RFC 4180: Common Format and MIME Type for CSV Files Checked 2026-10-11.
    • There may be an optional header line as the first line, with the same number of fields as the rest of the file.
    • Fields containing line breaks, double quotes or commas should be enclosed in double quotes, and a double quote inside such a field is escaped by another double quote.
    • Each line should contain the same number of fields throughout the file, and the RFC defines the comma as the only separator; semicolon, tab and pipe are common variants the RFC does not define.
  • MDN: FileReader Checked 2026-10-11.
    • FileReader lets web applications read the contents of files stored on the user's computer, and can only read files the user has explicitly selected.
  • MDN: TextDecoder constructor Checked 2026-10-11.
    • By default a TextDecoder skips the byte order mark; the ignoreBOM option controls whether it is kept in the decoded text, which this tool uses so a byte order mark can be reported.