- Quoted fields, doubled quotes and embedded newlines
- A field wrapped in double quotes may contain the delimiter, line breaks, or quotes escaped by doubling them. So "x, y" is one value containing a comma, "has ""quotes""" becomes has "quotes", and a quoted field spanning two lines stays one value with a newline in it. This is RFC 4180 behaviour and the reason splitting on commas by hand fails on real data.
- Duplicate headers are made unique, not dropped
- CSV allows the same column name twice; a JSON object cannot have duplicate keys. Rather than silently losing a column, the second occurrence is suffixed: name and name become name and name_2. The warning lists every rename so you know which column is which. Blank header cells become column_1, column_2 and so on by position.
- Type inference is deliberately conservative
- Values that are unambiguously numbers become numbers, true and false become booleans, and the literal null becomes null. Anything that could be an identifier stays a string: 007 keeps its leading zero, +44 7700 stays text, and an integer too large for JavaScript to represent exactly is left as a string rather than silently rounded. This is the behaviour you want for IDs, phone numbers and postal codes.
- Ragged rows are reported, never silently truncated
- If a row has more fields than the header, the extras cannot be named and are dropped, with a warning saying so. If it has fewer, the missing keys are filled with empty strings, or null if you prefer. Either case usually means a quoting problem upstream, so the warning is worth reading.
- Empty cells: string or null?
- By default an empty cell becomes an empty string, which matches what the file literally contains. Switch on "Treat empty cells as null" when a blank means "no value" rather than "the empty string", useful when the JSON feeds a database import.
- Delimiters and line endings
- The delimiter is detected automatically, so comma, semicolon, tab and pipe files all work; European exports using semicolons are common. Both LF and CRLF line endings are handled, and a UTF-8 byte order mark at the start of the file is stripped so the first column name is not corrupted.