Headers become keys
The first row is treated as the header row, and each cell in it becomes a key in every output object. This is the convention almost every CSV follows, and when a file does not have a header row the conversion has to be told so, otherwise the first row of real data disappears into the key names.
Header cells are trimmed before becoming keys, so ` First Name ` becomes `First Name`; spaces inside the name remain. That is valid JSON and slightly awkward to work with in JavaScript, where it cannot be accessed with dot notation. Renaming headers in the source is usually cleaner than post-processing the JSON.
Duplicate headers are the case worth watching. The converter disambiguates later repeated trimmed names with a suffix, such as `id` and `id_2`; blank trimmed headers become positional names such as `column_1`. The suffix algorithm does not reserve names that already look suffixed: headers `id,id,id_2` can still create two `id_2` keys and the later column overwrites the earlier one in the JSON object. Check generated keys rather than assuming every source header survived exactly as written.
Type inference, and when to switch it off
With type inference on, lowercase `true`, `false` and `null` become JSON scalars; uppercase `TRUE` stays text. Empty cells remain empty strings by default; use the separate empty-as-null option when null is the intended meaning. Numeric inference accepts decimal and exponent forms without leading zeros or a leading plus, but preserves unsafe integer results as strings. Decimals still use JavaScript floating-point precision.
Identifiers still deserve deliberate review. A postcode of `01234`, a phone number, an account number, and a version string are text even when some of them look numeric. Inference cannot know your business meaning, so turn it off when a column must remain text regardless of its spelling.
The safe approach for data you do not control is to convert with inference off, so every value is a string, and then coerce the specific fields you know are numeric. That is the opposite of convenient, and it is the only way to be sure an identifier survived. For a file you wrote and understand, inference on is fine.
Quoting, commas and line breaks inside fields
A CSV field may contain the delimiter, a double quote, or a line break, as long as the field is quoted, and a double quote inside a quoted field is escaped by doubling it. A correct parser handles all of this, which is precisely why using one matters more here than almost anywhere else.
Splitting a CSV line on commas is the classic mistake, and it fails on the first address field containing `Berlin, Germany`. The symptom is a row with more columns than the header, and every value after the offending field lands under the wrong key. On a large file it may corrupt only a handful of rows, which is worse than failing outright because it looks like it worked.
Multi-line fields, a quoted field containing a newline, break line-by-line processing in the same way. A file where the row count does not match what you expect almost always has one.
Ragged rows and missing values
A row with fewer cells than the header is missing trailing values; a row with more has an extra field, usually from a quoting problem upstream. Neither is valid CSV, and both occur constantly in real files.
This converter keeps every object key present: missing trailing cells become empty strings by default, or null when the explicit empty-as-null option is selected. That distinction matters if a consumer treats an empty string differently from a missing value. An over-long row is reported as an extra-field warning rather than silently treated as a clean record.
An over-long row is a signal rather than something to paper over. It means the file is malformed, and the most likely cause is an unquoted delimiter inside a value, so the row you are looking at is probably not the only damaged one.
Dotted headers stay flat
A CSV column named `address.city` becomes the literal JSON key `address.city`; this converter does not interpret dots as a path separator. Likewise, `tags.0` remains a string key rather than becoming an array element.
That matters when a JSON-to-CSV export used dotted names to represent nested objects. The CSV output is readable and useful as a table, but converting it back does not automatically unflatten those columns or rebuild arrays. If the original tree matters, keep the JSON source or use a deliberately designed schema transformation.
Treat dotted headers as ordinary column names unless the next tool explicitly documents an unflattening convention. Assuming a round trip reconstructs the original tree can create a result that looks plausible while having a different shape.
Encoding problems
A CSV file carries no declaration of its own encoding, so a parser has to assume, and UTF-8 is the right assumption. When a file is actually Windows-1252 or Latin-1, which is what Excel produces on a Western Windows install unless told otherwise, accented characters and symbols arrive mangled.
The recognisable symptom is a single character becoming two: `é` appearing as `é`, or a curly apostrophe as `’`. That specific pattern is UTF-8 bytes being read as single-byte characters, and it means the file needs converting to UTF-8 before parsing rather than fixing afterwards.
A byte-order mark at the start of the file is the related nuisance: it makes the first header key begin with an invisible character, so a lookup for `id` fails against a key that looks identical. A parser should strip it, and a key that mysteriously does not match is the sign that something did not.