Guide

How to convert CSV to JSON

CSV is a table and JSON is a tree, so this direction is the easy one: every row becomes an object and every column becomes a key. The shape of the result is predictable before you start.

The difficulty is not structure but interpretation. CSV has no types, every value is text, so the conversion has to decide whether `42` is a number or a string, whether an empty cell is an empty string or null, and whether `TRUE` is a boolean. Those decisions are where a conversion goes wrong, and they are almost always recoverable if you know what to look for.

Open the CSV to JSON

Use this when


  • A spreadsheet export needs to go into an API that expects JSON
  • You have a CSV from a database dump and want to query it with jq
  • A JavaScript application needs to read data that arrived as CSV
  • You want to validate tabular data against a JSON Schema
  • A CSV needs converting to a JSON fixture for tests

Headers become keys


The first row is treated as the header row, and each cell in it becomes a key in every output object. This is the convention almost every CSV follows, and when a file does not have a header row the conversion has to be told so, otherwise the first row of real data disappears into the key names.

Header cells are trimmed before becoming keys, so ` First Name ` becomes `First Name`; spaces inside the name remain. That is valid JSON and slightly awkward to work with in JavaScript, where it cannot be accessed with dot notation. Renaming headers in the source is usually cleaner than post-processing the JSON.

Duplicate headers are the case worth watching. The converter disambiguates later repeated trimmed names with a suffix, such as `id` and `id_2`; blank trimmed headers become positional names such as `column_1`. The suffix algorithm does not reserve names that already look suffixed: headers `id,id,id_2` can still create two `id_2` keys and the later column overwrites the earlier one in the JSON object. Check generated keys rather than assuming every source header survived exactly as written.

Type inference, and when to switch it off


With type inference on, lowercase `true`, `false` and `null` become JSON scalars; uppercase `TRUE` stays text. Empty cells remain empty strings by default; use the separate empty-as-null option when null is the intended meaning. Numeric inference accepts decimal and exponent forms without leading zeros or a leading plus, but preserves unsafe integer results as strings. Decimals still use JavaScript floating-point precision.

Identifiers still deserve deliberate review. A postcode of `01234`, a phone number, an account number, and a version string are text even when some of them look numeric. Inference cannot know your business meaning, so turn it off when a column must remain text regardless of its spelling.

The safe approach for data you do not control is to convert with inference off, so every value is a string, and then coerce the specific fields you know are numeric. That is the opposite of convenient, and it is the only way to be sure an identifier survived. For a file you wrote and understand, inference on is fine.

Quoting, commas and line breaks inside fields


A CSV field may contain the delimiter, a double quote, or a line break, as long as the field is quoted, and a double quote inside a quoted field is escaped by doubling it. A correct parser handles all of this, which is precisely why using one matters more here than almost anywhere else.

Splitting a CSV line on commas is the classic mistake, and it fails on the first address field containing `Berlin, Germany`. The symptom is a row with more columns than the header, and every value after the offending field lands under the wrong key. On a large file it may corrupt only a handful of rows, which is worse than failing outright because it looks like it worked.

Multi-line fields, a quoted field containing a newline, break line-by-line processing in the same way. A file where the row count does not match what you expect almost always has one.

Ragged rows and missing values


A row with fewer cells than the header is missing trailing values; a row with more has an extra field, usually from a quoting problem upstream. Neither is valid CSV, and both occur constantly in real files.

This converter keeps every object key present: missing trailing cells become empty strings by default, or null when the explicit empty-as-null option is selected. That distinction matters if a consumer treats an empty string differently from a missing value. An over-long row is reported as an extra-field warning rather than silently treated as a clean record.

An over-long row is a signal rather than something to paper over. It means the file is malformed, and the most likely cause is an unquoted delimiter inside a value, so the row you are looking at is probably not the only damaged one.

Dotted headers stay flat


A CSV column named `address.city` becomes the literal JSON key `address.city`; this converter does not interpret dots as a path separator. Likewise, `tags.0` remains a string key rather than becoming an array element.

That matters when a JSON-to-CSV export used dotted names to represent nested objects. The CSV output is readable and useful as a table, but converting it back does not automatically unflatten those columns or rebuild arrays. If the original tree matters, keep the JSON source or use a deliberately designed schema transformation.

Treat dotted headers as ordinary column names unless the next tool explicitly documents an unflattening convention. Assuming a round trip reconstructs the original tree can create a result that looks plausible while having a different shape.

Encoding problems


A CSV file carries no declaration of its own encoding, so a parser has to assume, and UTF-8 is the right assumption. When a file is actually Windows-1252 or Latin-1, which is what Excel produces on a Western Windows install unless told otherwise, accented characters and symbols arrive mangled.

The recognisable symptom is a single character becoming two: `é` appearing as `é`, or a curly apostrophe as `’`. That specific pattern is UTF-8 bytes being read as single-byte characters, and it means the file needs converting to UTF-8 before parsing rather than fixing afterwards.

A byte-order mark at the start of the file is the related nuisance: it makes the first header key begin with an invisible character, so a lookup for `id` fails against a key that looks identical. A parser should strip it, and a key that mysteriously does not match is the sign that something did not.

Common problems


Each of these is something that actually happens, with the cause rather than a generic suggestion to check your input.

Leading zeros disappeared from IDs or postcodes

Cause
A spreadsheet or an earlier processing step may have converted the identifier to a number. This converter preserves a leading-zero value such as `01234` as text.
Fix
Check the original CSV first. Disable inference to preserve all identifier columns as strings; digits already removed upstream cannot be recovered here.

Values are under the wrong keys in some rows

Cause
A field contained the delimiter and was not quoted, so the row split into more columns than the header has.
Fix
Fix the file at the source with proper quoting. The rows you have spotted are unlikely to be the only affected ones.

The first row of data became the keys

Cause
The file has no header row, but the conversion assumed one.
Fix
Tell the conversion there is no header, or add one. Generic keys are better than losing a row of data.

Accented characters are mangled

Cause
The file is Windows-1252 or Latin-1 rather than UTF-8, a typical Excel export.
Fix
Re-save it as UTF-8, or in Excel use "CSV UTF-8" rather than plain CSV when exporting.

A key exists but a lookup for it fails

Cause
A byte-order mark is attached to the first header name, invisibly.
Fix
Strip the BOM from the start of the file. It is three bytes before the first character.

Fewer objects than the file has lines

Cause
A quoted field contains a line break, so one record spans several lines.
Fix
Nothing to fix. This is correct CSV and the record count is right. It only looks wrong if you counted lines.

Questions


Should the first row be treated as headers?
Almost always yes, nearly every CSV in circulation has a header row, and the keys have to come from somewhere. If your file does not have one, say so before converting, or the first record of real data will be consumed as key names.
Why did my phone number become a number?
CSV has no types, so the conversion infers them, and a cell of digits looks numeric. Any identifier (phone number, postcode, account number, anything with a leading zero) should be converted with inference off so it stays a string.
Can I get nested JSON out of a flat CSV?
Not automatically with this converter. `address.city` and `tags.0` remain literal flat keys, and JSON text inside a cell remains a string. Use a schema-aware transformation when you need nested objects or arrays, and keep the original JSON if you need an exact round trip.
What is the largest CSV I can convert?
Around 12 MB, which is the point where in-browser processing stops feeling immediate. For much larger files a streaming command-line tool is the better instrument. Nothing is uploaded here, so the work happens on your own machine either way.