CSV Deduplicator

Remove duplicate rows from a CSV file.

Input
CSV input
Output
Result
Options

Leave blank to compare entire rows.

Detected automatically unless you pick one.

About this tool


Duplicate rows arrive from repeated exports, merged files and re-run imports. This removes them, either matching on the entire row or on a single key column when only that needs to be unique.

Matching on a key column is the more useful mode in practice: records that share an ID but differ in a timestamp are logically duplicates even though their rows are not identical.

How to use it

  1. Paste or upload your csvDrop a file onto the input pane, use the file picker, or paste the text directly.
  2. Adjust the options if neededThe defaults suit most input; open Options to change the behaviour.
  3. Remove duplicatesPress Remove duplicates, or use Ctrl+Enter (Cmd+Enter on macOS).
  4. Copy or downloadCopy the result, or download it as a .csv file.

Worked examples


Each example below is executed against this tool by the test suite, so what you see is what the tool actually produces.

Exact duplicate removed

Input

a,b
1,2
1,2
1,3

Output

a,b
1,2
1,3

The repeated row is removed; the row differing in column b is kept.

What to watch for


The details that decide whether a conversion is correct, and where information can be lost without any error being raised.

Whole-row versus key-column matching
By default every column must match for a row to count as a duplicate. Naming a column instead treats any rows sharing that value as duplicates, which is what you want for deduplicating by ID or email where other fields may vary.
Keeping the first or the last occurrence
When rows differ outside the key column, the choice matters. Keeping the last occurrence is right for chronological data where later rows are updates; keeping the first preserves the original record. Row order is otherwise unchanged.
Case sensitivity
Matching is case-sensitive by default, so Alice and alice are distinct. Turn it off for deduplicating email addresses or usernames, where case is usually not meaningful.
Whitespace is significant
A value with a trailing space does not match one without. If your data may have inconsistent padding, run it through the CSV Formatter first or the near-duplicates will survive.

Limitations


  • Matches exact values; no fuzzy matching.
  • Whitespace differences prevent a match.
  • Processing happens in your browser, so very large inputs are bounded by available memory. Files above roughly 10 MB are handled but will feel slower, and multi-hundred-megabyte files are better suited to a command-line tool.

Questions


How do I deduplicate by ID only?
Name that column in Options. Any rows sharing its value are treated as duplicates regardless of their other fields.
Should I keep the first or last occurrence?
Last for chronological data where later rows supersede earlier ones; first to preserve the original record.
Why did near-identical rows survive?
Probably differing whitespace or case. Turn off case sensitivity, or normalise the file with the CSV Formatter first.