Remove duplicates from a CSV (exact, by key, and near matches)
Duplicate rows are the classic reason an import creates the same customer twice, overwrites a product, or doubles a mailing list. The hard part is not deleting rows: it is deciding which rows are actually duplicates.
A row that looks unique can be a duplicate that differs only by a capital letter, a trailing space, or an N/A in one column.
Check your file before you import it
Free preflight and fixes in your browser. Nothing is uploaded.
Exact duplicates are the easy case
Two rows with identical values in every column are always safe to collapse: keep the first occurrence and drop the rest. This is what almost every tool does — and where most tools stop.
Key duplicates are the ones that hurt
In real lists, the identity of a row is one or two columns: the email for contacts, the SKU for products. Two rows can agree on the key and differ everywhere else — one has the phone, the other the address.
That is why matching on a key column matters more than matching whole rows. And it is why the decision must stay visible: keep-first is a safe default, but you should see exactly which rows are being removed before it happens.
- Contacts: match on email (lowercased, trimmed).
- Products: match on SKU or handle.
- Transactions: match on order id or invoice number.
Why Excel's Remove Duplicates is not enough
Excel compares values exactly: "Mario@Example.com" and "mario@example.com" are different, "Black" and "Black " are different, "N/A" and an empty cell are different. Every normalization you would do by hand — trimming, lowercasing, unifying empty values — happens before the comparison in a proper cleaning tool.
BeforeImport normalizes those variants first, then finds duplicates on the full row or on the key column you choose, shows a preview of what will be removed, and lets you undo before exporting.
Frequently asked questions
Which row is kept when duplicates are removed?
The first occurrence is kept by default, and later copies are removed. Because every removal is previewed and reversible, you can verify the choice before exporting.
Does it find near-duplicates with typos?
BeforeImport normalizes case, spaces and empty-value variants before comparing, which removes the most common near-duplicates. Purely fuzzy name matching is deliberately conservative: it is better to leave a doubtful row than to delete a real customer.
Can I deduplicate only part of the file?
Yes — choose the key column for the scope you care about (for example email only), and the rest of the row is ignored for matching.
Is the file uploaded for deduplication?
No. Deduplication runs in your browser on your device; nothing is sent anywhere.
Related guides
Try it on your messiest file
Deterministic fixes, preview before every change, export CSV/XLSX and a cleaning report.
Open the free tool