Images, documents & data
How to find and remove duplicate rows from a CSV
Clean exact duplicates without deleting legitimate customers, products, or orders.
Quick answer
First fix parsing, then decide what counts as a duplicate. orders-source.csv contains four fictional order records in six columns, including a quoted comma, a line break, an accented name and SKU 000123. A physical line is not necessarily a record: the quoted gift note spans two lines.
Step by step
Download the source, select it in CSV tools and choose Open. Check six headers and four records before removing exact duplicates. With the filter empty, remove duplicates and export CSV. Open that export again: three records should remain. Test orders-eu.csv separately; its semicolon separates columns while 12,50 remains one text value, not two columns and not an automatically converted number.
Common mistakes
If Café is garbled, re-export from the originating app as UTF-8; this tool reads local text as UTF-8 and is not a legacy-encoding repair service. Never replace all commas with periods: commas also separate fields or belong to quoted text. Missing leading zeros cannot be reconstructed reliably from a numeric value. Import SKU columns as text in the receiving spreadsheet.
Completion check
Confirm that order 1001 appears once while order 1003 remains: it has the same SKU but a different quantity and is not a duplicate. Verify 000123, Café cup, the multiline note and the quoted word after reopening. Export uses the currently filtered rows, so clear a filter before downloading the full dataset. Keep the source until the destination system accepts the new file.
Make the right call
Use the order, not the product, to decide what is duplicate
In the supplied file, one order was exported twice and another order legitimately contains the same SKU. The safe operation is exact-row deduplication after verifying parsing—not deleting all repeated product codes.
| Check | Example or action | How to judge the result |
|---|---|---|
| Order 1001 is present twice with identical fields | Remove one exact duplicate after confirming all six columns. | The cleaned file contains three data records; count parsed records, not physical text lines. |
| Order 1003 uses the same SKU with a different quantity | Keep it as a separate order. | A shared SKU identifies a product, not a duplicated sale. Customer names and shared emails are similarly unsafe keys. |
| The file opens as one column or breaks at 12,50 | Check delimiter and quoted fields before cleaning. Compare the semicolon sample. | Never globally replace punctuation. A decimal comma, separator and comma inside a quoted name have different jobs. |
Worked example
Boundarysline sample materials use fictional data. Download them to follow the steps; they are not real customer records or certified outputs.
| orders-source.csv | 4 × 6 |
|---|---|
| orders-clean.csv | 3 × 6 |
| SKU | 000123 |
| UTF-8 | Café cup |