Remove Duplicate Lines: How It Works
Removing duplicate lines is a one-click operation with three decisions hiding inside it: whether case matters, whether surrounding whitespace matters, and whether the original order must survive. Getting those wrong silently corrupts data.
The three decisions
| Option | When it matters |
|---|---|
| Case sensitivity | Apple and apple — the same entry in an email list, different entries in a list of identifiers |
| Trim whitespace | john@x.com and john@x.com look identical and are not |
| Preserve order | Sorting reorders your data; keeping first occurrence does not |
Trailing whitespace is the one that catches people. A trailing space is invisible, and a list of a thousand entries copied from a spreadsheet frequently contains dozens. Without trimming, every one survives deduplication as a phantom duplicate.
Order matters more than it looks
Deduplicating by sorting is the fastest approach and destroys the sequence. If your list is a log, a ranked set, or paired with anything positional, that is data loss. Keeping the first occurrence and discarding later ones preserves the original order and is almost always the safer default.
There is also a choice between keeping the first and keeping the last occurrence. For a log of updates where later entries supersede earlier ones, keeping the last is correct; for a priority-ordered list, keeping the first is.
Normalisation before comparison
Some duplicates are not textually identical. Depending on the data, you may want to normalise before comparing:
- Email addresses — lowercase the domain always; the local part is technically case-sensitive but treated as insensitive by essentially every provider.
- URLs — trailing slashes,
httpversushttps, andwwwprefixes produce four spellings of one address. - Phone numbers — spaces, hyphens, brackets and country-code formats.
- Unicode — an accented character can be one code point or a letter plus a combining mark. These look identical and compare as different, which is what Unicode normalisation forms exist to resolve.
Before you deduplicate
- Keep the original. Deduplication is not reversible, and a bad option choice is only visible afterwards.
- Note the line count before and after — the number removed is a sanity check. Removing 90% of a list usually means an option was wrong, not that the list was 90% duplicates.
- Decide whether blank lines should be collapsed, kept, or treated as data.
- Spot-check the output against a few known entries.
Related operations
Sometimes you want the opposite: only the lines that appear more than once, to see what was duplicated rather than to remove it. And sometimes you want the difference between two lists rather than within one — entries in A but not in B — which is a comparison task rather than a deduplication one.
All of this runs locally in your browser. Lists of email addresses, customer records and identifiers are exactly the kind of data that should not be pasted into a server-side tool.