Remove Duplicate Lines
Paste a list to delete repeated lines. You choose whether case and extra spaces count, and nothing leaves your browser.
For repeated lines. Adds the count and a tab before each line, so the result pastes into a spreadsheet as two columns.
A to Z puts item2 before item10. Most repeated first works with Only repeated lines.
Result
How to remove duplicate lines
Paste your list into the box, one item per line, or press Load example to try a sample mailing
list. The result updates as you type and keeps the first copy of every line in its original position, with
counts of the lines in, the lines out and the duplicates removed. Copy the result or download it as a
.txt file.
Everything runs in your browser, so nothing you paste is uploaded. Windows (\r\n), Unix
(\n) and old Mac (\r) line endings are all read as line breaks, and a final line break
doesn't add an empty line. Lists over about 2 MB update when you press Remove duplicates
instead of as you type.
Ignore case and ignore extra spaces
By default, two lines are duplicates only when they match character for character.
Ignore case treats Apple and apple as the same line.
Ignore extra spaces skips spaces at the start and end of a line and treats a run of spaces or
tabs inside it as a single space. Neither option changes the text you get back: the kept copy is shown exactly
as you typed it. In the table, ␣ marks a space.
| Two lines | Default | Ignore case | Ignore extra spaces | Both |
|---|---|---|---|---|
| apple / apple | Same | Same | Same | Same |
| Apple / apple | Different | Same | Different | Same |
| apple␣ / apple | Different | Different | Same | Same |
| red␣␣apple / red␣apple | Different | Different | Same | Same |
| ␣APPLE / apple | Different | Different | Different | Same |
| redapple / red␣apple | Different | Different | Different | Different |
Trim lines in the output removes spaces and tabs at both ends of every line before lines are compared, so the result never holds two lines that look identical. Remove empty lines drops blank and whitespace-only lines. When it is off, all the empty lines collapse into one.
Keep the first or the last copy
With First copy, each line stays where it first appears. With Last copy, it moves to where it last appears, and if the copies differ in case or spacing you get the last spelling. That suits lists that grow over time, where the newest entry is the one that counts.
| Input | Keep first copy | Keep last copy |
|---|---|---|
| a, b, a, c, b | a, b, c | a, c, b |
| Red, blue, red | Red, blue | blue, red |
Commas separate the lines here. The second row uses Ignore case, so Red and red count as one line.
Repeated lines, single lines and sorting
Only repeated lines lists one copy of each line that appears two or more times. Tick
Show counts to put the number of copies in front of each line, followed by a tab, for example
3, a tab, then apple. The tab makes the result paste into a spreadsheet as two
columns. Lines that appear once does the opposite and lists only the lines with no copies.
A to Z and Z to A sort in natural order, so item2 comes before
item10. Most repeated first is available with Only repeated lines and orders by
count; lines with the same count keep their original order.
Common uses
- Email lists. Merge exports from several tools and dedupe with Ignore case and Ignore extra spaces. The domain part of an address is not case-sensitive, and most providers ignore case before the @ too.
- Keyword lists. Combine suggestions from several keyword tools. Only repeated lines with counts shows which keywords more than one source agrees on.
- URL lists. Clean up crawl exports before building a redirect map or a sitemap. Leave Ignore case
off: paths are case-sensitive on many servers, so
/Pageand/pagecan be different URLs. - Log lines. Only repeated lines, Show counts and Most repeated first give you the most frequent messages. Timestamps make every line unique, so cut them off first. To look at one kind of message only, keep the matching lines with Filter Lines.
To find what one list has that another doesn't, rather than repeats inside a single list, use Compare Lists.
Removing duplicates in Excel and Google Sheets
In Excel, select the column and choose Data > Remove Duplicates. It keeps the first copy, ignores case and
edits the sheet in place, so work on a copy. In Excel 2021 and Microsoft 365, =UNIQUE(A2:A100)
returns the deduplicated list without touching the original, and =UNIQUE(A2:A100,,TRUE) returns the
values that appear exactly once. Neither ignores extra spaces: clean the column with TRIM first,
which removes spaces at the ends and reduces runs of spaces inside to one.
In Google Sheets, =UNIQUE(A2:A100) keeps the first copy of each value in order, but unlike Excel it
is case-sensitive. To ignore case, wrap the range:
=UNIQUE(ARRAYFORMULA(LOWER(A2:A100))). The result comes back in lower case.
Removing duplicates on the command line
sort -u list.txt removes duplicates but sorts the output, so the original order is lost. Add
-f to ignore case. To keep the order, use awk:
awk '!seen[$0]++' list.txt > unique.txt
It prints a line only the first time it sees it. Use seen[tolower($0)] to ignore case. To keep
the last copy instead, reverse the file, dedupe it and reverse it back with
tac list.txt | awk '!seen[$0]++' | tac (tac is part of GNU coreutils).
sort list.txt | uniq -d lists repeated lines, the same pipeline with uniq -u lists
the lines that appear once, and sort list.txt | uniq -c | sort -rn counts every line, most
repeated first. One catch: a file with Windows line endings keeps a hidden \r at the end of each
line, so a line from that file won't match the same line from a Unix file. This tool reads both the same way.
Frequently asked questions
How do I remove duplicate lines from a list?
Paste the list into the box with one item per line. Repeated lines are removed as you type, and the first copy of each line stays in its original position. Copy the result or download it as a text file.
Why are some duplicates not removed?
The lines usually differ in something hard to see, such as a trailing space, a tab or a capital letter. Tick Ignore extra spaces and Ignore case to treat those lines as the same.
Is removing duplicate lines case-sensitive?
By default, yes: Apple and apple are kept as two lines. Tick Ignore case to treat them as one. The copy you keep is shown exactly as it was typed.
How do I find which lines are duplicated and how many times?
Choose Only repeated lines to list one copy of each line that appears more than once. Tick Show counts to add the number of copies, and sort by Most repeated first to see the most common lines at the top.
Does removing duplicates keep the original order?
Yes, unless you pick a sort. Each line stays where it first appears, or where it last appears if you choose to keep the last copy.
Is my list uploaded anywhere?
No. The tool runs in your browser and nothing you paste is sent to a server. Lists under 200 KB are saved in your browser's local storage so they are still there when you come back.