Filter Lines
Keep only the lines you want from a list, by words, regular expression or length. Nothing is uploaded.
Whole words only: "cat" matches "the cat." but not "category". To drop lines that contain none of the words, choose Keep matching lines with Any of these.
Unicode mode lets \p{L} match a letter in any script and makes . match a whole
emoji. Turn it off only for patterns written for older engines.
Length counts characters (Unicode code points), so an emoji such as 😀 counts as 1.
The text is used exactly as typed, including spaces.
Kept lines
How to filter lines
Paste your list into Your lines, one item per line. Choose Keep matching lines
or Remove matching lines, then pick what a match is: words, a regular expression, a line length,
or text at the start or end of the line. The result updates as you type, with counts of input, kept and removed
lines. Tick Show removed lines to see what the filter threw away, which is the quickest way to
check that a rule does what you meant. Copy the result or download it as a .txt file.
Everything runs in your browser. Nothing you paste is uploaded, so it is safe to use with server logs, customer lists and unpublished keyword research.
Options
- Ignore case (on by default) treats
Erroranderroras the same in words, regex and starts-or-ends-with matching. - Whole words only makes
catmatch "the cat." but not "category". Letters in any script, digits and underscores count as part of a word, similar togrep -w, socafdoes not match "café". Leave it off for languages written without spaces, such as Chinese or Japanese. - Trim lines removes spaces at both ends before matching, and the output is trimmed too.
- Remove empty lines drops blank and whitespace-only lines before filtering.
- Remove duplicates from the result keeps the first copy of each kept line. The comparison is exact and case-sensitive; turn on Trim lines if copies differ only in surrounding spaces.
Examples for each mode
Each row filters these five lines with Keep matching lines, unless it says otherwise:
https://example.com/pricinghttps://blog.example.com/serp-api-guidehttp://example.org/old-pagehttps://example.com/files/report.pdfhttps://www.example.net/a
| Match by | Setting | Lines kept |
|---|---|---|
| Words, any of these | blog and .pdf | 2, 4 |
| Words, all of these | https and example.com | 1, 2, 4 |
| Words, with Remove matching lines | example.com | 3, 5 |
| Regular expression | ^https://(www\.)?example\.com/ | 1, 4 |
| Length, at most | 27 characters | 1, 3, 5 |
| Starts with | http:// | 3 |
| Ends with | 4 |
To keep lines that contain none of your words, use Remove matching lines with Any of these. To keep lines that lack at least one of them, use Remove matching lines with All of these.
Keep only URLs from one domain
A words filter for example.com is quick, but it also keeps blog.example.com,
notexample.com and example.com.evil.net, because it matches the text anywhere in the
line. For an exact host, switch to Regular expression and anchor the pattern:
^https?://(www\.)?example\.com([/?#]|$) keeps example.com and
www.example.com over HTTP or HTTPS and nothing else. To include every subdomain, use
^https?://([^/]+\.)?example\.com([/?#]|$). This is handy for trimming a crawl export
or a sitemap down to one site before you compare it with another list in
Compare Lists.
Remove lines with certain words from keyword lists
Keyword exports are full of terms you will never target. Choose Remove matching lines, Words and Any of these,
and list the negatives one per line: free, jobs, salary, competitor brand
names. Turn on Whole words only so that free removes "free seo tools" but keeps "freelance seo
writer". Add Trim lines and Remove duplicates to tidy the result in the same pass, or run it through
Remove Duplicate Lines for more control over what counts as a copy.
Regex cheat sheet
Patterns use JavaScript syntax, and each line is tested on its own. Unicode mode is on by default: it adds
\p{L} and makes . match a whole emoji, but it rejects needless escapes
such as \- outside brackets. If you see that error, remove the backslash or turn Unicode mode off.
| Pattern | Matches |
|---|---|
| . | Any one character |
| \d \s \w | A digit, a whitespace character, an ASCII letter, digit or underscore |
| \. \? \( | A literal dot, question mark or parenthesis |
| ^ $ | Start and end of the line |
| [abc] [^abc] [a-z] | One of a, b or c; anything except them; a range |
| cat|dog | Either alternative |
| ? * + | Zero or one, zero or more, one or more of the previous item |
| {3} {2,} {2,5} | Exactly 3, 2 or more, 2 to 5 of the previous item |
| (www\.)? | A group, here made optional |
| \b | A boundary between \w and non-\w characters (ASCII only) |
| \p{L} | A letter in any script (Unicode mode) |
A group that repeats and contains a repeat, such as (\w+\s?)+$, can take minutes on a
single 35-character line, because the engine tries every way of splitting it. The tool doesn't run such patterns
as you type; it warns you and waits for Filter lines. It also checks the clock every thousand lines and stops
after about 2 seconds with a message, though a single match already in progress can't be interrupted. For inputs
over about 1 MB, regex filtering runs only when you press Filter lines, and above 2 MB every mode works that way.
Find lines by length
Length mode keeps lines that are at least, at most, exactly or between a number of characters. Use it to find
page titles longer than about 60 characters, meta descriptions over 160, or short junk lines in a scraped list.
Google shortens long titles by pixel width rather than a fixed count, so treat those numbers as rough guides.
Length counts Unicode code points: 😀 is 1, but a flag or an emoji with a skin tone is 2, and an
accent typed as a separate combining mark adds 1. With Trim lines on, the trimmed line is measured.
The same filters with grep and spreadsheets
On the command line, grep keeps matching lines, grep -v removes them and
grep -E enables extended regular expressions. In a spreadsheet, FILTER is available in
Google Sheets and in Excel 365 and 2021 or later; SEARCH ignores case, FIND does not,
and REGEXMATCH is a Sheets function.
| Task | grep | Excel or Google Sheets |
|---|---|---|
| Keep lines containing a word | grep -i 'error' log.txt | =FILTER(A2:A1000, ISNUMBER(SEARCH("error", A2:A1000))) |
| Remove lines containing a word | grep -v -i 'error' log.txt | =FILTER(A2:A1000, NOT(ISNUMBER(SEARCH("error", A2:A1000)))) |
| Any word from a list | grep -i -F -f words.txt log.txt | =FILTER(A2:A1000, REGEXMATCH(A2:A1000, "(?i)error|timeout")) (Sheets) |
| Whole words only | grep -w 'cat' pets.txt | =FILTER(A2:A1000, REGEXMATCH(A2:A1000, "\bcat\b")) (Sheets) |
| Regular expression | grep -E '^https?://(www\.)?example\.com/' urls.txt | =FILTER(A2:A1000, REGEXMATCH(A2:A1000, "^https?://(www\.)?example\.com/")) (Sheets) |
| Length of 60 or less | grep -E '^.{0,60}$' titles.txt | =FILTER(A2:A1000, LEN(A2:A1000)<=60) |
Two differences to watch for. In a grep pattern file, an empty line matches every line, while this tool ignores
blank lines in the word list. And grep has no "all of these"; you chain it instead, as in
grep -i 'red' list.txt | grep -i 'apple'. If you need the text between two markers rather than whole
lines, use Extract Text Between.
Frequently asked questions
How do I remove all lines that contain a certain word?
Choose Remove matching lines, match by Words and type the word. Every line containing it disappears from the result, and Show removed lines lists what was taken out so you can check it.
How do I delete lines that do not contain a word?
Choose Keep matching lines and type the word; everything else is dropped. With several words, Any of these keeps lines that have at least one of them and All of these keeps lines that have every one.
Can I filter lines with a regular expression?
Yes. Pick Regular expression and type a JavaScript pattern without the slashes. Each line is tested separately, the Ignore case option applies, and an invalid pattern shows the reason instead of a result.
Does whole word matching work with accented or non-Latin letters?
Yes. Letters from any script count as part of a word, so a whole-word search for caf does not match café and кот does not match котёнок. Digits and underscores also count as word characters, as they do in grep -w.
How many lines can I filter at once?
The filter runs in your browser and handles 100,000 lines without trouble. Above about 2 MB of text, or 1 MB with a regular expression, results update when you press Filter lines instead of as you type.
Is my text uploaded anywhere?
No. The filtering happens in your browser and nothing you paste is sent to a server.