Robots.txt Tester
Fetch or paste a robots.txt file, then test URLs against it for any crawler and see exactly which rule decides the result.
Source
Paste or type a robots.txt file below. It is tested in your browser and never uploaded.
Test URLs
Results
Add URLs or paths to test, one per line.
| URL | Result | Deciding rule | Group used |
|---|
File with line numbers
Sitemaps
Warnings
How to use this robots.txt tester
This robots.txt checker answers one question quickly: is a URL blocked by robots.txt for a given crawler, and which line decides it?
-
Choose Fetch from site and enter a domain. The tool requests
/robots.txtfrom that host and shows the HTTP status, size, redirects and what the status means for crawling. To check a draft before you upload it, choose Paste robots.txt. - Pick a crawler: Googlebot, Bingbot, one of the AI crawlers, or any custom user-agent token.
- Enter URLs or paths, one per line. Each gets an Allowed or Blocked result, the rule that decided it with its line number, and the group of rules that was used.
- Edit the file to try a fix; results update as you type. The warnings list works as a robots.txt validator for typos, unsupported fields and malformed lines.
Google retired the robots.txt Tester in Search Console in November 2023 and replaced it with the robots.txt report, which shows the robots.txt files Google found for your site, when it last crawled them, and any warnings or errors. For one URL on a site you own, the URL Inspection tool shows whether Google is allowed to crawl it.
Robots.txt syntax essentials
A robots.txt file is UTF-8 plain text at the root of a host. Each line is field: value, and anything
after # is a comment. Rules come in groups: one or more User-agent lines followed by
Allow and Disallow rules. Sitemap lines stand on their own and apply to the
whole file.
User-agent: *
Disallow: /cart/
Disallow: /*?sessionid=
Allow: /cart/help
Sitemap: https://example.com/sitemap.xml
Field names and user-agent values are case-insensitive, but paths are case-sensitive and must start with
/. An empty Disallow: blocks nothing. The sitemap URL must be absolute; if you need a
sitemap to point to, the Sitemap Generator builds one from a list of URLs.
How Google decides between conflicting rules
Google compares the path and query string of the URL with every rule in the group. The longest matching rule wins, measured by the length of its path. When an Allow and a Disallow rule are equally long, Google uses the less restrictive one, so Allow wins the tie. The order of rules in the file does not matter. These examples come from Google's documentation:
| URL | Rules | Result |
|---|---|---|
/page | Allow: /p, Disallow: / | Allowed: the Allow rule is longer |
/folder/page | Allow: /folder, Disallow: /folder | Allowed: a tie goes to Allow |
/page.htm | Allow: /page, Disallow: /*.htm | Blocked: the Disallow rule is longer |
/page.php5 | Allow: /page, Disallow: /*.ph | Allowed: both are five characters |
/ | Allow: /$, Disallow: / | Allowed |
/page.htm | Allow: /$, Disallow: / | Blocked: /$ only matches the root |
User-agent groups and specificity
A crawler obeys exactly one group: the one whose user-agent matches it most specifically. Googlebot-News uses a
googlebot-news group if there is one, then a googlebot group, and only then
*. Googlebot-Image and Googlebot-Video fall back to Googlebot in the same way, and Apple says Applebot
follows Googlebot rules when a file does not mention Applebot. Storebot-Google has no such fallback.
If several groups name the same crawler, their rules are merged. A named group is never combined with the
* group, though: once you add a User-agent: Googlebot section, Googlebot ignores every
rule under *, so repeat any rules it still needs. Version numbers are ignored, so
Googlebot/2.1 is read as googlebot.
Wildcards: * and $
* matches any sequence of characters, including none, and $ anchors the end of the URL.
Disallow: /*.pdf$ blocks URLs ending in .pdf but not /file.pdf?download=1, because the
query string is part of what is matched. A trailing * changes nothing: /fish* equals
/fish. To match a literal asterisk or dollar sign, write it percent-encoded as %2A or
%24.
What robots.txt does not do
Robots.txt controls crawling, not indexing. A blocked URL can still appear in Google results if other pages link to
it, usually without a description, because Google cannot read the page. To keep a page out of search results, let
it be crawled and add a noindex robots meta tag or an X-Robots-Tag HTTP header. You can
confirm the header is actually sent with the HTTP Header Checker. Google
stopped honouring noindex lines inside robots.txt in 2019.
Common mistakes
- Blocking CSS and JavaScript. Google renders pages much like a browser. If the files a page needs are disallowed, Google may not see the page the way visitors do.
- Trailing-slash errors.
Disallow: /privatealso blocks/private-offersand/private.html, whileDisallow: /private/does not block/privateitself. - Wrong case.
Disallow: /Admin/does not block/admin/. - Shipping the staging file. A
User-agent: *plusDisallow: /file that protected a staging site can block the whole live site after a deploy. Fetch the live file after each release, and protect staging with a password instead.
AI crawler tokens
Each AI company publishes its own tokens, and each is controlled separately, so you can allow search crawlers while opting out of model training:
- OpenAI:
GPTBotcollects training data,OAI-SearchBotcrawls for ChatGPT search, andChatGPT-Userfetches pages when a user asks; OpenAI says robots.txt rules may not apply to those user-initiated visits. - Anthropic:
ClaudeBotcollects training data,Claude-SearchBotcrawls for search, andClaude-Userfetches pages for user requests. - Google:
Google-Extendedis not a separate crawler. It controls whether content Google crawls may be used to train Gemini models and for grounding, and it does not affect Google Search. - Apple:
Applebot-Extendedcontrols use for training Apple's foundation models. It does not crawl, and pages that disallow it can still appear in Apple's search results. - Common Crawl:
CCBotbuilds the open Common Crawl web archive.
How 4xx and 5xx robots.txt responses are treated
The status code of the robots.txt file matters as much as its contents. Google follows at least five redirects and then treats the file as not found. Any 4xx response except 429, including 401 and 403, means there is no robots.txt, so everything may be crawled. A 429, a 5xx or a network error such as a timeout or DNS failure is treated as a server error: Google stops crawling the site for the first 12 hours, then relies on its last good copy for up to 30 days. This tester applies the same logic to the status of a fetched file and tells you which case you are in.
Frequently asked questions
Does robots.txt apply to subdomains?
No. A robots.txt file only covers the protocol, host and port it is served from. The file at https://example.com/robots.txt does not cover blog.example.com, www.example.com or http://example.com; each needs its own /robots.txt.
How long does Google take to notice a robots.txt change?
Google generally caches robots.txt for up to 24 hours, so changes usually take effect within a day. In an emergency you can request a recrawl of the file from the robots.txt report in Search Console.
Can I use robots.txt to hide private pages?
No. Anyone can read robots.txt and only well-behaved crawlers obey it, so listing private paths advertises them to everyone else. Protect private content with a login instead.
Does Google support Crawl-delay?
No. Google ignores the Crawl-delay line and adjusts its crawl rate based on how your server responds. Bing and some other crawlers do honour it.
How big can a robots.txt file be?
Google reads the first 500 KiB of a robots.txt file and ignores everything after that. If your file is larger, combine similar rules with wildcards and remove rules you no longer need.