Are you letting AI in?
Your robots.txt quietly decides whether the crawlers behind ChatGPT, Claude, Gemini, Perplexity and more are allowed to read your site. This tool checks it against 2026’s 13 major AI crawlers and shows a plain allow/block matrix, split between crawlers that fetch pages to answer live AI questions, crawlers that only gather data for model training, and crawlers that do both, so you can tell the difference between opting out of training and accidentally shutting yourself out of AI answers. It runs entirely in your browser, no API key, nothing sent to us.
What is AI Crawler Auditor?
The AI Crawler Auditor is a free, in-browser tool that reads your site's robots.txt and shows a plain allow/block matrix for 2026's 13 major AI crawlers; GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, Googlebot, CCBot, Bytespider, Amazonbot, Amzn-SearchBot, Amzn-User, and meta-externalagent. It separates crawlers that fetch pages to answer live AI questions from crawlers that only gather data for model training (and flags the ones, like Amazonbot, that do both), so you can see whether you are opting out of training or accidentally shutting yourself out of AI answers.
How it works
Enter your domain and the tool tries to fetch your public /robots.txt over HTTPS. Because browsers block most cross-site requests, if the fetch fails you simply open your robots.txt, copy it, and paste it in; the analysis is identical either way.
Everything is parsed on your device: each AI crawler's documented user-agent is checked against your rules (merging every matching robots.txt group and evaluating wildcard, terminal-$, and longest-match precedence) and marked allowed, partially allowed, or blocked, with a plain-English recommendation. Nothing is uploaded and no API key is needed.
Who it’s for
For site owners, marketers, and developers who want a fast, private read of which AI crawlers their robots.txt lets in. The outcome is clarity: keep live-retrieval crawlers allowed so you stay eligible to be cited in AI answers, and make the model-training opt-out (GPTBot, Google-Extended, CCBot, and similar) a deliberate, separate decision instead of an accident.
In practice
A clinic pastes in its robots.txt and sees a blanket Disallow rule is blocking OAI-SearchBot and PerplexityBot; telling those compliant crawlers not to fetch its pages, which can sharply reduce how often the clinic appears when people ask ChatGPT or Perplexity for a local provider. (robots.txt is voluntary guidance: engines may still mention a brand from previously indexed or third-party sources.) The matrix shows exactly which crawlers are blocked and why, and confirms its GPTBot training opt-out is fine to keep while the live-retrieval crawlers should be reopened.
Who can read your site, and who you shut out.
Enter your domain and we’ll try to fetch your /robots.txt. Browsers usually block cross-site requests to another domain’s files, so if the fetch fails, just open your robots.txt, paste it in, and analyze. Everything is parsed on your device against the documented user-agents behind ChatGPT, Claude, Perplexity, Google-Extended (which controls Gemini training and grounding, not AI Overviews), Googlebot (which powers Google Search and AI Overviews), Common Crawl, ByteDance, Amazon, and Meta. robots.txt governs only compliant bots, content already in a model’s training data persists regardless.
Paste your robots.txt instead (used if the browser can’t fetch it)
Open your robots.txt ↗, select all, copy, and paste it below.
Nothing is uploaded or stored, the file is parsed entirely in your browser. Remember: robots.txt is honored voluntarily (well-behaved crawlers respect it, some ignore it), and content already in a model’s training set stays there regardless. Want this watched for you? Talk to us.
Questions, answered.
Yes, it is completely free, with no sign-up and no API key. It reads your robots.txt rules and shows which AI crawlers you allow or block.
No. The tool runs entirely in your browser. When you enter your domain or paste your robots.txt, it is parsed on your device and never uploaded, logged, or stored by us.
Browsers block most cross-site requests to another domain’s files for security reasons, so a direct fetch of your robots.txt often fails. When that happens, you can open your robots.txt, copy it, and paste it in, and the tool analyzes it exactly the same way.
It checks the documented user-agents for 2026’s major AI crawlers, including GPTBot and OAI-SearchBot (OpenAI), ClaudeBot and Claude-SearchBot (Anthropic), PerplexityBot, Google-Extended and Googlebot (Google), CCBot (Common Crawl), Bytespider (ByteDance), Amazonbot, Amzn-SearchBot, and Amzn-User (Amazon), and meta-externalagent (Meta).
Training crawlers (like GPTBot, ClaudeBot, Google-Extended, CCBot) gather data that may be used to train models. Search-retrieval crawlers (like OAI-SearchBot, Claude-SearchBot, PerplexityBot) fetch pages to answer live questions. Some crawlers, like Amazonbot, do both. Blocking retrieval crawlers can remove you from AI answers now, while blocking training-only crawlers is a separate, valid opt-out choice.
No. Google-Extended only controls whether your content can be used for Gemini model training and grounding. AI Overviews are powered by the standard Googlebot index, so blocking Google-Extended does not remove you from AI Overviews, and allowing it does not add you to them.
No. It reports a site-level read of your robots.txt against documented user-agents. robots.txt is honored voluntarily, path-level rules can still restrict sections, bots can change, and content already in a model’s training data stays there regardless.
