AI crawler checker
See whether ChatGPT, Claude, Perplexity, Google and 13 other crawlers can reach and read your website. Enter a URL to check robots.txt rules for each crawler, content visible without JavaScript, noindex directives, structured data, sitemap and llms.txt.
What the checker looks at
- robots.txt, per crawler. The rules each of 17 crawlers would follow for your page, including which line decided it.
- Content without JavaScript. How many words are in the HTML your server sends, which is all most AI crawlers read.
- Server responses. What your server returns when asked by AI crawler user agents, including bot challenge pages.
- Indexing signals. noindex and nosnippet in meta tags and the X-Robots-Tag header, title, meta description, H1, canonical and language.
- Discovery. Structured data (JSON-LD), sitemap, llms.txt and Open Graph tags.
Questions
How do I know if AI crawlers can access my website?
Enter your address above. The checker reads your robots.txt the same way crawlers do (following the RFC 9309 standard) for each AI crawler, loads your page to see how much content is readable without JavaScript, looks for noindex and nosnippet directives, and tests whether your server answers requests from AI crawler user agents.
Why does it matter if my content only loads with JavaScript?
Google renders JavaScript, but most AI crawlers read only the HTML your server sends. If your content appears only after scripts run, those crawlers see a nearly empty page and can't quote or cite you. Server-side rendering or static generation fixes it.
What's the difference between AI search crawlers and AI training crawlers?
Search and user crawlers, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, fetch pages so AI tools can show and cite them in answers. Training crawlers, such as GPTBot and ClaudeBot, collect pages that may be used to train models. You can allow one group and block the other in robots.txt.
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and visits a user asks for use ChatGPT-User. To appear in ChatGPT search results, allow OAI-SearchBot.
Does Google-Extended affect Google Search rankings?
No. Google-Extended only controls whether Google may use your content for Gemini. Google Search and AI Overviews use Googlebot.
Why could my server block AI crawlers when robots.txt allows them?
Firewalls and bot-protection services can reject requests by user agent or show challenge pages. The checker requests your page with GPTBot, ClaudeBot, PerplexityBot and Googlebot user agents and reports what your server returns. Real crawlers come from their own IP addresses, so confirm any block in your firewall settings.
More free tools: robots.txt generator, llms.txt generator, FAQ schema generator.