robots.txt generator
Create a robots.txt that controls AI crawlers. Choose which AI search, answer and training crawlers may visit, block private folders, add your sitemap, then copy or download the file.
Start from a preset
AI search & answers
OAI-SearchBot · OpenAI
Finds pages to show and link in ChatGPT search results.
Claude-SearchBot · Anthropic
Indexes pages to improve Claude's search results.
PerplexityBot · Perplexity
Indexes pages that Perplexity can show and cite in answers.
Googlebot · Google
Google's main crawler. Pages it can't crawl can't appear in Search or AI Overviews.
Bingbot · Microsoft
Bing's crawler. Copilot answers are grounded in Bing's index.
DuckAssistBot · DuckDuckGo
Fetches pages used in DuckDuckGo's AI-assisted answers.
User-requested visits
ChatGPT-User · OpenAI
Visits a page when a ChatGPT user asks it to.
Claude-User · Anthropic
Visits a page when a Claude user asks it to.
Perplexity-User · Perplexity
Visits a page when a Perplexity user asks a question about it.
AI model training
GPTBot · OpenAI
Collects pages that may be used to train OpenAI's models.
ClaudeBot · Anthropic
Collects pages that may be used to train Anthropic's models.
Google-Extended · Google
Controls whether Google may use your pages for Gemini. It doesn't affect Google Search.
Applebot-Extended · Apple
Controls whether Apple may use your pages to train its models. Applebot still crawls for Siri and Spotlight.
CCBot · Common Crawl
Builds the open Common Crawl dataset, which many AI models are trained on.
Meta-ExternalAgent · Meta
Collects pages that may be used to train Meta's AI models.
Amazonbot · Amazon
Amazon's crawler, used for services such as Alexa.
Bytespider · ByteDance
ByteDance's crawler, used for its AI models.
Folders no crawler should visit
For example your admin area or internal API. Leave empty to allow everything.
robots.txt
User-agent: * Disallow: /admin Disallow: /api # OpenAI: OpenAI models User-agent: GPTBot Disallow: / # Anthropic: Anthropic models User-agent: ClaudeBot Disallow: / # Google: Gemini User-agent: Google-Extended Disallow: / # Apple: Apple Intelligence User-agent: Applebot-Extended Disallow: / # Common Crawl: Common Crawl User-agent: CCBot Disallow: / # Meta: Meta AI User-agent: Meta-ExternalAgent Disallow: / # Amazon: Amazon User-agent: Amazonbot Disallow: / # ByteDance: ByteDance User-agent: Bytespider Disallow: /
Upload it to the root of your site so it loads at yourwebsite.com/robots.txt. robots.txt is a request that well-behaved crawlers follow; it doesn't lock pages.
Questions
What is a robots.txt file?
A plain-text file at yourwebsite.com/robots.txt that tells crawlers which parts of your site they may visit. Each group starts with User-agent and lists Disallow or Allow rules. Well-behaved crawlers follow it; it doesn't password-protect anything.
How do I block AI crawlers from training on my site?
Add a group for each training crawler with Disallow: /, for example User-agent: GPTBot, User-agent: ClaudeBot, User-agent: Google-Extended and User-agent: CCBot. The "Allow AI search, block AI training" preset does this while keeping AI search crawlers allowed.
Will blocking AI crawlers hurt my Google rankings?
Not if you keep Googlebot allowed. Google-Extended controls only whether Google may use your content for Gemini; Google Search and AI Overviews use Googlebot.
Where do I put robots.txt?
At the root of your domain, so it loads at https://yourwebsite.com/robots.txt. Each subdomain needs its own file.
How do crawlers read the rules?
A crawler uses the group that names it, or the User-agent: * group if none does. Within that group, the longest matching path rule wins, and Allow wins a tie. That's why this generator gives blocked crawlers their own Disallow: / group.
More free tools: AI crawler checker, llms.txt generator, FAQ schema generator.