Robots.txt generator for AI crawlers
Decide which AI companies can train on your site and which can cite it. Pick a preset, adjust, and copy the file. Nothing leaves your browser.
Checked = blocked.
# AI crawlers blocked User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: Meta-ExternalAgent User-agent: Bytespider User-agent: Amazonbot Disallow: / # AI crawlers allowed User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: ChatGPT-User User-agent: Claude-User User-agent: Perplexity-User Allow: / # Everyone else User-agent: * Allow: /
Upload the file to your site root so it loads at yourdomain.com/robots.txt. Compliant crawlers follow it; it is a request, not access control.
Training bots and search bots are different
Training crawlers such as GPTBot and ClaudeBot collect pages to improve future models. Search crawlers such as OAI-SearchBot and PerplexityBot build the index an assistant pulls live answers and citations from. Blocking a training bot keeps your content out of model training. Blocking a search bot keeps you out of that assistant’s answers.
The default preset blocks training and leaves search open. Google-Extended only controls Gemini training and grounding: it does not affect Google Search or AI Overviews.
Once your robots.txt is live, check whether the assistants actually cite you.
Do AI assistants cite you?
Track Google rank and captured AI citations for the same keywords.