Free tool

Robots.txt generator for AI crawlers

Decide which AI companies can train on your site and which can cite it. Pick a preset, adjust, and copy the file. Nothing leaves your browser.

Model training

Collects pages to train future models. Blocking these does not remove you from AI answers.

AI search & citations

Builds the index AI assistants cite from. Block these and you can't be cited.

User-requested fetches

Fetches a page when someone pastes it into a chat. Some of these ignore robots.txt.

Checked = blocked.

robots.txt
# AI crawlers blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: Bytespider
User-agent: Amazonbot
Disallow: /

# AI crawlers allowed
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
Allow: /

# Everyone else
User-agent: *
Allow: /

Upload the file to your site root so it loads at yourdomain.com/robots.txt. Compliant crawlers follow it; it is a request, not access control.

Training bots and search bots are different

Training crawlers such as GPTBot and ClaudeBot collect pages to improve future models. Search crawlers such as OAI-SearchBot and PerplexityBot build the index an assistant pulls live answers and citations from. Blocking a training bot keeps your content out of model training. Blocking a search bot keeps you out of that assistant’s answers.

The default preset blocks training and leaves search open. Google-Extended only controls Gemini training and grounding: it does not affect Google Search or AI Overviews.

Once your robots.txt is live, check whether the assistants actually cite you.

Do AI assistants cite you?

Track Google rank and captured AI citations for the same keywords.

Try RankGhost free