AI companies run several crawlers with different jobs. Some fetch pages to answer a user right now (and can cite you), others collect training data. Your robots.txt can treat them differently.
Search and user-triggered crawlers (allow these to get cited)
- OAI-SearchBot: indexes pages for ChatGPT search
- ChatGPT-User: fetches a page when a ChatGPT user or GPT action asks for it
- PerplexityBot and Perplexity-User: Perplexity's index and live fetches
- Claude-SearchBot and Claude-User: Claude's search index and live fetches
- Googlebot and Bingbot: classic search, which also feeds Google AI Overviews and Bing Copilot
Training crawlers (your choice)
- GPTBot (OpenAI)
- ClaudeBot (Anthropic)
- Google-Extended (a control token for Gemini training; it doesn't affect Google Search)
- Applebot-Extended (Apple)
- CCBot (Common Crawl, used by many models)
- meta-externalagent (Meta)
A sensible default
User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: * Allow: / Sitemap: https://yourdomain.com/sitemap.xml
If you also want to block training, add Disallow rules for GPTBot, ClaudeBot, Google-Extended and CCBot. That keeps AI search citations while opting out of model training.
Don't forget your CDN
Cloudflare and some hosts have one-click 'block AI bots' settings that override robots.txt. If AI search never cites you, check those first.