# Which crawlers feed AI answers, and how do you allow them?

The crawlers that matter for being cited are the ones that build AI answer indexes: GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI), ClaudeBot and anthropic-ai (Anthropic), PerplexityBot (Perplexity), Google-Extended (Google's AI grounding) and CCBot (Common Crawl, which feeds many models). Allowing them is a robots.txt decision — and the decision is per crawler, so a blanket block on one can silently remove you from one engine's answers while leaving the others intact.

## Key points

- Allow by name, one block per crawler, and keep a permissive fallback for the rest. A file that names them is also a statement you can point at when asked how your content may be used.
- Check what your host already does. On 2026-10-07 this site's own robots.txt allowed every crawler above with no blocks, while two well-known publishing platforms did not: the crawler feeding AI answers was refused outright on both, and one of them refused all of the crawlers tested.
- robots.txt permission is necessary and not sufficient. A page that returns almost no text — a client-rendered shell, a login wall, an image — gives the crawler nothing to quote, which is the same outcome as being refused.
- Publishing platforms are not neutral carriers: where a platform refuses the crawlers, content hosted there cannot be cited from there, however good it is.
- Re-check robots.txt periodically. Sites that never mentioned these crawlers in the past are the ones most likely to add a blanket block later, and the change is silent.

## Frequently asked questions

**Q: Does allowing AI crawlers cost us traffic?**
A: It changes where the traffic comes from rather than removing it: a cited answer usually links the page it used. The honest trade-off is that a page may be summarised without a visit, which is why the decision belongs to the site owner and should be explicit rather than accidental.

**Q: What if our host sets robots.txt for us?**
A: Then the effective file is the host's, not yours, and the first step is to fetch it and read the rules for the crawlers above. Managed platforms frequently ship a template that blocks them without the site owner ever choosing that.

## Sources

- https://autoaeo.work/aeo/answers/what-is-answer-engine-optimization.md
- https://autoaeo.work/aeo/answers/citation-stability.md
- https://autoaeo.work/llms.txt

Updated 2026-10-07. Canonical source: https://autoaeo.work/aeo/answers/robots-for-ai-crawlers.md
