Direct answer: This is a maintained, living reference of every major AI crawler currently relevant to AI-search visibility, what each one actually does, and whether it's typically a training crawler or a real-time retrieval bot, the distinction that determines whether blocking it actually affects your AI citation eligibility. Bookmark this rather than a static list, since this landscape changes as companies split and rename crawlers.
I built and maintain this as a genuine reference, not just a one-time list, because the bot landscape keeps shifting and a stale list here would actively mislead people.
OpenAI
GPTBot: training crawler, scrapes content to help train future models, no direct connection to ChatGPT's live search results.
OAI-SearchBot: the crawler that actually powers ChatGPT Search indexing, blocking this removes a site from ChatGPT search results.
ChatGPT-User: triggered by a live user action inside ChatGPT, may not be fully governed by robots.txt the same way automated crawlers are.
Anthropic
ClaudeBot: Anthropic's primary crawler, handling both training and retrieval-adjacent functions.
anthropic-ai: a related Anthropic crawler identity, generally grouped with ClaudeBot for allow/block purposes.
Quick Knowledge Check
Which HTTP status code should be used for a permanent URL redirect?
Perplexity
PerplexityBot: Perplexity's primary crawler.
Perplexity-User: a separate real-time agent triggered by live user queries. Worth knowing: Perplexity's crawlers have been documented in some cases not fully complying with robots.txt directives, meaning real enforcement against non-compliant behavior requires server-level blocking, not just a robots.txt rule.
Google
Googlebot: the standard search crawler, unrelated to AI-specific opt-outs.
Google-Extended: a specific opt-out mechanism for Gemini and AI Overview training, separate from standard Googlebot, disallowing it doesn't affect regular Google search ranking.
Others Worth Knowing
CCBot: Common Crawl's crawler, widely used as a training data source by multiple AI companies beyond just one.
Bytespider: ByteDance/TikTok's crawler, also documented in some cases not fully respecting robots.txt.
Every AI Crawler You Should Know: The Complete 2026 List — self-check
Reviewed: 0/5 (0%)
The Training vs. Retrieval Split, Restated Simply
The single most important distinction across this entire list, covered in full depth in AI crawlers explained: training crawlers (GPTBot, Google-Extended, CCBot) don't directly power real-time AI search answers, blocking them limits training data use without necessarily affecting citation. Retrieval crawlers (OAI-SearchBot, ChatGPT-User, ClaudeBot in its retrieval capacity, PerplexityBot, Perplexity-User) directly power live AI search results, blocking these removes a site from those live answers entirely.
Community Poll
How has AI Overview affected your or clients' organic traffic quality?
Click to vote • Results shown after voting
What to Actually Do With This List
Cross-reference it against your current robots.txt configuration, checking specifically for the common mistakes covered in robots.txt mistakes that block AI bots, a blanket security-plugin rule catching bots that were never meant to be blocked, or a rule aimed at a training crawler that accidentally also caught its retrieval counterpart.
Frequently Asked Questions about Every AI Crawler You Should Know
How often does this list actually change?
Direct answer: Fairly often, new bots get introduced and existing ones sometimes get split into more specific variants, worth checking this page periodically rather than assuming a configuration set once stays correct indefinitely.
Are there AI crawlers not listed here that I should know about?
Direct answer: This covers the major, currently most consequential crawlers for AI-search visibility specifically, smaller or more specialized crawlers exist but matter less for most sites' practical visibility strategy.
Which crawlers should I definitely allow if AI-search visibility matters to my business?
Direct answer: At minimum, the retrieval-focused bots, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, and Perplexity-User, since blocking these directly removes a site from those platforms' live search results.
Is it safe to block every training-only crawler with no downside?
Direct answer: Research suggests blocking AI crawlers broadly can reduce overall traffic without reliably reducing citation elsewhere, covered in more depth in the fuller crawler explainer, a selective, informed approach is generally better than blocking everything by default.
📌Technical SEO Principles
How do I verify a crawler visiting my site is actually who it claims to be?
Direct answer: Through a reverse DNS lookup against the requesting IP address, not by trusting the user-agent string alone, since that string can be spoofed by anyone, a standard part of any real technical audit.
How often is this list actually updated?
Direct answer: As new crawlers become relevant, this is maintained as a living reference rather than a one-time publish, the same discipline applied to the AI Overview update tracker.
Is a training crawler worth blocking if I don't want my content used for AI training?
Direct answer: That's a legitimate, separate decision from real-time retrieval access, blocking a training-only crawler doesn't affect whether your content can still be cited in a live AI-search answer, the two are controlled independently.