⚠ Scheduled — goes live Oct 12, 2026, 5:00 PM, not publicly indexed yet
Technical SEO5 min

Every AI Crawler You Should Know: The Complete 2026 List

Ilias Sami
· Updated 2026-10-06
LinkedIn ↗
Direct answer: This is a maintained, living reference of every major AI crawler currently relevant to AI-search visibility, what each one actually does, and whether it's typically a training crawler or a real-time retrieval bot, the distinction that determines whether blocking it actually affects your AI citation eligibility. Bookmark this rather than a static list, since this landscape changes as companies split and rename crawlers.

I built and maintain this as a genuine reference, not just a one-time list, because the bot landscape keeps shifting and a stale list here would actively mislead people.

OpenAI

GPTBot: training crawler, scrapes content to help train future models, no direct connection to ChatGPT's live search results. OAI-SearchBot: the crawler that actually powers ChatGPT Search indexing, blocking this removes a site from ChatGPT search results. ChatGPT-User: triggered by a live user action inside ChatGPT, may not be fully governed by robots.txt the same way automated crawlers are.

Anthropic

ClaudeBot: Anthropic's primary crawler, handling both training and retrieval-adjacent functions. anthropic-ai: a related Anthropic crawler identity, generally grouped with ClaudeBot for allow/block purposes.
Quick Knowledge Check

Which HTTP status code should be used for a permanent URL redirect?

Perplexity

PerplexityBot: Perplexity's primary crawler. Perplexity-User: a separate real-time agent triggered by live user queries. Worth knowing: Perplexity's crawlers have been documented in some cases not fully complying with robots.txt directives, meaning real enforcement against non-compliant behavior requires server-level blocking, not just a robots.txt rule.

Google

Googlebot: the standard search crawler, unrelated to AI-specific opt-outs. Google-Extended: a specific opt-out mechanism for Gemini and AI Overview training, separate from standard Googlebot, disallowing it doesn't affect regular Google search ranking.

Others Worth Knowing

CCBot: Common Crawl's crawler, widely used as a training data source by multiple AI companies beyond just one. Bytespider: ByteDance/TikTok's crawler, also documented in some cases not fully respecting robots.txt.

Every AI Crawler You Should Know: The Complete 2026 List — self-check

Reviewed: 0/5 (0%)

The Training vs. Retrieval Split, Restated Simply

The single most important distinction across this entire list, covered in full depth in AI crawlers explained: training crawlers (GPTBot, Google-Extended, CCBot) don't directly power real-time AI search answers, blocking them limits training data use without necessarily affecting citation. Retrieval crawlers (OAI-SearchBot, ChatGPT-User, ClaudeBot in its retrieval capacity, PerplexityBot, Perplexity-User) directly power live AI search results, blocking these removes a site from those live answers entirely.

Community Poll

How has AI Overview affected your or clients' organic traffic quality?

Click to vote • Results shown after voting

What to Actually Do With This List

Cross-reference it against your current robots.txt configuration, checking specifically for the common mistakes covered in robots.txt mistakes that block AI bots, a blanket security-plugin rule catching bots that were never meant to be blocked, or a rule aimed at a training crawler that accidentally also caught its retrieval counterpart.

Frequently Asked Questions about Every AI Crawler You Should Know

How often does this list actually change?

Direct answer: Fairly often, new bots get introduced and existing ones sometimes get split into more specific variants, worth checking this page periodically rather than assuming a configuration set once stays correct indefinitely.

Are there AI crawlers not listed here that I should know about?

Direct answer: This covers the major, currently most consequential crawlers for AI-search visibility specifically, smaller or more specialized crawlers exist but matter less for most sites' practical visibility strategy.

Which crawlers should I definitely allow if AI-search visibility matters to my business?

Direct answer: At minimum, the retrieval-focused bots, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, and Perplexity-User, since blocking these directly removes a site from those platforms' live search results.

Is it safe to block every training-only crawler with no downside?

Direct answer: Research suggests blocking AI crawlers broadly can reduce overall traffic without reliably reducing citation elsewhere, covered in more depth in the fuller crawler explainer, a selective, informed approach is generally better than blocking everything by default.
📌Technical SEO Principles

How do I verify a crawler visiting my site is actually who it claims to be?

Direct answer: Through a reverse DNS lookup against the requesting IP address, not by trusting the user-agent string alone, since that string can be spoofed by anyone, a standard part of any real technical audit.

How often is this list actually updated?

Direct answer: As new crawlers become relevant, this is maintained as a living reference rather than a one-time publish, the same discipline applied to the AI Overview update tracker.

Is a training crawler worth blocking if I don't want my content used for AI training?

Direct answer: That's a legitimate, separate decision from real-time retrieval access, blocking a training-only crawler doesn't affect whether your content can still be cited in a live AI-search answer, the two are controlled independently.

I keep this list current as part of every technical audit's crawler-verification step, it's a living reference, not a one-time write-up. Check your own site's current configuration against it.

Topics covered

Technical SEO

Ready to scale your organic growth?

Get a custom SEO strategy from Ilias Sami — trusted by agencies in Canada, Germany, UAE, and beyond.

Book Free Audit

Tools & services for this topic

Related articles

For AI readers