# Every AI Crawler You Should Know: The Complete 2026 List

_2026-09-16 (updated 2026-10-06) · 5 min · by Ilias Sami · ~765 words_

> This is a maintained, living reference of every major AI crawler currently relevant to AI-search visibility, what each one actually does, and whether it's typically a training crawler or a real-time retrieval bot, the distinction that determines whether blocking it actually affects your AI citation eligibility. Bookmark this rather than a static list, since this landscape changes as companies split and rename crawlers.

**Direct answer:** This is a maintained, living reference of every major AI crawler currently relevant to AI-search visibility, what each one actually does, and whether it's typically a training crawler or a real-time retrieval bot, the distinction that determines whether blocking it actually affects your AI citation eligibility. Bookmark this rather than a static list, since this landscape changes as companies split and rename crawlers.

I built and maintain this as a genuine reference, not just a one-time list, because the bot landscape keeps shifting and a stale list here would actively mislead people.

# OpenAI

**GPTBot**: training crawler, scrapes content to help train future models, no direct connection to ChatGPT's live search results.
**OAI-SearchBot**: the crawler that actually powers ChatGPT Search indexing, blocking this removes a site from ChatGPT search results.
**ChatGPT-User**: triggered by a live user action inside ChatGPT, may not be fully governed by robots.txt the same way automated crawlers are.

# Anthropic

**ClaudeBot**: Anthropic's primary crawler, handling both training and retrieval-adjacent functions.
**anthropic-ai**: a related Anthropic crawler identity, generally grouped with ClaudeBot for allow/block purposes.

# Perplexity

**PerplexityBot**: Perplexity's primary crawler.
**Perplexity-User**: a separate real-time agent triggered by live user queries. Worth knowing: Perplexity's crawlers have been documented in some cases not fully complying with robots.txt directives, meaning real enforcement against non-compliant behavior requires server-level blocking, not just a robots.txt rule.

# Google

**Googlebot**: the standard search crawler, unrelated to AI-specific opt-outs.
**Google-Extended**: a specific opt-out mechanism for Gemini and AI Overview training, separate from standard Googlebot, disallowing it doesn't affect regular Google search ranking.

# Others Worth Knowing

**CCBot**: Common Crawl's crawler, widely used as a training data source by multiple AI companies beyond just one.
**Bytespider**: ByteDance/TikTok's crawler, also documented in some cases not fully respecting robots.txt.

{{BLOCK:0}}

# The Training vs. Retrieval Split, Restated Simply

The single most important distinction across this entire list, covered in full depth in [AI crawlers explained](/blog/ai-crawlers-explained): training crawlers (GPTBot, Google-Extended, CCBot) don't directly power real-time AI search answers, blocking them limits training data use without necessarily affecting citation. Retrieval crawlers (OAI-SearchBot, ChatGPT-User, ClaudeBot in its retrieval capacity, PerplexityBot, Perplexity-User) directly power live AI search results, blocking these removes a site from those live answers entirely.

# What to Actually Do With This List

Cross-reference it against your current robots.txt configuration, checking specifically for the common mistakes covered in [robots.txt mistakes that block AI bots](/blog/robots-txt-mistakes-block-ai-bots), a blanket security-plugin rule catching bots that were never meant to be blocked, or a rule aimed at a training crawler that accidentally also caught its retrieval counterpart.

# Frequently Asked Questions

**How often does this list actually change?**
**Direct answer:** Fairly often, new bots get introduced and existing ones sometimes get split into more specific variants, worth checking this page periodically rather than assuming a configuration set once stays correct indefinitely.

**Are there AI crawlers not listed here that I should know about?**
**Direct answer:** This covers the major, currently most consequential crawlers for AI-search visibility specifically, smaller or more specialized crawlers exist but matter less for most sites' practical visibility strategy.

**Which crawlers should I definitely allow if AI-search visibility matters to my business?**
**Direct answer:** At minimum, the retrieval-focused bots, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, and Perplexity-User, since blocking these directly removes a site from those platforms' live search results.

**Is it safe to block every training-only crawler with no downside?**
**Direct answer:** Research suggests blocking AI crawlers broadly can reduce overall traffic without reliably reducing citation elsewhere, covered in more depth in the fuller crawler explainer, a selective, informed approach is generally better than blocking everything by default.

**How do I verify a crawler visiting my site is actually who it claims to be?**
**Direct answer:** Through a reverse DNS lookup against the requesting IP address, not by trusting the user-agent string alone, since that string can be spoofed by anyone, a standard part of any real technical audit.

**How often is this list actually updated?**
**Direct answer:** As new crawlers become relevant, this is maintained as a living reference rather than a one-time publish, the same discipline applied to [the AI Overview update tracker](/blog/ai-overview-algorithm-update-tracker).

**Is a training crawler worth blocking if I don't want my content used for AI training?**
**Direct answer:** That's a legitimate, separate decision from real-time retrieval access, blocking a training-only crawler doesn't affect whether your content can still be cited in a live AI-search answer, the two are controlled independently.

---

*I keep this list current as part of every technical audit's crawler-verification step, it's a living reference, not a one-time write-up. [Check your own site's current configuration against it](/chat).*

---
_Canonical page: [https://iliassami.com/blog/complete-ai-crawler-list](https://iliassami.com/blog/complete-ai-crawler-list) · Markdown generated on request from the live site content._
