AI crawlers list: GPTBot, ClaudeBot and more explained
GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot and more. What each AI crawler does and how Orchly spots fake bots.
Every bot on the AI crawlers list visits your site for a different reason. GPTBot collects training data, OAI-SearchBot indexes pages for ChatGPT search, and ChatGPT-User fetches a page live because someone just asked ChatGPT a question. Orchly sorts every AI crawler into one of five bot types by what it came to do, and checks whether each bot is really who it says it is.
You’ll see these types in the Bot Types chart, the Pages tab columns and the Bot Type column of the Traffic log, all under AI Traffic.
The five AI bot types
Each bot’s category comes from its own company’s documentation, not from which company owns it. That’s why three OpenAI bots land in three different groups.
| Bot type | What it means | Why it matters to you |
|---|---|---|
| AI Citation | An AI fetching your page live to answer someone’s question | The visit behind a citation. The closest thing to AI recommending you |
| AI Agent | An AI acting for a person, such as browsing, shopping or coding | Agents may buy, book or sign up on a user’s behalf |
| AI Indexing | An AI search engine adding your page to the index it answers from | You need to be in the index before you can be cited |
| AI Training | An AI company’s crawler collecting content to train its models | Shapes what models know about you long term |
| Search Indexing | A classic search engine like Google, Bing or Apple indexing your page | Your normal SEO crawl |
Two more labels show up alongside them:
- Human Visit: a real person who landed on your site by clicking a link inside an AI answer.
- Other: a bot that isn’t AI or search, like a social media link preview.
Google and Apple may also use their search crawls for their AI features unless your robots.txt opts out. So Googlebot counts as Search Indexing, but what it fetches can still end up in Gemini or Google AI Overviews.
AI crawlers list: which bots belong to each type?
These are the bots Orchly recognizes today, grouped by type.
| Bot type | Bots |
|---|---|
| AI Citation | ChatGPT-User, Claude-User, Claude-Web, Perplexity-User, Gemini-Deep-Research, Google-NotebookLM, DuckAssistBot, Amazon-User, Meta-ExternalFetcher, MistralAI-User |
| AI Agent | OAI-Operator, Codex, Claude Code, Google-Agent, GoogleAgent-Mariner, AmazonBuyForMe |
| AI Indexing | OAI-SearchBot, Claude-SearchBot, PerplexityBot, Google-CloudVertexBot, Amazon-SearchBot, Meta-WebIndexer, GrokBot, YouBot |
| AI Training | GPTBot, ClaudeBot, Anthropic-AI, Google-Extended, Amazonbot, Meta-ExternalAgent, xAI-Bot, Bytespider, CCBot, Cohere-AI, MistralBot, AI21Bot, Together-Bot, Webzio-Extended |
| Search Indexing | Googlebot (plus Image, Video and News), GoogleOther, Storebot-Google, Bingbot, Applebot, DuckDuckBot, YandexBot, Baiduspider, PetalBot, Yahoo Slurp |
| Other | MicrosoftPreview, FacebookBot, Twitterbot, TikTokSpider |
When an AI assistant pulls a page just before its user opens it (a prefetch), Orchly counts it as an AI Citation too.
GPTBot vs ChatGPT-User vs OAI-SearchBot
This is the question people ask most. In short:
- GPTBot crawls for training. Blocking it keeps your content out of future OpenAI model training.
- OAI-SearchBot builds the index ChatGPT search answers from. Block it and you can drop out of ChatGPT search results.
- ChatGPT-User fetches a page live when a user’s question needs it. These are the visits behind ChatGPT citations.
The same pattern holds for Claude (ClaudeBot, Claude-SearchBot, Claude-User) and Perplexity (PerplexityBot, Perplexity-User). If you want to show up in AI answers but don’t want to be trained on, block the training crawler and leave the other two alone.
How does Orchly verify AI bots?
Anyone can send a request that says “GPTBot” in its user agent, and scrapers do it all the time. So each bot visit is checked the way the bot’s company documents, while Orchly still has the IP address in hand:
- Published IP ranges. OpenAI, Perplexity, Google, Bing, Apple, DuckDuckGo and Common Crawl publish the addresses their bots use. Orchly checks the request against them.
- Reverse DNS. For companies that document it, like Google, Bing, Apple, Amazon and Yandex, the IP’s hostname has to sit under the company’s domain and resolve back to the same IP.
Each visit then gets one of three results in the Traffic log’s Verified column:
| Result | Meaning | Counted? |
|---|---|---|
| Verified | The request came from an address the company publishes for this bot | Yes |
| Fake | The bot’s name came from an address the company doesn’t own | No, left out of every chart and count |
| Unchecked | The company publishes no way to check, or the check couldn’t finish | Yes, counted as normal |
Unchecked is never treated as fake. Anthropic and Meta, for example, publish no way to verify their bots, so ClaudeBot visits usually show as Unchecked.
Why fake bots matter
If you count spoofed GPTBot visits as real, you’ll think OpenAI crawls you far more than it does. Worse, you might block or allow the wrong thing in robots.txt based on traffic that was never OpenAI.
Orchly leaves fakes out of every chart and count, and the Bot Types card says how many it excluded. The individual rows stay in the Traffic log, tagged Fake, so you can see which paths scrapers go after.