The full list of AI crawler signatures we match.
We match 60+ answer-engine crawlers. This page is the source of truth. If your logs show one we miss, tell us — we’ll add it within 7 days and credit you.
68 signatures currently active. First-seen dates are not tracked per signature, so the table shows the date each entry was last verified against live traffic.
| Crawler | Signature type | Example payload | What it does | Last verified | Changelog |
|---|---|---|---|---|---|
| chatgpt | user-agent token | GPTBot | OpenAI training crawler | 2026-09-03 | entry |
| chatgpt | user-agent token | OAI-SearchBot | OpenAI search index crawler | 2026-09-03 | entry |
| chatgpt | user-agent token | ChatGPT-User | ChatGPT live browsing on user request | 2026-09-03 | entry |
| chatgpt | user-agent token | OpenAI-Image | OpenAI image fetcher | 2026-09-03 | entry |
| chatgpt | user-agent token | openai.com | Generic OpenAI agent string | 2026-09-03 | entry |
| claude | user-agent token | ClaudeBot | Anthropic crawler | 2026-09-03 | entry |
| claude | user-agent token | Claude-Web | Anthropic web fetcher | 2026-09-03 | entry |
| claude | user-agent token | anthropic-ai | Anthropic legacy agent | 2026-09-03 | entry |
| claude | user-agent token | Claude-User | Claude live browsing | 2026-09-03 | entry |
| claude | user-agent token | Claude-SearchBot | Claude search index crawler | 2026-09-03 | entry |
| copilot | user-agent token | MicrosoftPreview | Microsoft Copilot preview fetcher | 2026-09-03 | entry |
| copilot | user-agent token | msnbot-media | Microsoft media fetcher | 2026-09-03 | entry |
| copilot | user-agent token | BingPreview | Bing preview fetcher | 2026-09-03 | entry |
| copilot | user-agent token | BingBot-AI | Bing AI answer crawler | 2026-09-03 | entry |
| copilot | user-agent token | Copilot | Microsoft Copilot agent | 2026-09-03 | entry |
| gemini | user-agent token | Google-CloudVertex-Bot | Vertex AI crawler | 2026-09-03 | entry |
| gemini | user-agent token | GoogleOther | Google experimental/AI crawler | 2026-09-03 | entry |
| gemini | user-agent token | GoogleAgent-Mariner | Google agentic browsing | 2026-09-03 | entry |
| gemini | user-agent token | Google-Firebase | Google AI product fetcher | 2026-09-03 | entry |
| gemini | user-agent token | Google-Extended | Google Gemini training opt-in token | 2026-09-03 | entry |
| gemini | user-agent token | Google-CloudVertexBot | Vertex AI grounding fetcher | 2026-09-03 | entry |
| other | user-agent token | NeevaBot | Neeva legacy AI crawler | 2026-09-03 | entry |
| other | user-agent token | Andibot | Andi search engine | 2026-09-03 | entry |
| other | user-agent token | Phindbot | Phind answer engine | 2026-09-03 | entry |
| other | user-agent token | Poe | Quora Poe agent | 2026-09-03 | entry |
| other | user-agent token | Applebot-Extended | Apple Intelligence training crawler | 2026-09-03 | entry |
| other | user-agent token | Applebot | Apple search/AI crawler | 2026-09-03 | entry |
| other | user-agent token | cohere-ai | Cohere crawler | 2026-09-03 | entry |
| other | user-agent token | cohere-training-data-crawler | Cohere training crawler | 2026-09-03 | entry |
| other | user-agent token | Meta-ExternalAgent | Meta AI crawler | 2026-09-03 | entry |
| other | user-agent token | Meta-ExternalFetcher | Meta AI live fetcher | 2026-09-03 | entry |
| other | user-agent token | FacebookBot | Meta training crawler | 2026-09-03 | entry |
| other | user-agent token | Amazonbot | Amazon Alexa/AI crawler | 2026-09-03 | entry |
| other | user-agent token | Bytespider | ByteDance/Doubao crawler | 2026-09-03 | entry |
| other | user-agent token | TikTokSpider | ByteDance crawler | 2026-09-03 | entry |
| other | user-agent token | CCBot | Common Crawl (LLM training corpus) | 2026-09-03 | entry |
| other | user-agent token | ai2bot | Allen Institute crawler | 2026-09-03 | entry |
| other | user-agent token | AI2Bot-Dolma | Allen Institute Dolma corpus | 2026-09-03 | entry |
| other | user-agent token | YouBot | You.com answer engine | 2026-09-03 | entry |
| other | user-agent token | Diffbot | Diffbot knowledge graph | 2026-09-03 | entry |
| other | user-agent token | ImagesiftBot | Imagesift/Hive AI crawler | 2026-09-03 | entry |
| other | user-agent token | PetalBot | Huawei Petal AI search | 2026-09-03 | entry |
| other | user-agent token | DataForSeoBot | DataForSEO crawler | 2026-09-03 | entry |
| other | user-agent token | SemrushBot-OCOB | Semrush AI content crawler | 2026-09-03 | entry |
| other | user-agent token | Timpibot | Timpi decentralized index | 2026-09-03 | entry |
| other | user-agent token | omgili | Webz.io LLM data crawler | 2026-09-03 | entry |
| other | user-agent token | omgilibot | Webz.io crawler | 2026-09-03 | entry |
| other | user-agent token | Webzio-Extended | Webz.io AI dataset crawler | 2026-09-03 | entry |
| other | user-agent token | ICC-Crawler | NICT AI research crawler | 2026-09-03 | entry |
| other | user-agent token | Kangaroo Bot | Kangaroo LLM crawler | 2026-09-03 | entry |
| other | user-agent token | img2dataset | LAION dataset builder | 2026-09-03 | entry |
| other | user-agent token | Scrapy | Generic scraping framework | 2026-09-03 | entry |
| other | user-agent token | python-requests | Generic scripted fetch | 2026-09-03 | entry |
| other | user-agent token | node-fetch | Generic scripted fetch | 2026-09-03 | entry |
| other | user-agent token | Go-http-client | Generic scripted fetch | 2026-09-03 | entry |
| other | user-agent token | MistralAI-User | Mistral Le Chat browsing | 2026-09-03 | entry |
| other | user-agent token | Mistral-Crawler | Mistral crawler | 2026-09-03 | entry |
| other | user-agent token | xAI-Bot | xAI Grok crawler | 2026-09-03 | entry |
| other | user-agent token | GrokBot | xAI Grok agent | 2026-09-03 | entry |
| other | user-agent token | DeepSeekBot | DeepSeek crawler | 2026-09-03 | entry |
| other | user-agent token | QwenBot | Alibaba Qwen crawler | 2026-09-03 | entry |
| other | user-agent token | SentiBot | Senti AI crawler | 2026-09-03 | entry |
| other | user-agent token | Firecrawl | Firecrawl agent scraper | 2026-09-03 | entry |
| other | user-agent token | Brightbot | Bright Data AI crawler | 2026-09-03 | entry |
| other | user-agent token | LinerBot | Liner answer engine | 2026-09-03 | entry |
| perplexity | user-agent token | Perplexity-ai | Perplexity legacy agent | 2026-09-03 | entry |
| perplexity | user-agent token | Perplexity-User | Perplexity live user fetch | 2026-09-03 | entry |
| perplexity | user-agent token | PerplexityBot | Perplexity index crawler | 2026-09-03 | entry |
How we classify
Classification runs server-side on every request, before any JavaScript executes, and in three passes. The first pass looks for an exact documented user-agent token such as the ones OpenAI, Anthropic, Perplexity and Google publish. A token match is the strongest signal available and is used directly. The second pass applies a regular expression for known variants — version suffixes, vendor-specific fetcher strings, and agents that embed a product name inside a generic library string.
The third pass handles requests whose agent string is weak, generic or absent by checking the source address against published crawler IP ranges. That pass also runs as a cross-check on token matches where the vendor publishes ranges, because a request claiming a crawler identity from outside the vendor’s network is an impersonator and is reported as such. Nothing in this pipeline reads or writes anything on the visitor’s device, which is why it works without cookies or a consent banner.
Frequently asked questions
How do you decide whether a request is an AI crawler?
We match the user-agent against a documented token first, then against a regex for variants, and finally against published IP ranges when the agent string is weak or absent.
What if a crawler is missing from this list?
Send us the raw log line. We add verified signatures within seven days and credit the reporter in the changelog entry.
Do you block these crawlers?
No. Citealytics only classifies and reports them. Whether to allow a crawler is a robots.txt decision you make on your own site.
Can someone fake a crawler user-agent?
Yes, which is why token matches are cross-checked against published IP ranges where the vendor provides them. A token match from outside those ranges is treated as an impersonator.
Measure the AI traffic your analytics is hiding.
One line of script, no cookies, no consent banner. Citealytics classifies every visit against a registry of answer-engine signatures and shows which pages get cited.
Start free — 10k events/mo →