AI Crawler & Bot Field Guide
Documented signatures, behaviors, robots.txt directives, and verification rules for the modern automated web.
GPTBot
OpenAILarge-scale ingestion for model pre-training. Scans robots.txt and sitemaps.
OAI-SearchBot
OpenAILive web indexing for ChatGPT Search features. Honors robots.txt separately from training.
ChatGPT-User
OpenAIFetched on-demand when a ChatGPT user provides a URL or search query.
ClaudeBot
AnthropicLarge-scale crawler for Claude model training. Follows standard links & robots.txt.
PerplexityBot
Perplexity AIPerplexity search crawler for indexing web content and summarizing citations.
Perplexity-User
Perplexity AILive on-demand web fetches triggered directly by user prompts in Perplexity.
Googlebot
GooglePrimary Google search indexer. Renders HTML and executes JS when queue permits.
Google-Extended
GoogleUsed to manage Gemini and Vertex AI training access without affecting Google Search.
Bingbot
MicrosoftMicrosoft search indexer powering Bing and Microsoft Copilot index data.
Bytespider
ByteDanceHigh-frequency crawler used for TikTok search and Doubao LLM ingestion.
CCBot
Common CrawlMonthly open web archive used by hundreds of open source AI models (Llama, etc.).
Meta-ExternalAgent
MetaMeta crawler for AI model development and link previews.
Amazonbot
AmazonAmazon web crawler for Alexa answer generation and search features.
Applebot
AppleApple search indexer and Apple Intelligence web dataset ingestion.
Cohere-AI
CohereCrawler gathering web text for Cohere Command and Embed models.
Automated Scraper / Script
Autonomous / DeveloperCustom script, automated library, or headless browser session.