Crawlers
A reference of the 65 crawlers tracked by tinytrack.
Search engine crawlers
Googlebot is the search engine crawler for Google Search.
The Google Images bot is the search engine crawler for Google Images Search.
The Google Videos bot is the search engine crawler for Google Video Search.
Generic crawler that may be used by various product teams for fetching publicly accessible content from sites. For example, it may be used for one-off crawls for internal research and development. https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers#googleother
APIs-Google is the user agent used by Google APIs to deliver push notification messages. Application developers can request these notifications to avoid the need for continually polling Google's servers to find out if the resources they are interested in have changed. To make sure nobody abuses this service, Google requires developers to prove that they own the domain before allowing them to register a URL with a domain as the location where they want to receive messages.
The Google StoreBot is a search-engine-based program that automatically 'crawls' through web pages to gather and analyse data. Google uses crawlers that go through product pages and checkout processes using machine learning algorithms to fill in forms with information such as delivery addresses, and help compile other information on price, delivery, payments and more.
Bingbot crawler and handles most of Bing's crawling needs each day.
DuckDuckBot is the search engine crawler for the DuckDuckGo search engine.
The main indexing robot for Yandex search.
Baiduspider is the search engine crawler for the search engine Baidu.
YisouSpider crawls the web for Yisou and Shenma, the Alibaba-backed Chinese mobile search engines. It is one of the highest-volume crawlers originating from China and a frequent target of rate limits.
Applebot data is used to power various features, such as the search technology that is integrated into many user experiences in Appleʼs ecosystem including Spotlight, Siri, and Safari.
Yeti is the web crawler for Naver, a South Korean search engine. It indexes websites to provide search results and power other services on the Naver platform.
Yahoo! Slurp was the search engine crawler for Yahoo's search engine.
Amazon Kendra is a highly accurate intelligent search service that enables your users to search unstructured data using natural language. It returns specific answers to questions, giving users an experience that's close to interacting with a human expert. It is highly scalable and capable of meeting performance demands, tightly integrated with other AWS services such as Amazon S3 and Amazon Lex, and offers enterprise-grade security.
AmazonProductDiscoveryBot crawls the web to discover products and product information, improving what Amazon can surface in its shopping experiences.
AI indexing and training crawlers
GPTBot is used to crawl content that may be used in training OpenAI's generative AI foundation models
OAI-SearchBot is used to link to and surface websites in search results in the SearchGPT prototype
OAI-AdsBot is used to validate the safety of web pages submitted as ads on ChatGPT.
ClaudeBot helps enhance the utility and safety of generative AI models by collecting web content that could potentially contribute to their training.
Claude-SearchBot navigates the web to improve search result quality for Claude users.
PerplexityBot builds the index behind Perplexity's answer engine. Pages it crawls can be cited as sources in Perplexity answers.
Google-Extended is not a separate crawler: it is a robots.txt control token. Disallowing it tells Google not to use your content for Gemini training and grounding, while normal Google crawling and indexing continue via Googlebot.
Applebot-Extended is a robots.txt control token rather than a crawler. Disallowing it opts your content out of Apple Intelligence foundation-model training; the actual fetching is still done by Applebot.
Use cases such as training AI models or improving products by indexing content directly.
The Meta-WebIndexer crawler navigates the web to improve Meta AI search result quality for users. In doing so, Meta analyzes online content to enhance the relevance and accuracy of Meta AI.
Amazonbot is Amazon's web crawler used to improve our services, such as enabling Alexa to answer even more questions for customers. Amazonbot is a polite crawler that respects standard robots.txt rules and robots meta tags.
The Amazon Q Business web crawler indexes web content that enterprises connect to Amazon Q, so the assistant can answer questions grounded in that content.
Bytespider is ByteDance's web crawler, collecting training data for its AI models including Doubao. It is known for aggressive crawl volumes and has been reported to ignore robots.txt at times.
TikTokSpider is a ByteDance crawler used for TikTok-related and internal content discovery, complementing Bytespider's model-training crawls.
CCBot builds the Common Crawl open web archive — a free, public dataset of web crawl data. Many AI labs train models on Common Crawl, so allowing or blocking CCBot indirectly affects a large share of model training pipelines.
AI2Bot crawls web content for the Allen Institute for AI to build open training datasets such as Dolma, which back its open language models.
ShapBot helps discover and index websites for Parallel's web APIs.
The assistants that fetch pages on demand.
ChatGPT-User is for user actions in ChatGPT and Custom GPTs. When users ask ChatGPT or a CustomGPT a question, it may visit a web page to help answer and include a link to the source in its response.
Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent.
Claude Code fetches pages when Anthropic's coding agent needs documentation or web context during a programming task. Fetches are user-initiated rather than bulk crawling.
Perplexity-User fetches a page on demand when a Perplexity user asks about it or clicks a citation. Because fetches are user-initiated, Perplexity states this agent generally ignores robots.txt rules.
Bot for user actions in le Chat by Mistral AI, for instance when asked to open a web page.
DuckAssistBot is a web crawler for DuckDuckGo
You.com Search Engine Crawler
ExaBot crawls the web to build Exa's neural search index, which is used by AI agents, developers and research tools to retrieve pages via semantic search.
FirecrawlAgent fetches pages for Firecrawl, an API that converts websites into LLM-ready data. Requests are triggered by Firecrawl customers scraping or crawling specific sites.
Crawl websites and extract content to feed AI apps. Convert web data to Markdown or HTML, download files, and more.
Crawler available to site owners to request crawls of their own sites for targeted AI training
Google-Agent is used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request (e.g., Project Mariner).
Gemini Deep Research fetches pages while compiling a research report a Gemini user asked for. It is a user-triggered fetcher rather than a bulk crawler.
Google NotebookLM fetches the specific web pages a user adds as sources to a NotebookLM notebook, so the assistant can ground its answers in them.
Google Read Aloud service enables reading web pages using text-to-speech (TTS).
Google FeedFetcher is the RSS reader for Google.
Crawler receives individual links at the user's initiative to support certain product features.
Verification is the process of proving that you own the property that you claim to own. Search Console needs to verify ownership because verified owners have access to sensitive Google Search data for a site, and can affect a site's presence and behavior on Google Search and other Google properties. A verified owner can grant full or view access to other people.
Fetches and processes feeds that publishers explicitly supplied for use in Google News landing pages.
The tools that audit and analyze.
AhrefsBot is a Web Crawler that powers the 12 trillion link database for Ahrefs online marketing toolset. It constantly crawls web to fill our database with new links and check the status of the previously found ones to provide the most comprehensive and up-to-the-minute data to our users.
AhrefsSiteAudit is used by website owners (paid and free) to look for issues on their websites
Semrushbot crawls your website to analyze it for different SEO and technical issues.
Check for over 130 common website issues and get special reports about your site’s crawlability, use of markups, internal linking, speed/performance, HTTPS, and international SEO.
SplitSignalBot crawls pages that take part in Semrush SplitSignal SEO A/B tests, verifying test variants as search engines would see them.
RyteBot crawls websites for Ryte's quality-assurance and SEO platform, which integrates with Semrush, checking pages for technical and content issues.
DataForSEO Bot is a driving force of our leading product - Backlinks API, which has been developed with a single purpose: providing website owners, webmasters, and SEO professionals with opportunities to analyze the key component of website optimization – backlink analytics. You can learn more about the DataForSEO Bot on this dedicated page: https://dataforseo.com/dataforseo-bot
Dotbot is Moz's web crawler, it gathers web data for the Moz Link Index. This data we collect through Dotbot is available in the Links section of your Moz Pro campaign, Link Explorer, and the Moz Links API.
SEO web crawler to identify web page popularity and thematic
ClarityBot fetches pages for Microsoft Clarity, the free analytics tool, to render the page snapshots behind its heatmaps and session recordings.
Google-InspectionTool is the crawler used by Search testing tools such as the Rich Result Test and URL inspection in Search Console. Apart from the user agent and user agent token, it mimics Googlebot.
Similarweb's crawler gathers page and site data that feeds its traffic estimates and competitive-intelligence reports.
The Screaming Frog SEO Spider is a desktop crawler run by SEO practitioners auditing a site. Traffic comes from individual users' machines rather than a central service, usually in short concentrated bursts.