Crawlers

A reference of the 65 crawlers tracked by tinytrack.

AI indexing and training crawlers

GPTBot
OpenAI

GPTBot is used to crawl content that may be used in training OpenAI's generative AI foundation models

AI crawler
OAI-SearchBot
OpenAI

OAI-SearchBot is used to link to and surface websites in search results in the SearchGPT prototype

Search engine crawler
OAI-AdsBot
OpenAI

OAI-AdsBot is used to validate the safety of web pages submitted as ads on ChatGPT.

Advertising
ClaudeBot
Anthropic

ClaudeBot helps enhance the utility and safety of generative AI models by collecting web content that could potentially contribute to their training.

AI crawler
Claude-SearchBot
Anthropic

Claude-SearchBot navigates the web to improve search result quality for Claude users.

AI search
PerplexityBot
Perplexity

PerplexityBot builds the index behind Perplexity's answer engine. Pages it crawls can be cited as sources in Perplexity answers.

AI search
Google-Extended
Google

Google-Extended is not a separate crawler: it is a robots.txt control token. Disallowing it tells Google not to use your content for Gemini training and grounding, while normal Google crawling and indexing continue via Googlebot.

AI training control
Applebot-Extended
Apple

Applebot-Extended is a robots.txt control token rather than a crawler. Disallowing it opts your content out of Apple Intelligence foundation-model training; the actual fetching is still done by Applebot.

AI training control
Meta-ExternalAgent
Meta

Use cases such as training AI models or improving products by indexing content directly.

AI crawler
Meta-WebIndexer
Meta

The Meta-WebIndexer crawler navigates the web to improve Meta AI search result quality for users. In doing so, Meta analyzes online content to enhance the relevance and accuracy of Meta AI.

Page preview
Amazonbot
Amazon

Amazonbot is Amazon's web crawler used to improve our services, such as enabling Alexa to answer even more questions for customers. Amazonbot is a polite crawler that respects standard robots.txt rules and robots meta tags.

AI crawler
Amazon Q Business
Amazon

The Amazon Q Business web crawler indexes web content that enterprises connect to Amazon Q, so the assistant can answer questions grounded in that content.

AI crawler
Bytespider
ByteDance

Bytespider is ByteDance's web crawler, collecting training data for its AI models including Doubao. It is known for aggressive crawl volumes and has been reported to ignore robots.txt at times.

AI crawler
TikTokSpider
ByteDance

TikTokSpider is a ByteDance crawler used for TikTok-related and internal content discovery, complementing Bytespider's model-training crawls.

AI crawler
CCBot
Common Crawl

CCBot builds the Common Crawl open web archive — a free, public dataset of web crawl data. Many AI labs train models on Common Crawl, so allowing or blocking CCBot indirectly affects a large share of model training pipelines.

Dataset crawler
AI2Bot
Allen Institute for AI

AI2Bot crawls web content for the Allen Institute for AI to build open training datasets such as Dolma, which back its open language models.

AI crawler
ShapBot
Parallel

ShapBot helps discover and index websites for Parallel's web APIs.

AI search

The assistants that fetch pages on demand.

ChatGPT-User
OpenAI

ChatGPT-User is for user actions in ChatGPT and Custom GPTs. When users ask ChatGPT or a CustomGPT a question, it may visit a web page to help answer and include a link to the source in its response.

AI assistantUser-triggered
Claude-User
Anthropic

Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent.

AI assistantUser-triggered
Claude Code
Anthropic

Claude Code fetches pages when Anthropic's coding agent needs documentation or web context during a programming task. Fetches are user-initiated rather than bulk crawling.

AI assistantUser-triggered
Perplexity-User
Perplexity

Perplexity-User fetches a page on demand when a Perplexity user asks about it or clicks a citation. Because fetches are user-initiated, Perplexity states this agent generally ignores robots.txt rules.

AI assistantUser-triggered
MistralAI-User
Mistral AI

Bot for user actions in le Chat by Mistral AI, for instance when asked to open a web page.

AI assistantUser-triggered
DuckAssistBot
DuckDuckGo

DuckAssistBot is a web crawler for DuckDuckGo

AI assistantUser-triggered
YouBot
You.com

You.com Search Engine Crawler

Search engine crawlerUser-triggered
ExaBot
Exa

ExaBot crawls the web to build Exa's neural search index, which is used by AI agents, developers and research tools to retrieve pages via semantic search.

AI searchUser-triggered
FirecrawlAgent
Firecrawl

FirecrawlAgent fetches pages for Firecrawl, an API that converts websites into LLM-ready data. Requests are triggered by Firecrawl customers scraping or crawling specific sites.

AI assistantUser-triggered
ApifyWebsiteContentCrawler
Apify

Crawl websites and extract content to feed AI apps. Convert web data to Markdown or HTML, download files, and more.

AI assistantUser-triggered
Google-CloudVertexBot
Google

Crawler available to site owners to request crawls of their own sites for targeted AI training

AI crawlerUser-triggered
Google-Agent
Google

Google-Agent is used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request (e.g., Project Mariner).

AI assistantUser-triggered
Gemini Deep Research
Google

Gemini Deep Research fetches pages while compiling a research report a Gemini user asked for. It is a user-triggered fetcher rather than a bulk crawler.

AI assistantUser-triggered
Google NotebookLM
Google

Google NotebookLM fetches the specific web pages a user adds as sources to a NotebookLM notebook, so the assistant can ground its answers in them.

AI assistantUser-triggered
Google Read Aloud
Google

Google Read Aloud service enables reading web pages using text-to-speech (TTS).

AccessibilityUser-triggered
Google Feedfetcher
Google

Google FeedFetcher is the RSS reader for Google.

Feed fetcherUser-triggered
Meta-ExternalFetcher
Meta

Crawler receives individual links at the user's initiative to support certain product features.

AI assistantUser-triggered
Google Site Verifier
Google

Verification is the process of proving that you own the property that you claim to own. Search Console needs to verify ownership because verified owners have access to sensitive Google Search data for a site, and can affect a site's presence and behavior on Google Search and other Google properties. A verified owner can grant full or view access to other people.

SecurityUser-triggered
Google Publisher Center
Google

Fetches and processes feeds that publishers explicitly supplied for use in Google News landing pages.

Feed fetcherUser-triggered

The tools that audit and analyze.

AhrefsBot
Ahrefs

AhrefsBot is a Web Crawler that powers the 12 trillion link database for Ahrefs online marketing toolset. It constantly crawls web to fill our database with new links and check the status of the previously found ones to provide the most comprehensive and up-to-the-minute data to our users.

SEO
AhrefsSiteAudit
Ahrefs

AhrefsSiteAudit is used by website owners (paid and free) to look for issues on their websites

SEO
SemrushBot
Semrush

Semrushbot crawls your website to analyze it for different SEO and technical issues.

SEO
SiteAuditBot
Semrush

Check for over 130 common website issues and get special reports about your site’s crawlability, use of markups, internal linking, speed/performance, HTTPS, and international SEO.

SEO
SplitSignalBot
Semrush

SplitSignalBot crawls pages that take part in Semrush SplitSignal SEO A/B tests, verifying test variants as search engines would see them.

SEO
RyteBot
Ryte

RyteBot crawls websites for Ryte's quality-assurance and SEO platform, which integrates with Semrush, checking pages for technical and content issues.

SEO
DataForSeoBot
DataForSEO

DataForSEO Bot is a driving force of our leading product - Backlinks API, which has been developed with a single purpose: providing website owners, webmasters, and SEO professionals with opportunities to analyze the key component of website optimization – backlink analytics. You can learn more about the DataForSEO Bot on this dedicated page: https://dataforseo.com/dataforseo-bot

SEO
DotBot
Moz

Dotbot is Moz's web crawler, it gathers web data for the Moz Link Index. This data we collect through Dotbot is available in the Links section of your Moz Pro campaign, Link Explorer, and the Moz Links API.

SEO
Barkrowler
Babbar

SEO web crawler to identify web page popularity and thematic

SEO
ClarityBot
Microsoft

ClarityBot fetches pages for Microsoft Clarity, the free analytics tool, to render the page snapshots behind its heatmaps and session recordings.

Analytics
Google Inspection Tool
Google

Google-InspectionTool is the crawler used by Search testing tools such as the Rich Result Test and URL inspection in Search Console. Apart from the user agent and user agent token, it mimics Googlebot.

Security
Similarweb
Similarweb

Similarweb's crawler gathers page and site data that feeds its traffic estimates and competitive-intelligence reports.

Analytics
Screaming Frog
Screaming Frog

The Screaming Frog SEO Spider is a desktop crawler run by SEO practitioners auditing a site. Traffic comes from individual users' machines rather than a central service, usually in short concentrated bursts.

SEO