CCBot

Common Crawl
Dataset crawler

About

CCBot builds the Common Crawl open web archive — a free, public dataset of web crawl data. Many AI labs train models on Common Crawl, so allowing or blocking CCBot indirectly affects a large share of model training pipelines.

Common Crawl, feeds many models

Details

Type
Dataset crawler
Operator
Common Crawl
Access
Direct — crawls on the operator’s own schedule
Traffic rank