Reference library

Dataset

Open datasets you can query yourself, from crawl corpora to real-user performance data.

7 references · 1 canonical · 6 free · Advanced / Intermediate

TitleAuthor / publisherYearLevelCost
Common Crawl ↗Canon
Free, openly licensed petabyte-scale archive of web crawl data with monthly releases of WARC, WAT and WET files plus columnar indexes. Heavily used for link graph research and LLM pretraining.
commoncrawl.org
Common Crawl Foundation, AdvancedFree
ai.robots.txt Community Blocklist ↗
Community-maintained open-source list of AI-related crawlers with ready-made robots.txt, .htaccess, nginx, Caddy, HAProxy and JSON outputs for blocking or auditing them.
github.com
ai-robots-txt project, IntermediateFree
Chrome UX Report on BigQuery ↗
The public BigQuery dataset containing monthly Core Web Vitals distributions for millions of origins, queryable for competitive performance benchmarking at scale.
console.cloud.google.com
Google Chrome team, AdvancedFreemium
ClueWeb22 Dataset ↗
Large research web corpus derived from commercial search engine crawls, distributed under licence for academic information retrieval and web search research.
lemurproject.org
Lemur Project / Carnegie Mellon University, AdvancedFree
HTTP Archive Reports ↗
Longitudinal trend reports on how the web is built and performs, backed by a public BigQuery dataset of monthly crawls of millions of URLs.
httparchive.org
HTTP Archive, AdvancedFree
MS MARCO ↗
Large-scale dataset of real anonymised Bing queries with human-generated answers and passage relevance labels. The standard benchmark for passage ranking and dense retrieval research.
microsoft.github.io
Microsoft, AdvancedFree
Wikimedia Downloads (Wikipedia Dumps) ↗
Complete periodic exports of Wikipedia and sister projects, including page content, pagelinks and pageview data. A standard corpus for entity, knowledge graph and query understanding work.
dumps.wikimedia.org
Wikimedia Foundation, AdvancedFree

Other formats