Overview
News
Technologies
Salaries
Products
People
Growth
Financials
Overview
We build and maintain an open repository of web crawl data that can be accessed and analyzed by anyone.
News
August 2026 Crawl Archive Now Available
We are pleased to announce that the crawl archive for August 2026 is now available, containing 2.14 billion web pages or 360 TiB of uncompressed content.
Read more
Report
Common Crawl Foundation at ACL 2026
The Common Crawl team attended the 64th Annual Meeting of the Association for Computational Linguistics in San Diego, California, presenting recent published work, and strengthening ties with the research community.
Read more
Report
Announcing the First Stable Release of CC-Downloader
Over a year ago we released cc-downloader, an experimental tool to politely download Common Crawl data. Today we're releasing its first stable version, with a Rust library and Python bindings.
Read more
Report
Notes from HTTP Workshop Basel and IETF 126 Vienna
Two weeks in Basel and Vienna, at the HTTP Workshop and IETF 126. Protocol adoption measured across the whole web, and an attempt to define what "machine readable" actually means.
Read more
Report
July 2026 Crawl Archive Now Available
The crawl archive for July 2026 is now available. The data was crawled between July 7th and July 25th, and contains 2.14 billion web pages (or 364.01 TiB of uncompressed content). We also announce some improvements and changes.
Read more
Report
Pro access
Upgrade to see all 37 mentions
Upgrade to a paid plan to read every media mention of this company - funding news, awards, product launches and press releases from all the outlets writing about it.
Every media mention and press release
Funding news, awards and product launches
Fresh coverage from every outlet writing about the company
Upgrade now
Cancel anytime. Secure checkout. Instant activation.
