All organizations

Non-profit

Common Crawl

Founded2007

Maintains a free, petabyte-scale archive of the open web used in thousands of research projects and as a major source of language-model training data. Its public crawls, text extracts, indexes, and web graphs make web-scale analysis broadly accessible.

See something inaccurate or outdated?