Non-profit
Common CrawlMaintains a free, petabyte-scale archive of the open web used in thousands of research projects and as a major source of language-model training data. Its public crawls, text extracts, indexes, and web graphs make web-scale analysis broadly accessible.