Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I spent years building a 103B-token Usenet corpus (1980–2013) and finally documented it [P]

Via r/MachineLearning
Friday, May 1, 2026 · 6:01PM
Summary

For the past several years I've been quietly assembling and processing what I believe is one of the larger privately held pretraining corpora around... a complete Usenet archive spanning 1980 to 2013. Here's what it ended up being: 103.1 billion tokens (cl100k_base) 408 million posts across 9 newsgr

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories