Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Training an LLM from scratch on 1800's texts (160GB dataset)

Via r/LocalLlama
Friday, Jul 10, 2026 · 6:51PM
Summary

Hi everyone, A year ago I began pre-training language models exclusively on 1800’s London data. Recently I have completed my largest dataset ever, containing 40B tokens or 160GB of 1800-1875 english data from England and the United States. I will soon train a 2B parameter model on it, but for now I’

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories