Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I trained a 75M parameter LLM from scratch on 18B tokens and it beats a model almost double its size

Via r/LocalLlama
Tuesday, Jun 2, 2026 · 5:41PM
Summary

I trained a small language model from scratch called KeyLM. It is 75M params, decoder-only, and there is a pretrained base, an instruction-tuned version, and a GGUF. On IFEval (instruction following) the 75M instruct model scores slightly higher than the original SmolLM-135M-Instruct at about half t

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories