Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

100 Trillion+ Pretraining data??? This is the largest data I've see a model being trained on.

Via r/LocalLlama
Monday, Jun 1, 2026 · 4:38AM
Summary

https://preview.redd.it/oss7g2gnll4h1.png?width=894&format=png&auto=webp&s=5d4295707a700ed7541c274b8be8ad75bbd0903d Edit: This is about Minimax-M3, I just realised I didn't mention it lol Usually we see 27-50 Trillion tokens in most models, kimi, mimo, deepseek. They seem to have doubled the pretrai

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories