Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

What a time to be alive from 1tk/sec to 20-100tk/sec for huge models

Via r/LocalLlama
Sunday, May 3, 2026 · 5:46PM
Summary

https://www.reddit.com/r/LocalLLaMA/comments/1eb6to7/llama_405b_q4_k_m_quantization_running_locally/ https://www.reddit.com/r/LocalLLaMA/comments/1ebbgkr/llama_31_405b_q5_k_m_running_on_amd_epyc_9374f/ Llama405b q4 at 1.2tk/sec 2 years ago was something to be excited about. That same hardware will n

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories