Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

1-bit llms on device?!

Via r/LocalLlama
Wednesday, Apr 1, 2026 · 12:22AM
Summary

everyone's talking about the claude code stuff (rightfully so) but this paper came out today, and the claims are pretty wild: 1-bit 8b param model that fits in 1.15 gb of memory ... competitive with llama3 8B and other full-precision 8B models on benchmarks runs at 440 tok/s on a 4090, 136 tok/s on

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories