Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

TurboQuant for weights: near‑optimal 4‑bit LLM quantization with lossless 8‑bit residual – 3.2× memory savings

Via r/LocalLlama
Friday, Mar 27, 2026 · 11:22AM
Summary

an adaptation of the recent TurboQuant algorithm (Zandieh et al., 2025) from KV‑cache quantization to model weight compression. It gives you a drop‑in replacement for nn.Linear with near‑optimal distortion. Benchmarks (Qwen3.5‑0.8B, WikiText‑103) Config Bits PPL Δ PPL Compressed Size Baseline bf16 1

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories