Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I found why KV cache INT4 breaks on some models (Qwen2-7B: ΔPPL +238) and built a 4-line fix no training, no calibration, 12 models tested up to 40B

Via r/LocalLlama
Wednesday, Apr 15, 2026 · 8:23PM
Summary

I've been playing with KV cache INT4 quantization and noticed something weird: it works perfectly on some models and completely destroys others. Examples: Falcon-40B: ΔPPL +0.08 ✅ (basically free compression) OPT-13B: ΔPPL +0.28 ✅ Qwen2-7B: ΔPPL +238 ❌ (output becomes incoherent garbage) Pythia-6.9B

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories