Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

KV cache compression on Qwen 3.6 — 1M context: 10.7GB → 6.9GB (V: 3.5× smaller)

Via r/LocalLlama
Friday, Apr 17, 2026 · 8:20PM
Summary

Quick demo of KV cache compression on Qwen 3.6 at 1M context. In this run: KV cache: 10.74 GB → 6.92 GB V cache: 5.37 GB → 1.55 GB (~3.5× reduction) Still seeing near-zero PPL change in early tests (3 seeds), but focusing mainly on memory + long-context behavior for now. Curious how people think abo

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories