Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

2.5x faster Qwen3.6 NVFP4 Unsloth quants

Via r/LocalLlama
Friday, Jul 10, 2026 · 1:20PM
Summary

Hey r/LocalLLaMA folks! We made NVFP4 quants 2.5x faster for Qwen3.6 27B and also 1.56x to 1.79x faster for 35B-A3B vs NVIDIA's NVFP4 quants without any accuracy degradation! We used W4A4 so actual 4bit tensor cores for matmuls, whilst NVIDIA's ones uses W4A16. FP8 KV Cache calibration is also provi

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories