Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise?

Via r/LocalLlama
Thursday, Jul 9, 2026 · 3:53PM
Summary

TLDR: 75B-total / 9B-active MoE is the perfect shape for multi-24GB rigs, and almost nobody ships it. Qwen 27B is a great model and punches way above its weight-class, it is a frequent fallback for me. Nemotron-3-Puzzle-75B-A9B, NVFP4, vLLM 0.22.1 (the new Marlin fallbacks run FP4 on Ampere), pipeli

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories