Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Gemma-4-31B NVFP4 inference numbers on 1x RTX Pro 6000

Via r/LocalLlama
Friday, Apr 3, 2026 ยท 4:48PM
Summary

Ran a quick inference sweep on gemma 4 31B in NVFP4 (using nvidia/Gemma-4-31B-IT-NVFP4). The NVFP4 checkpoint is 32GB, half of the BF16 size from google (63GB), likely a mix of BF16 and FP4 roughly equal to FP8 in size. This model uses a ton of VRAM for kv cache. I dropped the kv cache precision to

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories