Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

RTX 5090 gemma4-26b TG performance report

Via r/LocalLlama
Monday, Apr 6, 2026 · 1:21AM
Summary

Nothing exhaustive... but I thought I'd report what I've seen from early testing. I'm running a modified version of vLLM that has NVFP4 support for gemma4-26b. Weights come in around 15.76 GiB and the remainder is KV cache. I'm running full context as well. For a "story telling" prompt and raw outpu

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories