Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DiffusionGemma 26b on a 4090 at up to 475t/s... and some thoughts...

Via r/LocalLlama
Thursday, Jun 18, 2026 · 10:29PM
Summary

Figured I'd post up a bit of info for anyone else who was thinking about messing with this model on a 3090/4090. Obviously I can't use the nvfp4, but I got it up and running in vLLM using diffusiongemma-26B-A4B-it-AWQ-INT4. Had to run it in a custom vLLM docker they provide for the purpose, then loa

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories