Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

120 tok/s on 12GB VRAM with Gemma 4 12B QAT MTP

Via r/LocalLlama
Saturday, Jun 6, 2026 · 6:53PM
Summary

Google just released the QAT (Quantization-Aware Training) variant of their Gemma 4 models, including 12B, so it was only natural for me to benchmark it on my 12GB GPU since it fits entirely in VRAM. I was pleasantly surprised of the result! By using Google's QAT assistant / draft (gemma-4-12B-it-qa

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories