Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Speculative Decoding works great for Gemma 4 31B with E2B draft (+29% avg, +50% on code)

Via r/LocalLlama
Sunday, Apr 12, 2026 · 12:08PM
Summary

Following up on my previous Gemma 4 31B benchmark post, I tested speculative decoding with Gemma 4 E2B (4.65B) as the draft model. The results were much better than I expected, so I wanted to share some controlled benchmark numbers. Setup GPU: RTX 5090 (32GB VRAM) Main model: Gemma 4 31B UD-Q4_K_XL

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories