Following up on my previous Gemma 4 31B benchmark post, I tested speculative decoding with Gemma 4 E2B (4.65B) as the draft model. The results were much better than I expected, so I wanted to share some controlled benchmark numbers. Setup GPU: RTX 5090 (32GB VRAM) Main model: Gemma 4 31B UD-Q4_K_XL