Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Gemma 4 MoE hitting 120 TPS on Dual 3090s!

Via r/LocalLlama
Saturday, Apr 4, 2026 · 3:06AM
Summary

Thought I'd share some benchmark numbers from my local setup. Hardware: Dual NVIDIA RTX 3090s Model: Gemma 4 (MoE architecture) Performance: ~120 Tokens Per Second The efficiency of this MoE implementation is unreal. Even with a heavy load, the throughput stays incredibly consistent. It's a massive

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories