Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Dual 3090 Gemma 4 31B QAT: MTP performance worse than without MTP.

Via r/LocalLlama
Sunday, Jul 19, 2026 · 6:15AM
Summary

Been dealing with this issue for a while with no apparent explanation. My base TPS are around 39.7tps in tensor parallelism, about 31 to 33tps with --sm layer. However, using the MTP my TPS go wildly between 27 tps to 34 max. Dual 3090 No MTP: 0.33.712.373 I slot print_timing: id 3 | task 0 | n_deco

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories