Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Pushing the limit: minimax m2.7 q8_0 128k on 2x3090, 256GB DDR4

Via r/LocalLlama
Sunday, May 17, 2026 · 9:59PM
Summary

CPU is just a secondhand 10900x. Using 128k context, unquantized kv cache. Model is at q8_0 to mitigate some weird behavior I was seeing at lower quants. Speed is very slow at around 50tps pp, 10tps tg, but usable for coding agent workflows. Anybody else running MoE models in this size class on rela

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories