Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

"inference falls back to dense attention" for MiniMax M3 - does it mean 428B weights used at each step?

Via r/LocalLlama
Friday, Jun 12, 2026 · 9:28PM
Summary

So like 100x (or how much) slower vs. full implementation? https://huggingface.co/unsloth/MiniMax-M3-GGUF Note: MiniMax Sparse Attention is not supported yet, so inference falls back to dense attention.

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories