"inference falls back to dense attention" for MiniMax M3 - does it mean 428B weights used at each step?
Via r/LocalLlama
Friday, Jun 12, 2026 · 9:28PM
Summary
So like 100x (or how much) slower vs. full implementation? https://huggingface.co/unsloth/MiniMax-M3-GGUF Note: MiniMax Sparse Attention is not supported yet, so inference falls back to dense attention.