Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Special Architecture in AFM3 20B: Instruction Following Pruning

Via r/LocalLlama
Tuesday, Aug 4, 2026 · 1:26AM
Summary

https://openreview.net/forum?id=juARG7yu4P This is a model designed to activate ~20% of active MLP layers. It is also an MoE so it has some sparsity built-in. It's trained from scratch to use the same experts per prompt, not per token or switching per layer. Around two-thirds of a model's active par

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories