Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Liquid AI's LFM2-24B-A2B running at ~50 tokens/second in a web browser on WebGPU

Via r/LocalLlama
Wednesday, Mar 25, 2026 · 8:59PM
Summary

The model (MoE w/ 24B total & 2B active params) runs at ~50 tokens per second on my M4 Max, and the 8B A1B variant runs at over 100 tokens per second on the same hardware. Demo (+ source code): https://huggingface.co/spaces/LiquidAI/LFM2-MoE-WebGPU Optimized ONNX models: - https://huggingface.co/Liq

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories