Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Inkling-Small 276B-A12B at ~2.9 tok/s on

Via r/LocalLlama
Wednesday, Aug 5, 2026 · 6:35PM
Summary

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B total, ~12B active, 3.4 GB resident set, ~148 GB on disk. Measured on my M5, 24GB: Prompt Type Prompt /

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories