Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Turbo-fieldfare: Open-source engine running Gemma 4 26B in 2 GB RAM on Apple Silicon

Via r/LocalLlama
Thursday, Jul 30, 2026 · 12:46PM
Summary

Its a custom Swift/Metal inference engine that runs Gemma 4 26B-A4B-IT on M-series Macs with very low RAM. It uses ~2GB instead of ~14 GB. The result is reportedly 5–6 tok/s on an 8 GB M2 MacBook Air and 31–35 tok/s on an M5 MacBook Pro. It also includes an OpenAI-compatible local server with stream

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories