Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Benchmarked Needle 26M vs Qwen3-0.6B on CPU function calling, 50 queries across 5 difficulty tiers. The 23x smaller model wins on accuracy and is 4.4x faster.

Via r/LocalLlama
Saturday, May 23, 2026 · 3:38PM
Summary

Ran a head-to-head on two open-weight models for tool-calling on a 4-core CPU, no GPU, no cherry-picking. Wanted to see if the small specialist (Needle, 26M, distilled from Gemini 3.1 for function calls) actually holds up against a small generalist (Qwen3-0.6B) that also does tools. Setup: 50 querie

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories