Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

Via r/LocalLlama
Friday, Jul 24, 2026 · 5:18PM
Summary

I've been building a C99 inference engine from scratch (no Python, no BLAS, just gcc and make) that runs BitNet's ternary models on CPU. A few weeks ago I got obsessed with the matmul kernel - wrote a new one using AVX-512BW's vpermt2w to pack 5 ternary weights per byte instead of 4, benchmarked it

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories