Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

DFlash speculative decoding on Apple Silicon: 4.1x on Qwen3.5-9B, now open source (MLX, M5 Max)

Via r/LocalLlama
Monday, Apr 13, 2026 · 3:48PM
Summary

A few weeks ago I posted early results from a native MLX implementation of DFlash. Since then I rewrote the benchmark methodology, fixed numerical issues, and open sourced the whole thing. A small draft model generates 16 tokens in parallel via block diffusion, the target verifies them in one forwar

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories