Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

PFlash: 10x prefill speedup over llama.cpp at 128K on a RTX 3090

Via r/LocalLlama
Friday, May 1, 2026 · 2:53PM
Summary

Hey fellow Llamas, thank you for all the nice words and great feedback on the last post I made. We have something new we thought would be useful to share. As always your time is precious, so I'll keep it short. We built speculative prefill for long-context decode on quantized 27B targets, C++/CUDA o

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories