Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Deepseek V4 Flash ~105 t/s on two Nvidia 4090d 48G (ada) in vLLM

Via r/LocalLlama
Thursday, Jul 23, 2026 · 7:01PM
Summary

TLDR: I (with the help of AI) re-implemented every Blackwell-only kernel (DeepGEMM, FlashInfer sparse-MLA, block-scaled FP8) in Triton, because they simply don't exist for sm89. The performance is 2-3x more for parallel agentic workflows. Benchmark llama-server vs vLLM I was inspired by the post htt

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories