Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

vLLM PR adding native HIP W4A16 kernel was merged

Via r/LocalLlama
Friday, May 29, 2026 · 12:31PM
Summary

The performance increase introduced by the PR is awesome. Makes my ROCm rig a lot more useful. Numbers from the PR: Kernel dtype max-num-seqs=8 max-num-seqs=32 Triton W4A16 bf16 82.4 tk/s - Triton W4A16 fp16 83.2 tk/s - ExLlama (no bf16) fp16 255.0 tk/s 382.5 tk/s RDNA3 W4A16 (this PR) bf16 205.3 tk

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories