Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Intel Arc Pro B70 32GB performance on Qwen3.5-27B@Q4

Via r/LocalLlama
Saturday, Apr 11, 2026 · 5:52AM
Summary

Posted something when I initially got the GPU on r/IntelArc. Did not have vllm working at the time, so no real use case numbers. After many nights fighting with vllm, I finally got it to work. Here are some summery. both llama.cpp and llm-scaler-vllm produce ~12tps token generation rate. tensor para

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories