Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Fix: Dual Intel Arc GPUs using all system RAM during inference - found the cause and a working fix (llama.cpp SYCL)

Via r/LocalLlama
Wednesday, Apr 8, 2026 ยท 12:51AM
Summary

If you're running dual Intel Arc GPUs with llama.cpp and your system RAM maxes out during multi-GPU inference, even though the model fits in VRAM, this post explains why and how to fix it. I've been running dual Arc Pro B70s (32GB each, 64GB total VRAM) for local LLM inference with llama.cpp's SYCL

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories