Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

llama.cpp slower on P-Cores than on E-Cores with MoE Model and GPU+CPU offloading?

Via r/LocalLlama
Friday, Jul 24, 2026 · 6:29PM
Summary

I am currently experimenting with my setup: RTX 5090 + Intel 270K Plus CPU (8 Performance Cores + 16 Efficiency Cores) + 128 GB DDR5-6000 RAM, Ubuntu 26.04. I wanted to test the performance of Qwen 3.5 122b a10b with CPU offloading. Some mentioned that pinning llama.cpp to CPU performance cores coul

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories