Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Comparing dual-GPU inference speed between llama.cpp row/tensor split and ik_llama graph split

Via r/LocalLlama
Friday, Jun 12, 2026 · 4:43PM
Summary

Setup: +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 610.43.02 KMD Version: 610.43.02 CUDA UMD Version: 13.3 | +-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories