Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

backend-agnostic tensor parallelism has been merged into llama.cpp

Via r/LocalLlama
Thursday, Apr 9, 2026 · 2:46PM
Summary

if you have more than one GPU - your models can now run much faster -sm layer is the default behaviour, -sm tensor is the new thing to try "backend-agnostic" means you don't need CUDA to enjoy this This is experimental, and in your case the results may be poor (try different models). You have been w

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories