Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

llama-server router: a model pinned to one GPU still grabs a CUDA context on every card, so it OOMs when my others are full. Am I missing a flag or is this just how it is?

Via r/LocalLlama
Sunday, Jun 7, 2026 · 9:09PM
Summary

Running into something annoying with llama-server in router mode (`--models-preset`) and I can't tell if I'm missing a flag or if this is just how it works. My rig is 2x 3090, 2x 4060 Ti (one's unplugged at the moment, riser got repurposed) and a 5060 Ti. I run a single llama-server router that spaw

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Is this the dawn of the Tokenpocalypse?
TechCrunch AI · Industry & Money
Amazing Digital Dentures (a failed project)
Hugging Face Blog · Models & Research
OpenAI is still working on that ‘super app’
TechCrunch AI · Industry & Money
Back to all stories