Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Is there a way to disable reasoning per request in llama.cpp's llama-server, while leaving it on by default?

Via r/LocalLlama
Tuesday, May 19, 2026 · 11:02PM
Summary

Title. I've got a llama.cpp server running a model being accessed across a number of scripts, and some of them are easier for the model than others, and those easier ones are also latency dependent. Rather than host two different servers with different parameters, I'd rather just send something alon

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories