Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How does Pi coding agent control Qwen's thinking verbosity? (Qwen 35B A3B, llama-server)

Via r/LocalLlama
Sunday, May 17, 2026 · 6:07AM
Summary

I'm running Qwen 35B A3B via llama-server with reasoning budget set to -1 (unlimited) for testing. In every client I've tried, the model just thinks endlessly before responding. But with Pi, it does the bare minimum thinking and still responds fairly accurately - which is a stark difference. My firs

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories