Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

What do you consider to be the minimum performance (t/s) for local Agent workflows?

Via r/LocalLlama
Saturday, Apr 25, 2026 · 12:02PM
Summary

What would you say is the minimum amount of tokens per second you would tolerate for your local agent workflows? I have been trying pi.dev connected to a llama.cpp instance running Qwen3.6-27B-Q6_K_L with 200K context running on an RTX A6000. I get about 26 t/s and is surprisingly usable. About the

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories