Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Qwen 3.6 27B in Claude Code says it will do something then stops and prompts for user reply (not failing a tool call)

Via r/LocalLlama
Sunday, Apr 26, 2026 · 5:19PM
Summary

I'm running Qwen/Qwen3.6-27B-FP8 via vLLM using this command: vllm serve Qwen/Qwen3.6-27B-FP8 --tensor-parallel-size 4 --gpu-memory-utilization 0.95 --max-num-seqs 8 \ --enable-auto-tool-choice --tool-call-parser qwen3_xml \ --enable-prefix-caching --attention-backend flashinfer It works pretty well

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories