Hi everyone, I've been trying to optimize my setup to use OpenCode with Qwen 3.6 27B (Unsloth quant Q4_K_XL) on my RX 7900 XTX with ROCm in llama.cpp. And I'm confused, it can run ok for small prompt, it seems people are using for agentic coding, but I'm seeing a collapse in the prompt processing sp