Saw the Qwen3-TTS thread this morning and it finally pushed me to write this up. Background: ive been building a local voice assistant for a client over the past 3 weeks. Voice-first interface on top of a RAG backend -- use case is an AI assistant where they need responses that feel conversational,