Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How fast can I get a voice assistant to respond without a GPU? Qwen3-ASR and Kokoro-TTS ONNX on CPU.

Via r/LocalLlama
Friday, Jul 10, 2026 · 9:19AM
Summary

Been testing out the ONNX models to see how far I can push the CPU to take on ASR and TTS, so the GPU is completely free for running the LLM. The video attached shows me testing latency on a 2022 Macbook M2 and an AMD Ryzen 9 7900. This is just running the regex fast commands, so most of the latency

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories