Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Mistral Medium 3.5 on AMD Strix Halo

Via r/LocalLlama
Sunday, May 3, 2026 · 6:49PM
Summary

TLDR; it's slow as heck. Run overnight. I asked it a question about codebase architecture. For an end-to-end prompt of 48k tokens + 4k thinking tokens, it took about 2 hours. llama-server -hf unsloth/Mistral-Medium-3. 5-128B-GGUF:UD-Q5_K_XL --temp 0.7 --host 0.0.0.0 --port 8080 -c 80000 -fa on -ngl

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories