Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[Help] Running big dense models faster

Via r/LocalLlama
Saturday, May 2, 2026 ยท 2:19PM
Summary

I have been trying Mistral 3.5 on my 4x RTX 3090 rig with llama.cpp. Inference is slow (about 11 t/s) even without anything being offloaded to the CPU. Here is the llama-server command I used: ./llama-server --model ../downloaded_models/Mistral-Medium-3.5-128B-UD-Q4_K_XL-00001-of-00003.gguf --port 1

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories