Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Mistral Medium 3.5 128B and Qwen 3.5 122B A10B on 4x RTX 3080 20GB

Via r/LocalLlama
Monday, May 4, 2026 ยท 4:09AM
Summary

Mistral Medium 3.5 128B with 4x3080 20GB with layer split: CUDA_VISIBLE_DEVICES=0,1,2,3 ./build/bin/llama-bench --model /data/huggingface/Mistral-Medium-3.5-GGUF/Mistral-Medium-3.5-128B-IQ4_XS-00001-of-00003. gguf -ngl 99 -d 0,16384 -fa 1 --split-mode layer ggml_cuda_init: found 4 CUDA devices (Tota

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories