Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Devs - you have 64gb of VRAM - which model do you use for coding?

Via r/LocalLlama
Tuesday, Jun 30, 2026 · 8:03PM
Summary

I've currently settled on an unsloth version of Qwen 3.5 122b-a10b model (UD-IQ4_NL). With 100k bf16 context window, I only had to load a few layers into CPU/RAM, it runs around 30 tok/sec which is fine for me. I've tested many models, hours of testing but I am currently deeply impressed with this o

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories