Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Consolidated my homelab from 3 models down to one 122B MoE — benchmarked everything, here's what I found

Via r/LocalLlama
Friday, Mar 27, 2026 · 12:39AM
Summary

Been running local LLMs on a Strix Halo setup (Ryzen AI MAX+ 395, 128GB RAM, 96 GiB shared GPU memory via Vulkan/RADV) under Proxmox with LXC containers and llama-server. Wanted to share where I landed after way too much benchmarking. THE OLD SETUP (3 text models) - GLM-4.7-Flash: 30B MoE 3B active,

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories