Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Turboquant+MTP for ROCm(Llama CPP)

Via r/LocalLlama
Thursday, May 14, 2026 · 8:24AM
Summary

TL;DR: I got TBQ4 KV cache + MTP working on AMD ROCm for RX 7900 XTX / RDNA3 / gfx1100 in llama.cpp. Main win: 64k context fits on 24 GB VRAM and remains usable. Branch: tbq4-rdna3-experiment (https://github.com/DrBearJew/llama.cpp/tree/tbq4-rdna3-experiment) I dug into TurboQuant / TBQ4 + MTP on AM

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories