Hey reddit fam! If you have an AMD GPU and have ever tried to run a local LLM on it, you know the pain. ROCm doesn't support consumer cards. vLLM won't work. llama.cpp kind of works through Vulkan but treats your GPU like an afterthought - generic shaders, no architecture tuning, no real serving sto