Quick notes after running StepFun Step-3.7-Flash on AMD with ROCm. The two things that matter most: Do not run ROCm past ~94k context. On my setup, ROCm corrupts long context somewhere around 94k tokens. The model usually does not crash. It just loops, burns the token budget, and never gives a usabl