Took Qwen-35B-A3 and trained it with PPO — and honestly this is the first time I've ever seen PPO actually pull its weight (with verifiable reward). SO: On karpathy/autoresearch for parameter-golf → beats GLM-5.2 and Qwen-350B, and the ideas it spits out feel Opus4.8-like On bullshit-bench beats NEX