In this paper, Li, Li, and Zhou review formal theory describing how inference-time compute can be traded off for higher or lower inference competence, and apply that theory to a handful of familiar open-weight LLMs (Llama-3.2, Qwen1.5, Qwen2.5, and Qwen3): https://arxiv.org/abs/2606.06574v1 > Large