Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Finetuning a Reasoning LLM with Supervised or Reinforcement Learning? [D]

Via r/MachineLearning
Monday, Jun 1, 2026 · 4:23PM
Summary

Hello, I have a task to fine-tune small LLMs on annotated conversational data. The dataset contains not only the final answers, but also reasoning traces and tool-calling decisions (i.e., when the model should think and when it should call a tool). I am wondering what the best training approach woul

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories