Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

How to fine-tune an LLM for open-ended problems? [P]

Via r/MachineLearning
Saturday, May 30, 2026 · 2:42PM
Summary

I want to develop an LLM that can solve open-ended math problems (such as proof-only problems). This means that RLVR where we use the final answer alone as reward signal is not enough. Since SFT is useless here and GRPO/PPO methods will not have an appropriate reward function, what kind of fine-tuni

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories