Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Trained a Qwen2.5-0.5B-Instruct bf16 model on Reddit post summarization task with GRPO written from scratch in PyTorch - updates! [P]

Via r/MachineLearning
Wednesday, Apr 15, 2026 ยท 9:01AM
Summary

So, yesterday run was a success and I did get an avg rollout length of about 64 tokens as attached in the image! This was with quality_reward + length_penalty (more info below!) Next, I'll be going with length penalty as the reward and with the mistake of counting characters as tokens fixed and see

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories