Best AI News β€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I RL-trained Qwen3.6-35B-A3B to RL-train small task-specific Qwen models. Fully open source! πŸ€“

Via r/LocalLlama
Tuesday, Jul 14, 2026 Β· 12:46PM
Summary

πŸ‘‹ Training my first RL model last year was super fun, now I've RL-trained a model that RL-trains other models... wild times! The agent gets a task, writes the full training job (environment, reward, dataset, hyperparameters), and submits it to real GPUs. When the model it trained scores higher on a

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories