π± Course: https://github.com/anakin87/llm-rl-environments-lil-course | π₯ Video: https://www.youtube.com/watch?v=71V3fTaUp2Q I've been deep into RL for LLMs lately. Over the past year, we've seen a shift in LLM Post-Training. Previously, Supervised Fine-Tuning was the most important part: making mode