Tutorials
My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-tra…
My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me f
My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024. Transferring as much of the intuitions of building Olmo as I possibly can in the book format. The book is launching with an over 10 hour, full course with slidedecks, functional code for the training chapters, an example model completions library, and of course the free online web version. Physical orders from Manning will ship in 1-2 weeks, and Amazon a week or so after. Thanks for your support!
Related
- Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling
- Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model
- A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions
- Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
Source: Swyx (X) | 2026-07-21