Overview
News
Technologies
Salaries
Products
People
Growth
Financials
Overview
A free online book and course on RLHF, preference tuning, reward models, RLVR, and post-training language models.
News
Post-Training Course by Nathan Lambert
Free course lectures on RLHF, reward models, preference tuning, RLVR, and modern LLM post-training.
Read more
Report
Reinforcement Learning RLHF and Post-Training Book by Nathan Lambert
Policy gradient methods for RLHF and LLM post-training, including PPO, REINFORCE, RLOO, GRPO, and implementation details.
Read more
Report
