Overview
News
Technologies
Salaries
Products
People
Growth
Financials

Overview

A free online book and course on RLHF, preference tuning, reward models, RLVR, and post-training language models.

News

News RLHF Book 2 months ago
Post-Training Course by Nathan Lambert
Free course lectures on RLHF, reward models, preference tuning, RLVR, and modern LLM post-training.
Read more
Report
News RLHF Book 6 months ago
Reinforcement Learning RLHF and Post-Training Book by Nathan Lambert
Policy gradient methods for RLHF and LLM post-training, including PPO, REINFORCE, RLOO, GRPO, and implementation details.
Read more
Report