RLHF & DPO
Aligning the model to human preferences with reinforcement learning (PPO) or the simpler direct approach (DPO).
advanced#rlhf#ppo#dpo#alignment
Full write-up in progress
This topic is on the roadmap and its detailed page — theory, math, code, quizzes and projects — is being authored. Its metadata, prerequisites and links are ready below.