Knowledge BasePost-training: Making it Helpful

RLHF & DPO

Aligning the model to human preferences with reinforcement learning (PPO) or the simpler direct approach (DPO).

advanced#rlhf#ppo#dpo#alignment
Full write-up in progress

This topic is on the roadmap and its detailed page — theory, math, code, quizzes and projects — is being authored. Its metadata, prerequisites and links are ready below.