Knowledge BasePost-training: Making it Helpful

Reward Models & Human Preferences

Learning a model of what humans prefer, so the assistant can be optimized toward helpful, honest, harmless answers.

advanced#reward-model#preferences#rlhf
Full write-up in progress

This topic is on the roadmap and its detailed page — theory, math, code, quizzes and projects — is being authored. Its metadata, prerequisites and links are ready below.