Knowledge BaseShipping the Model

Evaluating LLMs

How we measure capability and safety — benchmarks, held-out evals, red-teaming, and their real limitations.

intermediate#evaluation#benchmarks#red-teaming
Full write-up in progress

This topic is on the roadmap and its detailed page — theory, math, code, quizzes and projects — is being authored. Its metadata, prerequisites and links are ready below.