Inference Optimization
KV caching, continuous batching, speculative decoding — serving tokens fast and cheap.
advanced#inference#kv-cache#batching#serving
Full write-up in progress
This topic is on the roadmap and its detailed page — theory, math, code, quizzes and projects — is being authored. Its metadata, prerequisites and links are ready below.