Knowledge BaseServing & Ops

Quantization & Compression

int8/int4, GPTQ/AWQ, distillation — shrinking models for cheaper, faster inference.

advanced#quantization#int8#distillation
Full write-up in progress

This topic is on the roadmap and its detailed page — theory, math, code, quizzes and projects — is being authored. Its metadata, prerequisites and links are ready below.