VisualizationsDeep Learning

Self-Attention

Hover a token to see where it attends. Adjust the temperature to sharpen or soften the attention distribution.

ThecatsatonthematThecatsatonthemat0.370.170.080.110.130.150.150.270.140.100.200.140.060.120.390.100.210.120.110.120.130.400.160.090.100.160.200.120.290.130.150.150.150.090.170.29
Hover a row to see where that token attends.

Each row is a query token; each cell is a softmax attention weight over the keys (rows sum to 1). Low temperature sharpens attention onto one token; high temperature spreads it out. Weights here come from random embeddings — it's the mechanism, not linguistics.