AI Academy / Simulation
See what attention actually calculates
Edit three tiny vectors and inspect dot products, normalized weights and the weighted output. No model downloads or API calls.
You will learn to
- Compute scaled dot-product attention
- Explain causal masking
- Observe temperature and normalization
Before you start
Vectors, dot products and weighted averages
A deliberately small model
Three positions each have a two-dimensional vector. This experiment uses the same vectors as queries, keys and values, equivalent to identity projections. Real transformers learn different projections and combine multiple heads. The position labels are just A, B and C; they have no learned language meaning.
From similarity to weights
For each query, multiply matching coordinates with a key and sum them. Divide by the square root of the dimension, here √2. We additionally divide by the temperature control to explore concentration. Softmax exponentiates these scores and divides each result by the sum. Each row of weights sums to one. The implementation subtracts the largest available score before exponentiation for numerical stability.
Keep future positions out
With the causal mask enabled, row A can attend only to A, row B to A and B, and row C to all three. Masked scores receive zero weight after normalization. Without the mask, the first position can use later positions. A model trained to predict the next token must not peek at the future answer in its context.
Test an invariant
Set all vectors to 0,0 and disable the mask. Every score is zero, so each row has three weights of one third. Enable the mask: the rows become [1,0,0], [0.5,0.5,0], and [1/3,1/3,1/3]. The output is still zero because every value is zero. This is a real computation of one operation, not a language model, training run or generated response.
Try it yourself
Enable JavaScript for this interactive activity. You can read all lesson explanations above without it.