Attention and Self-Attention
Queries, keys, and values built from the ground up: why attention exists, how scaled dot-product attention is computed, why it is divided by the square root of the dimension, and how causal masking works, with every matrix computed and checked.
Prerequisites: Backpropagation and Gradient Descent
Self-attention is the mechanism that made modern language models possible. It lets every position in a sequence look at every other position and decide, for itself, which ones matter. This article builds it from the problem it solves up to the full matrix computation, working every number on a three-token example small enough to check by hand.
A. The problem attention solves
Consider resolving what "it" refers to in a sentence. The information needed sits elsewhere in the sequence, and where it sits depends entirely on the sentence. A fixed-window approach cannot express "look back to whichever earlier word this one actually depends on".
Attention makes that dependency learned and data-dependent: each position computes, from the content of the tokens themselves, how much to draw from every other position.
Attention weights explorer
One query attending over four key/value pairs.
- Largest weight
- 29.7%
- Attention spread
- 0.990
| Token | q · k | ÷ √dₖ | Weight | |
|---|---|---|---|---|
| the | 0.40 | 0.20 | 22.0% | |
| cat | 1.00 | 0.50 | 29.7% | |
| sat | 0.92 | 0.46 | 28.5% | |
| down | 0.20 | 0.10 | 19.9% |
Query vector
Output - the weighted sum of the value vectors
[0.220, 0.297, 0.285, 0.199]
Attention is a weighted average, and the softmax decides the weights. Lower the temperature and it sharpens toward picking a single token; raise it and attention smears evenly across all of them. Turning off the √dₖ scaling spreads the raw scores further apart, which pushes the softmax toward saturation - the reason the factor is there at all.
B. Queries, keys, and values
Each input token embedding is projected into three vectors by three learned weight matrices:
The database analogy is genuinely apt:
- the query is what this position is looking for;
- the key is what each position advertises about itself;
- the value is what each position actually contributes if attended to.
Matching a query against a key measures relevance; the values are what get mixed. Separating "what identifies a token" (key) from "what it contributes" (value) is what gives the mechanism its flexibility - and are learned by the backward pass from Backpropagation and Gradient Descent.
C. Scaled dot-product attention
The whole mechanism, for all positions at once:
Read it in four steps: score every query against every key (), scale by , normalise each row to sum to 1 (softmax), then take the correspondingly weighted average of the values.
Raschka notes this is called scaled dot-product attention, and it is the mechanism used in the original transformer and in the GPT family.
D. Working it through completely
Three tokens, . To keep the arithmetic checkable we take and already projected and equal, with distinct values:
Step 1 - attention scores. Entry is :
Check one: row 3, column 3 is , the largest score in the matrix - token 3's query matches its own key best.
Step 2 - scale. Divide by :
Step 3 - softmax each row. For row 1, exponentiating gives , , , summing to . Dividing:
All three rows:
Every row sums to 1 - these are genuine weightings. Row 3 puts on position 3, matching the strongest score from step 1.
Step 4 - weighted values. Multiply by . Row 1, first component:
Each output row is a blend of all three value vectors, mixed by learned relevance.
Runs in your browser. The first run downloads the Python runtime (~10 MB), then it is cached.
Running it reproduces both matrices and prints matches: True True.
E. Why divide by the square root of the dimension
The scaling is not cosmetic. A dot product of two -dimensional vectors sums terms, so its magnitude grows with - roughly as for independent, unit-variance components.
Large scores are a problem for the softmax. As inputs grow, softmax approaches a one-hot vector: one weight near 1 and the rest near 0. In that saturated regime its gradients are minuscule, so learning stalls. Dividing by keeps the scores in a range where the softmax stays sensitive and gradients keep flowing.
With , unscaled scores would be about eight times larger than the scaled ones - comfortably enough to saturate.
The figure below is this same head, with two things you can move. Turning token three's query rotates it through the plane its keys live in, and only the third row of the matrix responds - keys and values do not move, which is what makes a query a query. Then raise the dimension. With the division in place, nothing happens at all: the matrix at 512 dimensions is the matrix above, to the last decimal. Switch the division off and raise it again, and watch the row collapse onto a single token. That collapse is the whole reason the square root is in the formula.
Interactive: steer a query, then change the dimension
Only token three’s query moves. Keys and values stay where the lesson put them.
Attention weights, row by row
- Divisor
- 1.4142
- Largest weight
- 0.5035
- Row 3 spread
- 1.496 bits
- Output row 3
- 1.76, 1.00
With the division in place the attention matrix does not depend on d_k at all - drag the dimension from 2 to 512 and not one weight moves. That invariance is the whole content of the square root: it holds the scaled scores at a constant size while the raw dot products grow like sqrt(d_k), so the softmax stays in the range where it still has a gradient to give back. Row three is spread across 1.496 bits of its possible 1.585.
F. Causal masking
A model that generates text left to right must not see the future. If position 2 could attend to position 3, the model would be trained with access to the answer and would fail at generation time, when the future genuinely does not exist.
Causal attention prevents this by masking: before the softmax, every score at is set to , so and those positions receive zero weight. The remaining weights renormalise to sum to 1.
Applying it to our scaled scores:
Row 1 attends only to itself, so its weight is forced to . Row 2 splits between positions 1 and 2 - note these are not the unmasked row-2 values and ; with position 3 removed, the remaining two are renormalised by their own sum, , giving and . (Carrying the rounded values through by hand gives and ; the figures above come from the full-precision computation.) Row 3, which was never allowed to see anything beyond position 3, is unchanged.
Raschka also notes a dropout mask is often applied to the attention weights during training, to reduce overfitting.
G. Multiple heads
One attention computation captures one kind of relationship. Multi-head attention runs several in parallel with separate , then concatenates the outputs and projects them back down. Different heads can specialise - one tracking syntactic dependencies, another longer-range topical links - and the model is not forced to squeeze every relation into a single weighting.
Key takeaways
- Attention lets each position decide, from content, how much to draw from every other position.
- Each token is projected into a query, key, and value; queries match against keys, and values are what get mixed.
- The mechanism is - score, scale, normalise, blend.
- The divisor prevents softmax saturation and keeps gradients usable.
- Causal masking sets future scores to so weights renormalise over the past only.
- Multi-head attention runs several of these in parallel to capture different relationships.
What's next
Attention is the core of the transformer, but a working language model also needs tokenization, embeddings, positional information, feed-forward blocks, and a training objective. Those pieces are covered across the rest of this track, and the classical-AI counterpart to "search the space of possibilities" is developed in Adversarial Search and Minimax.
References & further reading
- Sebastian Raschka, Build a Large Language Model (From Scratch), Manning, 2025· Kudos AI reference library
Copyrighted works are cited for reference only and are not hosted here; please consult the publisher for access.