Skip to content
Kudos AI

Self-Attention

A mechanism that lets every position in a sequence attend to every other, computing each output as a weighted sum of values whose weights come from query-key similarity.

Also known as: Scaled dot-product attention

One token gathering context from the rest of the sequence, with the attention weights forming and re-forming as the query moves.

Understanding Self-Attention

Interpreting a word requires context that may sit anywhere in the sentence. Self-attention gives every position direct access to every other position, letting the model decide, for each token, which other tokens are relevant to it, and read from those.

The mechanism projects each token’s representation into three vectors. The query expresses what this position is looking for, the key advertises what this position offers, and the value carries the content actually retrieved. The relevance of position j to position i is the dot product of i’s query with j’s key. These scores are passed through a softmax to become weights that sum to one, and the output at i is the weighted sum of all value vectors.

The scaling factor is not cosmetic. Raschka explains that dot products grow with the embedding dimension, which for GPT-style models is typically over a thousand; large dot products push the softmax toward behaving like a step function, and its gradients toward zero, which slows or stalls learning. Dividing by the square root of the key dimension keeps the scores in a range where the softmax stays responsive, and this normalization is what gives scaled dot-product attention its name.

The structural gain over recurrence is twofold. Any two positions are one operation apart regardless of distance, so long-range dependencies do not have to survive a long chain of intermediate steps. And because every position’s output is computed from the same matrices independently, the whole sequence is processed in parallel during training, rather than one step at a time. The cost is that comparing all pairs scales quadratically with sequence length, which is why long-context efficiency remains an active research problem.

How to Calculate

Attention(Q, K, V) = softmax(QKᵀ / √dₖ) V

where

Q
queries: what each position is looking for
K
keys: what each position offers for matching
V
values: the content read out once weights are decided
dₖ
the key dimension; dividing by its square root keeps softmax from saturating
softmax
normalizes the scores into weights summing to one

Example of Self-Attention

In "the animal did not cross the street because it was too tired", resolving "it" requires knowing whether it refers to the animal or the street. Self-attention lets the position holding "it" form a query that matches strongly against the key at "animal", so the value read into that position carries animal-related content.

Changing the final word to "wide" should shift the reference to the street. Nothing in the architecture is hard-coded for pronoun resolution; the query and key projections that produce this behaviour are learned from data alone.

In practice several attention heads run in parallel with separate projections, letting one head track syntactic agreement while another tracks reference and another topical similarity. Their outputs are concatenated, which is why the mechanism is normally deployed as multi-head attention.

Frequently Asked Questions

What is the difference between attention and self-attention?

Attention generally lets one sequence read from another, as when a translation decoder reads the encoded source. Self-attention is the case where queries, keys, and values all come from the same sequence, so it relates positions within one input to each other.

Why divide by the square root of the key dimension?

Dot products grow with dimensionality, and large scores make softmax approach a step function whose gradients are near zero, which stalls training. Scaling keeps the scores in a range where softmax remains smooth and gradients remain useful.

What is causal masking?

For generative models, positions must not attend to later positions, or the model would see the answer it is meant to predict. A causal mask sets those scores to negative infinity before the softmax, so their weights become zero.

The Bottom Line

Self-attention computes each position’s output as a learned weighted read over the whole sequence, with weights from scaled query-key similarity. It removes the distance penalty of recurrence and enables parallel training, at a quadratic cost in sequence length.