A weighted combination of vectors
Self-attention transforms each input representation into a query, key and value using learned projections. For one position, its query is compared with the keys. The resulting scores are scaled, masked where needed and normalized with softmax. Those weights combine the values into a contextual vector for that position.
Query (Q)
The projected vector used to score keys for the selected position.
Key (K)
The projected vector at each candidate position, compared with the query by a dot product.
Value (V)
The projected information combined according to the attention weights. A key determines a weight; a value contributes to the output.
Calculate one attention head
This example contains three words. Each word is treated as one token here. Input vectors and projection weights are chosen for easy arithmetic, not learned from language. Attention weights and the output below are computed from them.
Fixed projections (row vectors: Q = XWQ, K = XWK, V = XWV)
WQ
[1.00, 0.50]
[0.00, 1.00]
WK
[0.50, 1.00]
[1.00, 0.00]
WV
[1.00, -1.00]
[0.50, 1.00]
3. Inspect all token vectors
| Token | x | Q | K | V |
|---|---|---|---|---|
| The | [1.00, 0.00] | [1.00, 0.50] | [0.50, 1.00] | [1.00, -1.00] |
| cat | [0.00, 1.00] | [0.00, 1.00] | [1.00, 0.00] | [0.50, 1.00] |
| sleeps | [1.00, 1.00] | [1.00, 1.50] | [1.50, 1.00] | [1.50, 0.00] |
4. Dot product → scale → mask → softmax
The selected query scores every key. Divide by √dₖ, set forbidden positions to −∞, then apply softmax across the row. The weights sum to 1; masked positions contribute zero.
| Token | Q · K | ÷ √2 | Attention weight | Weight × V |
|---|---|---|---|---|
| The | 1.000 | 0.707 | 0.670 | [0.67, -0.67] |
| cat | 0.000 | 0.000 | 0.330 | [0.17, 0.33] |
| sleeps | 1.000 | −∞ | 0.000 | [0.00, 0.00] |
5. Sum the weighted values
Σⱼ aⱼVⱼ = [0.83, -0.34]
This output is a contextual vector for the selected position, not a next-token probability. Later projections and layers process it. Large attention weight alone does not establish grammar, importance or a causal explanation of the model’s answer.
One head, in matrix notation
Attention(Q,K,V) = softmax(QKᵀ / √dₖ + M) V
Q and K have width dₖ. The dot products form a score for each query-key pair. Scaling by √dₖ controls score magnitude; softmax runs across the allowed keys in each row. M contains 0 for allowed positions and −∞ for masked positions.
Visibility depends on the task
A causal decoder masks future tokens so training cannot reveal the answer it must predict. A bidirectional encoder can use positions on both sides. Padding masks and other attention patterns impose additional restrictions. The mask changes which positions are available; it does not assign linguistic roles to heads.
What scales quadratically?
Dense full attention computes n² query-key scores per head during prefill. The table counts scores, not total operations. Dot products also depend on head width, and models have multiple heads and layers. The displayed storage is only one materialized score matrix at two bytes per element.
| Sequence length | Score entries per head | Naive score storage |
|---|---|---|
| 4,096 tokens | 16,777,216 | 32 MiB |
| 8,192 tokens | 67,108,864 | 128 MiB |
| 16,384 tokens | 268,435,456 | 512 MiB |
Memory-efficient attention
FlashAttention reorganizes exact attention into blocks to reduce transfers between GPU memory levels and avoid storing the full n × n score matrix. It does not turn dense attention into a linear-arithmetic algorithm. Speed and memory savings depend on shapes, precision, kernels and hardware. Cached single-token decoding has a different workload: one query reads the existing keys and values.
Attention weights are not an explanation by themselves
Weights describe one operation in one head and layer. Other heads, value vectors, residual paths and later computation also affect the answer. A prominent link does not prove that a word caused the final decision, and heads do not have permanently assigned grammar or facts jobs.