Multi-Head Attention, MQA & GQA

Intermediate

How attention runs many learned views in parallel, and how MQA/GQA share key-value heads to reduce inference memory.

Last updated: Jun 12, 2026

Multi-Head Attention, MQA & GQA

How attention runs many learned views in parallel, and how MQA/GQA share key-value heads to reduce inference memory.

Attention observatory

Many query lenses, fewer key-value shelves: from MHA to MQA

Switch the telescope lens to see how each query head reads the prompt, then step through the head grouping. MHA gives every query head its own key-value head, GQA shares one per group, and MQA keeps a single shared key-value head for all queries.

Shared key-value shelves
Every query head needs keys and values to read. MHA stores one K/V head per query head; GQA and MQA let several query heads share the same K/V head.
Radar sweep for this head
Query B
TherobotthatLenabuiltexplainsattention
Current token: attention

Each query head keeps its own Q projection, even when it shares K/V projections. This pattern only illustrates a distinct view.

query beams read shared key-value shelves
attention intensity
The
8%
robot
84%
that
18%
Lena
76%
built
22%
explains
36%
attention
90%
Lens array
Four constructed views out of 8 query heads. Intensities are illustrative, not measured attention or assigned linguistic roles.

Shared key-value shelves

MHA · 8 KV heads
Head grouping

Multi-head attention: every query head owns its own K/V head.

Q1
↓
K1
V1
Q2
↓
K2
V2
Q3
↓
K3
V3
Q4
↓
K4
V4
Q5
↓
K5
V5
Q6
↓
K6
V6
Q7
↓
K7
V7
Q8
↓
K8
V8

More sharing means fewer KV heads to store and read at inference time. The model keeps all 8 query lenses, but the KV cache shrinks by the sharing factor.

query heads
8
KV heads stored
8
KV memory saving
×1

Heads are learned projections

Each head has its own learned Q, K, and V projection matrices. It is not a fixed slice of the original vector; it is a learned view of the full residual stream.

  • Different heads can track syntax, copying patterns, local order, or long-range references.
  • Do not assume every head has a clean human-readable job. Many behaviors are distributed.

Concatenate, then mix

The head outputs are concatenated and passed through another learned output projection. That final projection mixes the parallel views back into one update.

  • Multi-head attention gives the layer several relationship detectors at once.
  • The result is written back through the residual connection.

MQA and GQA share key-value heads

Standard multi-head attention can give every query head its own key/value heads. MQA and GQA keep many query views while sharing fewer key-value projections.

  • MQA lets many query heads share one key-value set.
  • GQA groups query heads over fewer key-value heads.
  • This reduces inference memory and bandwidth while preserving multiple query lenses.

Key Takeaways

  • Heads are learned views, not fixed chunks.
  • MQA/GQA are attention-architecture choices: many query views, fewer key-value heads.
  • Sharing key-value heads reduces inference memory and bandwidth.