Positional Encoding & RoPE

Expert

How transformer LLMs know token order, why RoPE is common, and why long context can still miss the middle.

Last updated: Jul 12, 2026

Positional Encoding & RoPE

How transformer LLMs know token order, why RoPE is common, and why long context can still miss the middle.

RoPE

RoPE phase orbits: position becomes rotation

Each token rides around a phase wheel. Shift both query and key together and their planets move, but the angular separation RoPE compares stays locked.

relative offset
3same phase gap
current orbit
Token orbit map
Δθ 1.35 rad = 1.35 rad
012345678qk
q vector
k vector
ghosted shifted pair
Token orbit map
relative phase survives translation
query planet5
key planet2
shifted pair7 → 4
RoPE similarity cue61%

For aligned unit vectors: angle(qᵢ) − angle(kⱼ) = (i − j) × rotation step

Two initially aligned unit vectors in 2D. The cue maps cosine from [−1, 1] to [0, 100]; it is not an attention probability. Real RoPE rotates many dimension pairs at different frequencies, with content-dependent input vectors.

Why position has to be added

Self-attention sees token vectors, but order is not built in. Without a position signal, the model has no direct way to distinguish “dog bites man” from “man bites dog”.

  • Early transformers added fixed sinusoidal patterns to token embeddings.
  • Learned absolute position embeddings assign a learned vector to each slot.
  • Both approaches give order information, but long-context generalization can be fragile.

RoPE in modern LLMs

Rotary Position Embeddings rotate parts of the query and key vectors by an angle based on position. When attention compares two tokens, relative distance is encoded in the rotation difference.

  • RoPE is common in LLaMA-, Mistral-, Gemma-, and Qwen-style models.
  • It does not add a separate learned vector for every absolute position.
  • Long-context models often extend RoPE with scaling or interpolation tricks.

Lost in the middle

More context length does not mean perfect recall. Models often use information near the beginning or end of a prompt more reliably than information buried in the middle.

  • Position encodings, attention patterns, training mix, and prompt structure all matter.
  • Important facts should be placed clearly and retrieved or verified when stakes are high.

Key Takeaways

  • Transformers need explicit position information; self-attention alone is order-agnostic.
  • RoPE encodes relative distance through rotations in Q/K space.
  • Long context improves capacity, not guaranteed recall or reasoning.