Positional Encoding & RoPE
How transformer LLMs know token order, why RoPE is common, and why long context can still miss the middle.
RoPE phase orbits: position becomes rotation
Each token rides around a phase wheel. Shift both query and key together and their planets move, but the angular separation RoPE compares stays locked.
For aligned unit vectors: angle(qᵢ) − angle(kⱼ) = (i − j) × rotation step
Two initially aligned unit vectors in 2D. The cue maps cosine from [−1, 1] to [0, 100]; it is not an attention probability. Real RoPE rotates many dimension pairs at different frequencies, with content-dependent input vectors.
Why position has to be added
Self-attention sees token vectors, but order is not built in. Without a position signal, the model has no direct way to distinguish “dog bites man” from “man bites dog”.
- Early transformers added fixed sinusoidal patterns to token embeddings.
- Learned absolute position embeddings assign a learned vector to each slot.
- Both approaches give order information, but long-context generalization can be fragile.
RoPE in modern LLMs
Rotary Position Embeddings rotate parts of the query and key vectors by an angle based on position. When attention compares two tokens, relative distance is encoded in the rotation difference.
- RoPE is common in LLaMA-, Mistral-, Gemma-, and Qwen-style models.
- It does not add a separate learned vector for every absolute position.
- Long-context models often extend RoPE with scaling or interpolation tricks.
Lost in the middle
More context length does not mean perfect recall. Models often use information near the beginning or end of a prompt more reliably than information buried in the middle.
- Position encodings, attention patterns, training mix, and prompt structure all matter.
- Important facts should be placed clearly and retrieved or verified when stakes are high.
Key Takeaways
- Transformers need explicit position information; self-attention alone is order-agnostic.
- RoPE encodes relative distance through rotations in Q/K space.
- Long context improves capacity, not guaranteed recall or reasoning.