Residual Stream & LayerNorm
The running workspace that transformer layers read from and write to, and the normalization that keeps deep stacks trainable.
Follow the residual branch
A four-dimensional toy block with fixed weights. These are real vector calculations, not a simulation of training stability. Channels have no assigned linguistic meaning.
y = x + F(Norm(x))
| x | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Residual input x | 1.000 | -1.000 | 2.000 | 0.000 |
| Input to F | 0.447 | -1.342 | 1.342 | -0.447 |
| Branch update F | -0.107 | -0.306 | 0.468 | -0.107 |
| Residual sum | 0.893 | -1.306 | 2.468 | -0.107 |
| Block output | 0.893 | -1.306 | 2.468 | -0.107 |
Vector entering normalization: Mean = 0.500; Variance = 1.250.
LayerNorm subtracts the channel mean and divides by √(variance + ε). Here γ = 1, β = 0 and ε = 0.00001.
F(v)ᵢ = gain × tanh(0.6vᵢ + 0.3v₍ᵢ₊₁₎ mod 4). This fixed toy branch stands in for attention or an MLP.
Pre-Norm preserves a direct residual path around normalization. Post-Norm normalizes the residual sum. Their training behavior depends on initialization, depth and optimization; these numbers do not establish which model trains better.
Layers add updates
A transformer layer does not replace a token representation from scratch. Attention and MLP blocks compute updates that are added to a shared residual stream.
- The residual stream carries the current representation for every token.
- Each block reads it, computes an update, and writes back by addition.
Why residual connections matter
Residual connections make very deep networks easier to optimize because gradients and information can flow through many layers without every block having to perfectly preserve them.
- This is one reason transformer stacks can be dozens or hundreds of layers deep.
- The stream is a vector workspace, not a literal database or memory file.
LayerNorm and pre-norm
Layer normalization keeps activations in a stable range. Modern decoder-only LLMs often use pre-norm: normalize before attention or MLP, then add the block output back.
- Pre-norm improves training stability for deep models.
- Different model families vary in exact normalization choices, such as RMSNorm.
Key Takeaways
- The residual stream is the running vector workspace of the model.
- Transformer blocks add updates instead of replacing representations.
- LayerNorm/RMSNorm keeps deep stacks numerically stable.