Transformer Architecture

Intermediate

Understand the original Transformer and a modern decoder variant, including attention, feed-forward blocks, residual paths and normalization.

Last updated: Sep 13, 2026

What is the Transformer?

The Transformer is a neural network architecture introduced in the landmark 2017 paper "Attention is All You Need" by Vaswani et al. It replaced recurrent and convolutional approaches with a purely attention-based mechanism, enabling massive parallelization during training and capturing long-range dependencies far more effectively. Nearly every modern large language model -- GPT, BERT, LLaMA, Claude -- is built on the Transformer.

“Attention is all you need.”

-- Vaswani et al., "Attention Is All You Need" (2017, Google Brain)

LLM Visualization by Brendan Bycroft

The best interactive 3D visualization of transformer internals available. Explore how GPT-style models process tokens through embedding, attention, and feed-forward layers -- step by step, parameter by parameter. Highly recommended.

Interactive 3DBy Brendan Bycroftbbycroft.net/llm

Transformer Layer Stack

Compare a Pre-Norm decoder example with the historical Post-Norm structure from the 2017 paper. These are architecture sketches; tensor dimensions, masks and exact blocks depend on the model.

Example following LLaMA (2023): RMSNorm before attention and FFN, RoPE in queries/keys, and a SwiGLU FFN. Other models can use different blocks.

1x = token embeddings
2a = CausalAttention(RMSNorm(x)); RoPE(Q, K)
3x₁ = x + a
4f = SwiGLU(RMSNorm(x₁))
5x₂ = x₁ + f

Repeat the block, then apply final normalization and the LM head. Both additions preserve a direct residual path. RoPE acts on Q and K, rather than adding a position vector to x.

Architecture Variants

The original Transformer used both an encoder and decoder. Modern models often use just one. Toggle between the three main variants to see which components each uses.

Encoder-Decoder Variants

Toggle to compare architectures

Encoder
Self-Attention
Add & Norm
FFN
Add & Norm
Cross-Attention
Decoder
Masked Self-Attn
Cross-Attention
Add & Norm
FFN
Add & Norm

Example models:

T5BARTmBART

The original architecture. The encoder processes the full input bidirectionally, then the decoder generates output tokens one at a time, attending to the encoder's representations via cross-attention. Used for translation, summarization, and sequence-to-sequence tasks.

Token Dataflow

Follow a single token as it flows through the entire Transformer pipeline, from raw text to output probabilities. Watch how the tensor shape changes at each step.

Step-by-Step Dataflow

Play or scrub through the processing pipeline

1
Tokenize
Raw text is split into token IDs using BPE or similar. Each token maps to an integer.
[batch, seq_len]
2
Embed
[batch, seq_len, d_model]
3
Add Position
[batch, seq_len, d_model]
4
Compute Attention
[batch, heads, seq_len, seq_len]
5
Attention Output
[batch, seq_len, d_model]
6
Feed-Forward
[batch, seq_len, d_ff]
7
FFN Output
[batch, seq_len, d_model]
8
Output Logits
[batch, seq_len, vocab_size]
TokenizeOutput Logits

Key Concepts

Residual Connections

Residual paths provide a direct route for information and gradients around sublayers and can make deep networks easier to optimize. They reduce optimization difficulties but do not alone guarantee stable training.

Layer Normalization

LayerNorm normalizes across a token’s feature dimension. The original Transformer uses Post-Norm; many decoder architectures use Pre-Norm or RMSNorm. The order changes the residual path.

Positional Encoding

Since attention has no inherent notion of order, position must be explicitly injected. The original paper used fixed sinusoidal functions; modern models typically use learned position embeddings or relative position encodings like RoPE.

Why It Matters

The Transformer architecture is arguably the most impactful innovation in AI of the past decade. It unlocked the scaling laws that make modern LLMs possible.

  • 1The Transformer replaced RNNs and LSTMs by enabling full parallelization during training, reducing training time from weeks to days
  • 2Its attention mechanism captures long-range dependencies that sequential models struggled with, enabling understanding of entire documents
  • 3The architecture scales remarkably well -- from 100M parameter BERT to 1.8T parameter GPT-4, performance improves predictably with scale
  • 4Every major LLM today (GPT, Claude, Gemini, LLaMA, Mistral) is built on the Transformer, making it the foundation of modern AI