What is the Transformer?
The Transformer is a neural network architecture introduced in the landmark 2017 paper "Attention is All You Need" by Vaswani et al. It replaced recurrent and convolutional approaches with a purely attention-based mechanism, enabling massive parallelization during training and capturing long-range dependencies far more effectively. Nearly every modern large language model -- GPT, BERT, LLaMA, Claude -- is built on the Transformer.
“Attention is all you need.”
-- Vaswani et al., "Attention Is All You Need" (2017, Google Brain)
LLM Visualization by Brendan Bycroft
The best interactive 3D visualization of transformer internals available. Explore how GPT-style models process tokens through embedding, attention, and feed-forward layers -- step by step, parameter by parameter. Highly recommended.
Transformer Layer Stack
Compare a Pre-Norm decoder example with the historical Post-Norm structure from the 2017 paper. These are architecture sketches; tensor dimensions, masks and exact blocks depend on the model.
Example following LLaMA (2023): RMSNorm before attention and FFN, RoPE in queries/keys, and a SwiGLU FFN. Other models can use different blocks.
Repeat the block, then apply final normalization and the LM head. Both additions preserve a direct residual path. RoPE acts on Q and K, rather than adding a position vector to x.
Architecture Variants
The original Transformer used both an encoder and decoder. Modern models often use just one. Toggle between the three main variants to see which components each uses.
Encoder-Decoder Variants
Toggle to compare architectures
Example models:
The original architecture. The encoder processes the full input bidirectionally, then the decoder generates output tokens one at a time, attending to the encoder's representations via cross-attention. Used for translation, summarization, and sequence-to-sequence tasks.
Token Dataflow
Follow a single token as it flows through the entire Transformer pipeline, from raw text to output probabilities. Watch how the tensor shape changes at each step.
Step-by-Step Dataflow
Play or scrub through the processing pipeline
Key Concepts
Residual Connections
Residual paths provide a direct route for information and gradients around sublayers and can make deep networks easier to optimize. They reduce optimization difficulties but do not alone guarantee stable training.
Layer Normalization
LayerNorm normalizes across a token’s feature dimension. The original Transformer uses Post-Norm; many decoder architectures use Pre-Norm or RMSNorm. The order changes the residual path.
Positional Encoding
Since attention has no inherent notion of order, position must be explicitly injected. The original paper used fixed sinusoidal functions; modern models typically use learned position embeddings or relative position encodings like RoPE.
Why It Matters
The Transformer architecture is arguably the most impactful innovation in AI of the past decade. It unlocked the scaling laws that make modern LLMs possible.
- 1The Transformer replaced RNNs and LSTMs by enabling full parallelization during training, reducing training time from weeks to days
- 2Its attention mechanism captures long-range dependencies that sequential models struggled with, enabling understanding of entire documents
- 3The architecture scales remarkably well -- from 100M parameter BERT to 1.8T parameter GPT-4, performance improves predictably with scale
- 4Every major LLM today (GPT, Claude, Gemini, LLaMA, Mistral) is built on the Transformer, making it the foundation of modern AI