How LLMs Are Trained
Large language models go through multiple training stages, each with different objectives and techniques. Understanding this pipeline is crucial for understanding model capabilities and limitations.
The training process fundamentally shapes what LLMs can and cannot do. Different training approaches produce models with different strengths, weaknesses, and behaviors.
Interactive Training Pipeline
Explore the decisions and outputs of each stage. This overview does not assign universal cost, duration or GPU counts.
One possible training pipeline
Select a stage. Projects do not all use every stage or the same order. Cost and duration require a concrete model, data budget and hardware configuration.
Collect data
Record provenance, version, licences and intended use. Public accessibility alone is not permission for every use.
Output: an auditable source inventory.
Complete Training Pipeline (8 Stages)
Modern LLM training involves 8 major stages, each with different objectives, data requirements, and compute costs. Understanding this pipeline is essential for grasping the complexity and expense of training frontier models.
Data Collection & Curation
Gather massive text corpora from diverse sources including web crawls, books, code repositories, and scientific papers.
Data Cleaning & Deduplication
Remove duplicates, filter low-quality content, detect languages, remove PII, and normalize formatting.
Tokenization
Convert cleaned text into numerical token sequences using BPE, SentencePiece, or Unigram tokenizers.
Pre-training
Train the base model on trillions of tokens using next-token prediction objective. This is the most expensive and compute-intensive stage.
Supervised Fine-Tuning (SFT)
Fine-tune the base model on curated instruction-response pairs to teach it to follow instructions and respond helpfully.
RLHF / Preference Tuning
Align model outputs with human preferences using reinforcement learning (PPO) or direct preference optimization (DPO).
Safety & Evaluation
Red-team the model, run adversarial tests, apply Constitutional AI principles, and benchmark on standard evaluation suites.
Deployment Optimization
Optimize the model for production deployment through quantization, distillation, and inference infrastructure setup.
PPO, DPO and GRPO
Separate the policy objective from the source of rewards. PPO and GRPO optimize sampled responses; DPO can optimize a fixed dataset of preferred and rejected response pairs.
RLHF: The Traditional Approach
RLHF uses a separate reward model trained on human preferences, then optimizes the LLM using reinforcement learning (typically PPO) to maximize that reward.
1. Step 1: Collect Preferences
2. Step 2: Train Reward Model
3. Step 3: RL Optimization
DPO: The Simplified Alternative
DPO skips the reward model entirely, directly optimizing the LLM on preference data using a clever mathematical reformulation.
1. Step 1: Collect Preferences
2. Step 2: Direct Optimization
3. Step 3: No RL Required
GRPO: Group Relative Policy Optimization
GRPO estimates an advantage relative to a group of sampled responses. It avoids a separate value/critic model. Rewards can come from rules or a learned reward model; DeepSeekMath used learned reward models.
1. Generate Response Group
2. Score the response group
3. Relative Gradient Update
| Property | PPO | DPO | GRPO |
|---|---|---|---|
| Learning signal | Reward per rollout | Preferred / rejected pairs | Rewards for a rollout group |
| Separate critic | Usually yes | No | No; group-based baseline |
| Reward model | Possible, task-dependent | No separately trained model | Possible; rules can also provide rewards |
| Workload | Rollouts and policy/value updates | Training on fixed pairs in the offline recipe | Grouped rollouts and policy updates |
Key Takeaways
- 1LLM training has distinct stages: pretraining → SFT → RLHF → specialized alignment
- 2The RL paradigm (e.g., DeepSeek R1-Zero) shows reasoning can emerge from pure RL without human demonstrations
- 3RLHF aligns models with human preferences; pure RL optimizes for verifiable outcomes
- 4Modern models often combine multiple techniques: SFT for instruction following, RLHF for preferences, RL for reasoning
- 5Understanding the training pipeline helps you understand model behavior and limitations
- 6The field is rapidly evolving—new paradigms like pure RL are changing how we think about training