LLM Training

Intermediate

How large language models are trained: from pretraining to RLHF.

Last updated: Sep 13, 2026

How LLMs Are Trained

Large language models go through multiple training stages, each with different objectives and techniques. Understanding this pipeline is crucial for understanding model capabilities and limitations.

The training process fundamentally shapes what LLMs can and cannot do. Different training approaches produce models with different strengths, weaknesses, and behaviors.

Interactive Training Pipeline

Explore the decisions and outputs of each stage. This overview does not assign universal cost, duration or GPU counts.

One possible training pipeline

Select a stage. Projects do not all use every stage or the same order. Cost and duration require a concrete model, data budget and hardware configuration.

Collect data

Record provenance, version, licences and intended use. Public accessibility alone is not permission for every use.

Output: an auditable source inventory.

Complete Training Pipeline (8 Stages)

Modern LLM training involves 8 major stages, each with different objectives, data requirements, and compute costs. Understanding this pipeline is essential for grasping the complexity and expense of training frontier models.

1

Data Collection & Curation

Gather massive text corpora from diverse sources including web crawls, books, code repositories, and scientific papers.

2

Data Cleaning & Deduplication

Remove duplicates, filter low-quality content, detect languages, remove PII, and normalize formatting.

3

Tokenization

Convert cleaned text into numerical token sequences using BPE, SentencePiece, or Unigram tokenizers.

4

Pre-training

Train the base model on trillions of tokens using next-token prediction objective. This is the most expensive and compute-intensive stage.

5

Supervised Fine-Tuning (SFT)

Fine-tune the base model on curated instruction-response pairs to teach it to follow instructions and respond helpfully.

6

RLHF / Preference Tuning

Align model outputs with human preferences using reinforcement learning (PPO) or direct preference optimization (DPO).

7

Safety & Evaluation

Red-team the model, run adversarial tests, apply Constitutional AI principles, and benchmark on standard evaluation suites.

8

Deployment Optimization

Optimize the model for production deployment through quantization, distillation, and inference infrastructure setup.

PPO, DPO and GRPO

Separate the policy objective from the source of rewards. PPO and GRPO optimize sampled responses; DPO can optimize a fixed dataset of preferred and rejected response pairs.

RLHF: The Traditional Approach

RLHF uses a separate reward model trained on human preferences, then optimizes the LLM using reinforcement learning (typically PPO) to maximize that reward.

1. Step 1: Collect Preferences

2. Step 2: Train Reward Model

3. Step 3: RL Optimization

DPO: The Simplified Alternative

DPO skips the reward model entirely, directly optimizing the LLM on preference data using a clever mathematical reformulation.

1. Step 1: Collect Preferences

2. Step 2: Direct Optimization

3. Step 3: No RL Required

GRPO: Group Relative Policy Optimization

GRPO estimates an advantage relative to a group of sampled responses. It avoids a separate value/critic model. Rewards can come from rules or a learned reward model; DeepSeekMath used learned reward models.

1. Generate Response Group

2. Score the response group

3. Relative Gradient Update

PropertyPPODPOGRPO
Learning signalReward per rolloutPreferred / rejected pairsRewards for a rollout group
Separate criticUsually yesNoNo; group-based baseline
Reward modelPossible, task-dependentNo separately trained modelPossible; rules can also provide rewards
WorkloadRollouts and policy/value updatesTraining on fixed pairs in the offline recipeGrouped rollouts and policy updates

Key Takeaways

  • 1LLM training has distinct stages: pretraining → SFT → RLHF → specialized alignment
  • 2The RL paradigm (e.g., DeepSeek R1-Zero) shows reasoning can emerge from pure RL without human demonstrations
  • 3RLHF aligns models with human preferences; pure RL optimizes for verifiable outcomes
  • 4Modern models often combine multiple techniques: SFT for instruction following, RLHF for preferences, RL for reasoning
  • 5Understanding the training pipeline helps you understand model behavior and limitations
  • 6The field is rapidly evolving—new paradigms like pure RL are changing how we think about training

Primary sources