Distillation

Intermediate

How a student model learns from a teacher’s generated examples, probability distributions or intermediate representations.

Last updated: Sep 13, 2026

What is Knowledge Distillation?

Knowledge distillation trains a student to reproduce useful behavior of a teacher. Supervision can be generated text, soft output distributions or intermediate features. A smaller student can be cheaper to serve, but quality and compression depend on the model, data and task.

"Imagine a master chef teaching an apprentice—not just the recipes, but all the subtle intuitions: why this spice almost works, why that technique is close but not quite right."

A student can learn from worked examples or richer probability targets. These are different distillation recipes, with different access requirements.

The Teacher-Student Paradigm

Distillation follows a straightforward two-phase process: first train a large, powerful teacher model, then use its outputs to train a smaller, efficient student.

🎓

Teacher Model

A model that supplies training signals. Text generation may be sufficient for sequence distillation; matching logits or hidden states requires access to those quantities.

📖

Student Model

The model being trained. It may start from a pretrained checkpoint and learn from generated examples, softened probabilities or features, depending on the objective.

Teacher Model
↓
Generated examples / soft targets
↓
Student Model Learns

A closer look at logit distillation

💡

Soft targets carry additional information

In logit distillation, the student matches a teacher distribution over the vocabulary. Besides the highest-scoring token, the relative probabilities of alternatives can provide a useful learning signal.

This is one form of distillation. Sequence distillation instead uses teacher-generated responses as training examples, often with standard supervised fine-tuning. It does not require full probability distributions.

Hard Labels (Traditional Training)

"Paris" = 1.0, everything else = 0.0

Binary: either right or wrong. No nuance. The model learns nothing about the relationships between outputs.

Soft Labels (Distillation)

"Paris" = 0.92, "Lyon" = 0.03, "Marseille" = 0.015, "Berlin" = 0.008, ...

Rich signal: every probability encodes a relationship. The student learns that Lyon is more similar to Paris than Berlin is.

One logit-distillation step

A four-token toy vocabulary and fixed teacher logits. The student has independent logits. Each click applies the gradient of T² × KL(teacher || student); this is a real update of four numbers, not a trained language model.

TokenOne-hot targetTeacherStudent
Paris1.0000.517
0.212
Lyon0.0000.244
0.273
London0.0000.148
0.350
Rome0.0000.090
0.165

Scaled distillation loss: 1.00964 · Updates: 0

Both distributions use the same temperature. One-hot labels remain one-hot; a softmax at T = 1 is still a probability distribution. This example uses only the soft loss. In practice, it can be mixed with cross-entropy on labelled tokens. Sequence distillation instead trains on teacher-generated text and does not require access to logits.

Why Distillation Works

Soft targets can add information beyond a single target token. Whether they help depends on teacher quality, student capacity, temperature and the training data.

1

Richer Gradient Signal

Each training example provides information about all output classes simultaneously, not just the correct one. This means each example effectively teaches the student about thousands of relationships at once.

2

Dark Knowledge Transfer

The teacher's "mistakes" are informative. When the teacher assigns 3% probability to "Lyon" for a question about France's capital, it tells the student that Lyon is relevant to France—knowledge that hard labels completely discard.

3

Better Generalization

Students trained via distillation often generalize better than models trained on hard labels alone, even when the student has much fewer parameters. The soft labels act as a powerful regularizer.

4

Sample Efficiency

Because each training example carries more information (a full distribution vs. a single label), the student needs fewer examples to learn effectively. This reduces training time and data requirements.

The Distillation Loss

One common logit-distillation objective combines cross-entropy on labelled tokens with KL divergence between softened teacher and student distributions:

L = (1 - α) · CE(y, pstudent) + α · T² · KL(pteacherT || pstudentT)
  • CECross-Entropy with ground truth: ensures the student still learns from real labels
  • KLKL Divergence: measures how different the student's distribution is from the teacher's. The student is penalized for deviating from the teacher's soft probabilities.
  • TTemperature: controls how soft/smooth the distributions are. Higher T reveals more inter-class relationships.
  • αAlpha: balances the two loss terms. Typical values range from 0.1 to 0.9, with higher values placing more weight on matching the teacher.

The T² factor compensates for the approximate 1/T² gradient scaling at high temperature. The mixing weight still needs to be chosen and evaluated; it does not guarantee balanced losses at every temperature.

Types of Distillation

Different approaches depending on what knowledge is transferred from teacher to student:

Response-Based

The student mimics the teacher's final output distribution. This is the original and most common form, used by Hinton et al. (2015). Simple to implement and effective for classification and language modeling.

Feature-Based

The student learns to match intermediate representations (hidden states) of the teacher, not just the output. Captures deeper structural knowledge. Used in models like DistilBERT and TinyBERT.

Relation-Based

Transfers the relationships between different examples or layers, rather than individual outputs. Preserves how the teacher structures its internal representations and how it relates different inputs to each other.

Online Distillation

Teacher and student train simultaneously, learning from each other. No pre-trained teacher required. Useful when you cannot afford to train a massive teacher model first.

Real-World Examples

Distillation is used extensively in production AI systems:

DistilBERT (Hugging Face)

The DistilBERT paper (2019) reports a model 40% smaller and 60% faster, retaining 97% of BERT’s language-understanding performance in its evaluation. These numbers belong to that study and are not a general distillation guarantee.

Text supervision versus logit access

An API that returns generated text can support sequence distillation, subject to its terms. It does not automatically expose the full logits needed for a KL-matching objective. A low price or small model name is not evidence of a particular training recipe.

DeepSeek R1 Distillation

The DeepSeek-R1 report describes fine-tuning Qwen and Llama models on about 800,000 samples curated with R1. This is sequence/data distillation through supervised fine-tuning, not evidence that those students matched full R1 logits.

Key Takeaways

  • 1Distillation transfers behavior using generated examples, soft distributions or features.
  • 2Soft targets are useful for logit distillation; they are not required by every distillation method.
  • 3Use the same temperature for teacher and student probabilities when evaluating a softened KL objective.
  • 4Compression and retained performance must be measured on specified models and tasks. There is no universal retention percentage.
  • 5Use published training reports to identify a model’s distillation method; do not infer it from size, speed or price.

Primary sources