Multi-Token Prediction (MTP)

Expert

How training a model to predict several future tokens at once can improve representation learning, planning, and inference options.

Last updated: Sep 13, 2026

What is Multi-Token Prediction?

Classic language-model training asks the model to predict only the next token. Multi-Token Prediction adds auxiliary heads or objectives for the next several tokens, so the model learns not just what comes immediately next, but how the nearby future unfolds.

Standard LM: predict x[t+1]. MTP: predict x[t+1], x[t+2], x[t+3] ... x[t+k].

Why it matters

A denser learning signal

Each training position supervises multiple future targets. That can make scarce or expensive training tokens teach more than a single next-token label.

Better future awareness

Predicting several steps ahead pressures the representation to encode structure that survives beyond the immediate next word.

New decoding possibilities

If the model already has future-token heads, systems can explore block generation, verification, and speculative-style decoding variants.

Next-token vs multi-token learning

Next-token objective

Input: The capital of France is

→ Paris

Multi-token objective

Input: The capital of France is

→ Paris · and · it · is

The model still consumes the same prefix, but the loss asks more questions about the continuation. In practice, implementations may share the trunk and attach several lightweight prediction heads, or train with shifted targets across multiple future offsets.

Tiny loss visualizer

This example sums cross-entropy over three future targets. Equal weights follow the basic multi-head example; the optional decreasing weights are an explicitly chosen alternative, not a definition of MTP.

Toy model: the sliders set the probability of the correct token at three future positions. We calculate the sum of their weighted cross-entropy terms in nats. No language model is run.

−λ log(p) = 0.357

−λ log(p) = 0.916

−λ log(p) = 1.609

L = 2.882 nats

Bar scale: 0 to −log(0.01), shared across positions. Gloeckle et al. sum head losses with equal weights. Other architectures such as DeepSeek-V3 use their own MTP modules and loss weights.

How it is usually implemented

1

Shared trunk

The transformer body produces a hidden state for each position, just like a normal language model.

2

Several future heads

Additional heads predict token t+2, t+3, and so on. Their losses are combined with the normal next-token loss.

3

Weighted objective

Loss weights and dependencies between prediction heads are architecture-specific. Compare the stated objective in each paper rather than assuming future offsets must have decreasing weights.

MTP vs speculative decoding

Speculative decoding

An inference-time trick: a smaller draft model proposes tokens and a larger model verifies them. It can speed up serving without changing how the large model was trained.

Multi-Token Prediction

A training-time objective: the model itself learns to predict several future offsets. It may enable better block-style inference, but only if the runtime uses those heads.

Benefits

  • More supervision per context token during training.
  • Representations can become less myopic and more sequence-aware.
  • Can pair naturally with block generation and verification methods.

Caveats

  • Extra heads and losses add training complexity and tuning knobs.
  • Predicting farther tokens is harder and noisier, so loss weighting matters.
  • MTP is not automatically faster at inference unless the serving stack uses it.

Where this shows up

Recent frontier-model papers and open-model reports discuss MTP-style objectives as a way to make training signals denser and improve downstream generation behavior. The exact implementation details differ, so treat MTP as a family of objectives rather than one fixed architecture.

Key Takeaways

  • 1MTP extends the next-token objective to several future offsets.
  • 2The main benefit is a richer training signal and less short-sighted representations.
  • 3It can enable block-style generation ideas, but only with matching inference machinery.
  • 4Think of MTP as training the model to glance down the road, not just at the next footstep.