Next-Token Prediction

Beginner

What an LLM actually outputs: logits over the vocabulary, decoded one token at a time into text.

Last updated: Sep 13, 2026

From score to next token

Two steps with five readable example tokens and hand-set logits each. Probabilities are calculated here. No trained language model runs; the second distribution is also supplied.

Context so far

The cat sleeps on the ▌

Softmax: exp(logit / T), divided by the sum. Top-p retains the most likely tokens until at least p is reached, then renormalizes.

Step 1: logits → softmax → top-p
TokenLogitSoftmaxTop-p
mat3.252.1%57.4%
blanket2.525.9%28.5%
lawn1.812.8%14.1%
staircase1.16.4%0%
couch0.32.9%0%

Draw from one cumulative distribution

u is a number between 0 and 1. The interval containing u determines the token. Keep the same draw for comparisons or draw again.

u
01

Selection: blanket · Interval (rounded endpoints) ≈ [0.574, 0.859)

Logits over the vocabulary

At the end of the network, the final hidden vector is projected to one score for every token in the vocabulary. These raw scores are called logits.

  • A logit is not a word; it is a score for a token ID.
  • Softmax turns logits into a probability distribution over possible next tokens.

Decoding choices

Generation depends on how the model chooses from those probabilities. Greedy decoding picks the top token, while sampling methods introduce controlled randomness.

  • Temperature reshapes the probability distribution.
  • Top-k keeps only the k most likely tokens.
  • Top-p keeps the smallest set whose probability mass reaches p.

The generation loop

The chosen token is appended to the context. The model runs again, predicts the next token, appends it, and repeats until it stops.

  • Training can predict many sequence positions in parallel with causal masking.
  • Inference generation is sequential because each new token depends on the previous one.
  • High probability means likely continuation, not guaranteed truth.

Key Takeaways

  • LLMs output next-token probabilities, not direct truth claims.
  • Decoding strategy changes style, determinism, and error profile.
  • Text is produced by repeating predict → choose → append.