From score to next token
Two steps with five readable example tokens and hand-set logits each. Probabilities are calculated here. No trained language model runs; the second distribution is also supplied.
Context so far
The cat sleeps on the ▌
Softmax: exp(logit / T), divided by the sum. Top-p retains the most likely tokens until at least p is reached, then renormalizes.
| Token | Logit | Softmax | Top-p |
|---|---|---|---|
| mat | 3.2 | 52.1% | 57.4% |
| blanket | 2.5 | 25.9% | 28.5% |
| lawn | 1.8 | 12.8% | 14.1% |
| staircase | 1.1 | 6.4% | 0% |
| couch | 0.3 | 2.9% | 0% |
Draw from one cumulative distribution
u is a number between 0 and 1. The interval containing u determines the token. Keep the same draw for comparisons or draw again.
Selection: blanket · Interval (rounded endpoints) ≈ [0.574, 0.859)
Logits over the vocabulary
At the end of the network, the final hidden vector is projected to one score for every token in the vocabulary. These raw scores are called logits.
- A logit is not a word; it is a score for a token ID.
- Softmax turns logits into a probability distribution over possible next tokens.
Decoding choices
Generation depends on how the model chooses from those probabilities. Greedy decoding picks the top token, while sampling methods introduce controlled randomness.
- Temperature reshapes the probability distribution.
- Top-k keeps only the k most likely tokens.
- Top-p keeps the smallest set whose probability mass reaches p.
The generation loop
The chosen token is appended to the context. The model runs again, predicts the next token, appends it, and repeats until it stops.
- Training can predict many sequence positions in parallel with causal masking.
- Inference generation is sequential because each new token depends on the previous one.
- High probability means likely continuation, not guaranteed truth.
Key Takeaways
- LLMs output next-token probabilities, not direct truth claims.
- Decoding strategy changes style, determinism, and error profile.
- Text is produced by repeating predict → choose → append.