What is Temperature?
In LLMs, Temperature is a hyperparameter that scales the "logits" (raw scores) of the next token predictions before they are converted into probabilities. It essentially controls how much the model favors the most likely options versus exploring less likely ones.
Low Temperature
Concentrates probability on the highest logits. This makes sampling less varied, but does not make incorrect information true.
High Temperature
Flattens the distribution and gives lower-logit tokens more probability. The effect on task quality depends on the model and task.
Interactive Distribution
Adjust the temperature to see its effect
Compare the calculated distribution with actual draws from five example logits. Changing temperature clears the old sample counts.
Five example tokens with fixed logits [3, 2, 1, 0, -1]. Samples come from exactly the distribution shown here. No generated text or quality score is implied.
Purple: calculated probability. Cyan: observed sample frequency. Finite samples fluctuate. T = 0 uses greedy decoding.
How it Works Mathematically
The model generates a score for every possible token. To get probabilities, we use the Softmax function, modified by temperature:
Dividing by a small T amplifies differences between scores. The highest logit dominates exponentially.
Dividing by a large T compresses all scores toward zero, making them nearly equal after exponentiation.
Practical Guidelines
Start with guidance for the specific model and mode. Then compare settings on the same tasks and reference answers over multiple runs. Facts need sources and code needs tests; temperature does not replace these checks.
For example, Qwen3 recommends T = 0.6 for thinking and warns against greedy decoding. A universal recipe such as "always use T = 0 for math" would not fit this model.
Qwen3-32B: Best practicesKey Takeaways
- 1T = 0 commonly selects greedy decoding. Repeated API calls can still differ because other parts of inference may vary.
- 2Temperature changes the distribution, not the model’s stored knowledge.
- 3No universal temperature threshold guarantees truth, creativity or incoherence.
- 4Compare task results and follow model-specific guidance, especially for reasoning modes.