LLM temperature is the setting people change most and understand least. Mechanically it is one division, applied to the model's scores before it picks the next token, and almost everything people say about it follows from that.
The one line of math
At every step, a language model produces a score (a logit) for every token in its vocabulary. Softmax turns those scores into probabilities. Temperature divides the scores first:
p(token) = softmax(logits / T)
- T = 1 leaves the distribution as the model learned it.
- T < 1 stretches the gaps between scores. Likely tokens get likelier, unlikely ones fade.
- T > 1 squeezes the gaps. The distribution flattens and rare tokens get a real chance.
- As T approaches 0, all all the probability goes to the top token. That's greedy decoding.
A small example makes it concrete. Suppose three candidate tokens have logits 2.0, 1.0 and 0.0.
| Temperature | p(top) | p(second) | p(third) |
|---|---|---|---|
| 0.5 | 0.87 | 0.12 | 0.02 |
| 1.0 | 0.67 | 0.24 | 0.09 |
| 1.5 | 0.56 | 0.29 | 0.15 |
Computed from the formula above; rounded.
The model is the same in every row; only the sampling changed.
What it does to a whole answer
A single token's probabilities shift a little. Over a few hundred tokens, the effect compounds. At low temperature the model keeps taking the most expected path, so repeated runs look alike and phrasing turns generic. At high temperature each step has a small chance of an odd choice, and once the text takes an odd turn, every later token is conditioned on it. That's why very high temperatures drift into incoherence instead of producing steadily more creative text.
Temperature doesn't work alone
Most servers also truncate the distribution: top-k keeps the k most likely tokens, top-p (nucleus sampling) keeps the smallest set whose probabilities add up to p, and min-p keeps tokens whose probability is at least a fraction of the top token's. Min-p was proposed specifically to keep text coherent at high temperatures, because it cuts the long tail harder when the model is confident and relaxes when it isn't (Nguyen et al., arXiv 2407.01082).
So "temperature 1.2" describes several different samplers. Temperature 1.2 with top-p 0.9 behaves very differently from temperature 1.2 with no truncation. When you compare runs, record all of the sampling parameters, not just temperature.
Temperature 0 isn't a promise
Setting temperature to 0 makes decoding greedy. It doesn't make an API deterministic. Servers batch requests together, and the floating-point results of matrix multiplications and normalizations can differ depending on batch size. Thinking Machines showed this clearly and built batch-invariant kernels that made 1,000 runs of the same prompt identical, which also shows how much engineering determinism takes (Defeating Nondeterminism in LLM Inference). If you need reproducibility, store the outputs you depend on instead of counting on regenerating them.
Greedy decoding also has a quality cost for some models. Qwen's model cards, for example, warn against greedy decoding for Qwen3 because it can cause endless repetition, and they recommend temperature 0.6 in thinking mode and 0.7 without it (Qwen3-30B-A3B model card). The vendor's recommended settings are a better starting point than 0.
How I choose
- Extraction, classification, anything checked by code: low temperature (0 to 0.3), because I want the most likely answer and variety is noise.
- Drafts a person will pick from: the model's recommended default, several samples, and a selection step. Variety is useful when something chooses among the results.
- Brainstorming: higher temperature with min-p or top-p truncation, and often thinking turned off (I wrote about that separately).
Try it yourself
Take one prompt that asks for a short answer, for example "Name a sailing term that would make a good band name." Run it 20 times at each of three temperatures, say 0.2, 0.8 and 1.4, with the same top-p. Count distinct answers and read the strangest ones. The counts will depend on the model and server, so I won't give you mine; the shape is what to look for: near-identical answers at the low end, a useful spread in the middle, and a growing share of odd ones at the top.
The practical next step
Map one real execution and one failure.
That will reveal more about the right architecture than a tool comparison or model demo.
Let's build something real