In short

At every generated word, an LLM does not choose “the right word” — it draws at random from candidates, according to a probability distribution it has computed. Temperature and sampling parameters (top-k, top-p) control how this random draw is organized.

Low temperature: the model plays it safe, almost always choosing the most probable token. High temperature: it takes risks, explores less obvious choices. The result may be brilliant or incoherent.

This is neither magic nor raw randomness. It is applied statistics, and understanding the mechanism allows you to calibrate a model’s behavior to the task at hand.


Explanation

What the model computes at each step

An LLM generates text token by token. At each step, it produces a vector of logits — raw scores, one per token in its vocabulary (often between 50,000 and 130,000 tokens). These raw scores are not yet probabilities.

The softmax function converts these logits into probabilities: it exponentiates each score, then divides by the sum, so that the distribution sums to 1. The token with the highest logit gets the highest probability, but the others are not zero.

This is where temperature comes in.

Temperature: divide before normalizing

Temperature T is a positive number applied before softmax: each logit is divided by T, then normalization is applied normally.

  • T close to 0: dividing by a very small number amplifies the gaps between logits. The most probable token dominates all others — the distribution becomes nearly deterministic. In practice, at T=0, the model almost always chooses the same token (with slight variations due to floating-point numerical approximations).

  • T = 1: logits are not modified. The distribution reflects exactly what the model learned during training. This is the baseline behavior.

  • T > 1: dividing by a large number brings logits closer together. The distribution flattens — less probable tokens gain relative weight. Diversity increases, but so does the risk of producing something incoherent.

A useful metaphor: imagine a list of candidates ranked by relevance. At low temperature, only the top candidate matters. At high temperature, candidates at the bottom of the list have almost as much chance of being chosen as the first.

In short: temperature does not change what the model “knows”. It changes how it chooses among what it knows. At T=0, it always takes its first option. At T=2, it nearly draws lots between all its options, regardless of their merit. Neither is universally better — it all depends on what you want to obtain.

Top-k: narrowing the field of candidates

Top-k sampling is a filtering operation applied after probability computation. Only the k most probable tokens are kept; all others have their probability set to zero. The distribution is renormalized, then a draw is made from this reduced subset.

Typical values: k=10 to 20 to stay focused on the topic, k=50 to 100 to allow more variety.

Limit of this approach: k is fixed, regardless of the shape of the distribution. If a single token dominates with 95% probability, a top-k=50 still retains 49 very low-probability tokens. The selection can then be distorted by marginal candidates.

Top-p (nucleus sampling): an adaptive filter

Top-p sampling, also called nucleus sampling, resolves the top-k limitation elegantly. Rather than fixing a number of tokens, you fix a cumulative probability mass p. The smallest set of tokens whose probabilities sum to at least p is selected, the distribution is renormalized, and a draw is made from this set.

Concrete example: with p=0.9, if the model is very confident (one token at 85% probability), the nucleus contains only 2 or 3 tokens. If the model hesitates between ten close options, the nucleus expands to include them all.

This is a dynamic filter: it adapts at each generation step to the actual state of the distribution, which is more robust than top-k in most situations.

In short: top-k = “keep the K best options, no matter their actual scores”. Top-p = “keep the smallest set of options that cover X% of the probability mass”. Top-p adapts automatically to the situation. That is why it became the standard.

Softmax: the formula behind it all

For a vocabulary of N tokens with logits $z_1, z_2, \ldots, z_N$, the probability of token $i$ after applying temperature T is:

$$P(i) = \frac{e^{z_i / T}}{\sum_{j=1}^{N} e^{z_j / T}}$$

Dividing all logits by T before exponentiation compresses or stretches relative gaps, which changes the shape of the distribution without altering the ranking order of tokens.


Concrete examples

Code generation — low temperature

For code generation, you want predictable, syntactically correct responses. A temperature between 0.2 and 0.4, combined with a conservative top-p (0.85–0.9), keeps the model in the high-probability zones of its prediction space.

Concretely: if the model generates a Python function, the token following def will almost certainly be a space, then a valid function name. At low temperature, the model “knows” this is expected and sticks to it.

Creative writing — high temperature

For narrative text, a novel opening, or advertising slogans, a temperature between 0.7 and 1.0 produces less expected formulations, less common word associations. The model takes paths it would not have taken in safe mode.

Caution: above 1.2, narrative coherence can degrade. Low-probability tokens become too prevalent, and the text may lose its thread.

Brainstorming and factual questions — the extremes

For factual responses (the capital of a country, the definition of a technical term), T close to 0 is recommended: there are not multiple correct answers, and diversity adds no value.

For brainstorming, a high temperature (0.8–1.2) combined with a high top-k or top-p allows exploring less obvious associations. This is where the model can positively surprise you.

Summary table

Use caseTemperatureSampling
Code generation0.2 – 0.4top-p 0.85–0.9
Factual questions0.0 – 0.3low top-k
Creative writing0.7 – 1.0top-p 0.9–0.95
Brainstorming0.8 – 1.2high top-k or top-p

A practical recommendation from practitioners: use top-p or top-k in combination with temperature, but not both simultaneously. Combining top-k and top-p over-constrains the distribution with no demonstrated benefit.

Decision matrix by use case

ContextRecommendationWhy
Code generation, structured JSON, deterministic outputT = 0.0–0.3, top-p ≥ 0.9Eliminates variance, guarantees syntactic compliance. Cost: no variation between runs.
Factual questions (capitals, definitions, formulas)T = 0.0, top-k = 1One correct answer, always take the same. Diversity has no value.
Descriptive writing, summary, translationT = 0.5–0.7, top-p 0.9Fluidity/coherence trade-off, avoids repetition without going off the rails.
Creative writing (storytelling, slogans, dialogue)T = 0.8–1.0, top-p 0.9–0.95Lexical surprise sought, the model dares unusual associations.
Brainstorming, idea exploration (generating N variants)T = 1.0–1.2, top-p 0.95Maximizes inter-run diversity. Run multiple draws and sort afterwards.

Key takeaways

  • Temperature is not a “creativity” dial in any vague sense. It is a precise mathematical parameter that modifies the shape of a probability distribution before the model draws its next token.
  • Top-k and top-p are complementary filters: they restrict the sampling space to plausible candidates. Top-p is generally preferred because it dynamically adapts to each situation.
  • These parameters have no universally correct value. They depend on the task, the model, and sometimes the expected content style. Understanding what they do mechanically allows you to tune them in an informed way rather than iterating blindly.