1Purpose first
How a token gets chosen.
At each step a model produces a probability for every possible next token. Sampling parameters decide how one token is chosen from that list. They change how predictable or varied the reply is, without changing the model.
2The settings at a glance
Six knobs, one line each.
| Setting | What it does | Effect |
|---|---|---|
| max tokens | Caps the length of the reply | Stops it running on; cuts it off if too small |
| Greedy | Always take the most probable token | Same input, same output; can be repetitive |
| Random sampling | Draw a token by its probability | Varied output |
| Temperature | Sharpens or flattens the probabilities | Low is focused, high is more varied |
| Top-k | Only consider the k most probable tokens | Cuts the long tail |
| Top-p | Only consider the smallest set adding up to p | Adapts to how confident the model is |
3See it work
Temperature, top-k and top-p on real numbers.
import math
logits = {"Paris": 4.0, "Lyon": 2.5, "Rome": 2.0, "banana": -1.0}
def softmax(d, T=1.0):
ex = {k: math.exp(v / T) for k, v in d.items()}
s = sum(ex.values())
return {k: round(v / s, 3) for k, v in ex.items()}
for T in (0.5, 1.0, 2.0):
print("temperature", T, softmax(logits, T))
p = softmax(logits)
ranked = sorted(p.items(), key=lambda kv: -kv[1])
print("greedy:", ranked[0][0])
print("top-k (k=2):", [w for w, _ in ranked[:2]])
kept, total = [], 0
for w, prob in ranked:
kept.append(w); total += prob
if total >= 0.9: break
print("top-p (p=0.9):", kept)temperature 0.5 {'Paris': 0.936, 'Lyon': 0.047, 'Rome': 0.017, 'banana': 0.0}
temperature 1.0 {'Paris': 0.732, 'Lyon': 0.163, 'Rome': 0.099, 'banana': 0.005}
temperature 2.0 {'Paris': 0.52, 'Lyon': 0.246, 'Rome': 0.191, 'banana': 0.043}
greedy: Paris
top-k (k=2): ['Paris', 'Lyon']
top-p (p=0.9): ['Paris', 'Lyon', 'Rome']
Read the output line by line. At temperature 0.5 "Paris" takes 94% of the probability, at 2.0 only 52%, and even the absurd "banana" gets 4%. Top-k keeps a fixed count of candidates. Top-p keeps as many as it takes to reach 90%, here three, so it keeps more options when the model is unsure and fewer when it is sure.
4Temperature
The one people reach for first.
Temperature divides the model's raw scores before they become probabilities. Below 1 it concentrates probability on the leaders, and above 1 it spreads it out. Close to 0 behaves like greedy.
5What changed on newer models
Do not assume the knobs exist.
temperature, top_p or top_k with an error, and the advice is to leave them out and steer through the prompt. Reasoning-style models from other providers have similar limits.So the habit of "set temperature to 0 for reliability" is dated for those models. What replaces it:
- Say exactly what you want in the prompt and show an example.
- Use enforced structured output (tutorial 4) for shape.
- Validate the result in code and retry if it fails (a later tutorial covers retries).
Check the documentation for the exact model you use before passing any sampling setting.
6Check yourself
Short questions and one line to remember.
Which gives repeatable output, greedy or sampling?
Greedy always picks the top token, so it is the most repeatable, though some providers still vary slightly.
How does top-p differ from top-k?
Top-k keeps a fixed number of candidates. Top-p keeps however many are needed to reach a probability total, so it adapts.
Should you always set temperature to 0 for factual work?
Not on newer models that reject the setting. Use clear prompts, structured output and validation instead.