← Home · Tutorial 05 · Working with LLMs

Sampling parameters

Max tokens, greedy versus random, top-k, top-p and temperature, and what newer models changed.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

How I checked this: the mechanics are standard, and the code is my own working version. The note about newer models is from provider documentation and developer reports at the time of writing (see the footer). Treat that part as the most likely to change.
Who this is for: anyone who has seen temperature and top-p in an API and not known what to set. You will learn what the settings do, shown with runnable examples, and what changed on newer models.

1Purpose first

How a token gets chosen.

At each step a model produces a probability for every possible next token. Sampling parameters decide how one token is chosen from that list. They change how predictable or varied the reply is, without changing the model.

2The settings at a glance

Six knobs, one line each.

SettingWhat it doesEffect
max tokensCaps the length of the replyStops it running on; cuts it off if too small
GreedyAlways take the most probable tokenSame input, same output; can be repetitive
Random samplingDraw a token by its probabilityVaried output
TemperatureSharpens or flattens the probabilitiesLow is focused, high is more varied
Top-kOnly consider the k most probable tokensCuts the long tail
Top-pOnly consider the smallest set adding up to pAdapts to how confident the model is

3See it work

Temperature, top-k and top-p on real numbers.

Python
import math
logits = {"Paris": 4.0, "Lyon": 2.5, "Rome": 2.0, "banana": -1.0}

def softmax(d, T=1.0):
    ex = {k: math.exp(v / T) for k, v in d.items()}
    s = sum(ex.values())
    return {k: round(v / s, 3) for k, v in ex.items()}

for T in (0.5, 1.0, 2.0):
    print("temperature", T, softmax(logits, T))

p = softmax(logits)
ranked = sorted(p.items(), key=lambda kv: -kv[1])
print("greedy:", ranked[0][0])
print("top-k (k=2):", [w for w, _ in ranked[:2]])

kept, total = [], 0
for w, prob in ranked:
    kept.append(w); total += prob
    if total >= 0.9: break
print("top-p (p=0.9):", kept)
Output (run to check)
temperature 0.5 {'Paris': 0.936, 'Lyon': 0.047, 'Rome': 0.017, 'banana': 0.0}
temperature 1.0 {'Paris': 0.732, 'Lyon': 0.163, 'Rome': 0.099, 'banana': 0.005}
temperature 2.0 {'Paris': 0.52, 'Lyon': 0.246, 'Rome': 0.191, 'banana': 0.043}
greedy: Paris
top-k (k=2): ['Paris', 'Lyon']
top-p (p=0.9): ['Paris', 'Lyon', 'Rome']

Read the output line by line. At temperature 0.5 "Paris" takes 94% of the probability, at 2.0 only 52%, and even the absurd "banana" gets 4%. Top-k keeps a fixed count of candidates. Top-p keeps as many as it takes to reach 90%, here three, so it keeps more options when the model is unsure and fewer when it is sure.

4Temperature

The one people reach for first.

Temperature divides the model's raw scores before they become probabilities. Below 1 it concentrates probability on the leaders, and above 1 it spreads it out. Close to 0 behaves like greedy.

When it fits: low for extraction, classification and anything you want repeatable. Higher for brainstorming and drafting where variety helps.
Low is not a guarantee. Even at the lowest setting, outputs can still differ slightly between runs on some providers, so do not build on exact repeatability.

5What changed on newer models

Do not assume the knobs exist.

The catch on newer models. These knobs are going away on some of them. Anthropic's newer models, from Opus 4.7 onward according to developer reports and migration notes, reject requests that set non-default temperature, top_p or top_k with an error, and the advice is to leave them out and steer through the prompt. Reasoning-style models from other providers have similar limits.

So the habit of "set temperature to 0 for reliability" is dated for those models. What replaces it:

  • Say exactly what you want in the prompt and show an example.
  • Use enforced structured output (tutorial 4) for shape.
  • Validate the result in code and retry if it fails (a later tutorial covers retries).

Check the documentation for the exact model you use before passing any sampling setting.

6Check yourself

Short questions and one line to remember.

Which gives repeatable output, greedy or sampling?

Greedy always picks the top token, so it is the most repeatable, though some providers still vary slightly.

How does top-p differ from top-k?

Top-k keeps a fixed number of candidates. Top-p keeps however many are needed to reach a probability total, so it adapts.

Should you always set temperature to 0 for factual work?

Not on newer models that reject the setting. Use clear prompts, structured output and validation instead.

One line to remember: "Sampling settings shape how a token is picked, but on newer models the prompt and the checks around it do that job."
Newer-model note based on Anthropic’s Messages API reference and developer reports of the change (for example this write-up); I could not confirm every model’s status in the official pages, so check the docs for your model. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.