← Home · Tutorial 08 · Building with APIs

Conversation state and the context window

Why a chatbot has no memory, and how to drop, summarise or cache your way past the limit.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

Where this comes from: my own handwritten notes on conversation state and the context window. I checked the behaviour and figures against Anthropic's context windows and prompt caching documentation as read in October 2026. Numbers such as window sizes and prices change often, so the ideas here matter more than the figures.
Who this is for: anyone building a chatbot or wondering why a model forgets. You will learn why models are stateless, how the context window works, and how to trim, summarise and cache.

1Purpose first

Why chatbots seem to remember.

A chatbot feels like it remembers you. It does not. Understanding why is the most useful single idea for building anything conversational, because every cost and limit follows from it.

2Stateless by design

The conversation is a list you re-send.

The model API is stateless. Each call is independent. The "conversation" is the list of messages you send back every time, growing a turn at a time.

Each turn re-sends everything
Turn 1you + reply
→
Turn 2turn 1 + new message
→
Turn 3turns 1 and 2 + new message

So memory is something you manage in code, using the same list-of-messages shape from tutorial 2. The model sees only what you put in the list.

3A model is not a chatbot

What the application adds around the model.

Your notes make a useful distinction: a language model and a chatbot are not the same thing. The model is a function: tokens go in, the next token comes out, and it forgets everything once it has answered. Training makes each single reply helpful and safe, but nothing carries over to the next turn. The application around it does everything else.

LayerJob
The modelOne helpful, safe reply to one input. No memory before or after
MemoryRe-sending the conversation on every call
PersistenceA database, so chats survive past one session
The productInterface, sign-in, billing, streaming, moderation, tool calls
Context managementDeciding what to keep when the history outgrows the window

The model is the engine. Everything that makes it feel like a product is the car built around it. That is why most of the work in a chat application is not the model call.

4The context window

The limit on one request.

The context window is the most the model can handle in one request: your system prompt, every message so far, tool definitions and results, and the reply it is writing. Per Anthropic's documentation, recent Claude models have 1 million tokens, with older ones at 200,000.

Bigger is not free. Everything you send is processed and billed on every call, and the documentation notes accuracy can degrade as context grows, sometimes called context rot. A long window is room, not a reason to fill it.

5Three strategies

Drop, summarise or pick.

Three ways to stay within the limit, from simplest to most careful:

StrategyHowTrade-off
DropRemove the oldest messagesSimple, but forgets early details
SummariseReplace old turns with a short summary, written by a modelKeeps the gist, costs a call, may lose detail
PickKeep only the messages relevant to the new questionEfficient, needs a way to judge relevance

6Trimming in code

The simplest version, run and checked.

Python
def est_tokens(msgs):                        # rough: about 4 characters per token
    return sum(len(m["content"]) for m in msgs) // 4

def trim(msgs, limit):
    msgs = list(msgs)
    while est_tokens(msgs) > limit and len(msgs) > 1:
        msgs.pop(0)                          # drop the oldest message
    while msgs and msgs[0]["role"] != "user":
        msgs.pop(0)                          # the list must start with a user turn
    return msgs

history = [
    {"role": "user",      "content": "a" * 40},
    {"role": "assistant", "content": "b" * 40},
    {"role": "user",      "content": "c" * 40},
    {"role": "assistant", "content": "d" * 40},
    {"role": "user",      "content": "e" * 40},
]
print("before:", est_tokens(history), "tokens, roles", [m["role"][0] for m in history])
kept = trim(history, 25)
print("after: ", est_tokens(kept), "tokens, roles", [m["role"][0] for m in kept])
Output (run to check)
before: 50 tokens, roles ['u', 'a', 'u', 'a', 'u']
after:  10 tokens, roles ['u']

The 50-token history is over a 25-token limit, so the oldest messages go one by one until it fits. Two things to notice. It stopped dropping while still having two messages, but one was an assistant turn, so that was removed too: the list has to start with a user message. And only the latest message survived, which is why dropping alone loses context fast.

7Prompt caching

Stop paying for the same prefix twice.

Re-sending a long, unchanged prefix, such as a big system prompt or a document, every turn is wasteful. Prompt caching lets the provider reuse the processing for a prefix that has not changed. You mark where the stable part ends, and later requests with an identical prefix are charged much less for that part.

Shape taken from the docs, not run here
response = client.messages.create(
    model="MODEL_ID_FROM_THE_DOCS",
    max_tokens=1024,
    cache_control={"type": "ephemeral"},   # as shown in Anthropic's documentation
    system=LONG_SYSTEM_PROMPT,
    messages=history,
)

From the documentation at the time of writing: a cache write costs about 1.25 times the normal input price for the default 5-minute lifetime, and a cache hit about a tenth of it. Prompts below a minimum length, which depends on the model, are silently not cached, and changing the tools or system prompt invalidates the cache.

Design tip: keep stable content (instructions, reference documents) at the start and changing content (the user's latest message) at the end. Anything that changes early in the prompt breaks the cache for everything after it.

8Let the platform help

Compaction and context editing.

Some platforms now offer to shorten long conversations for you. Anthropic describes server-side compaction (in beta at the time of writing) that summarises earlier turns automatically, and context editing that clears old tool results. These are the "summarise" strategy done by the platform. They are convenient, but they still lose detail, so test what your application needs to keep.

9Check yourself

Short questions and one line to remember.

Why does a chatbot seem to have no memory?

The API is stateless. Memory is the message list you re-send, so the model only knows what you include.

What counts toward the context window?

Everything: system prompt, all messages, tool definitions and results, and the reply being written.

What invalidates a prompt cache?

Changing an earlier part of the prompt, such as the tools or system prompt. Anything after the change is reprocessed.

Drop or summarise: when do you pick each?

Drop when old turns truly do not matter. Summarise when the gist of earlier turns does.

One line to remember: "The model remembers nothing, so the conversation is the list you send, and cost and quality follow from how you shape it."
Figures from Anthropic’s context windows and prompt caching pages, as read in October 2026. Check the current pages for model-specific numbers. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.