1Purpose first
Why chatbots seem to remember.
A chatbot feels like it remembers you. It does not. Understanding why is the most useful single idea for building anything conversational, because every cost and limit follows from it.
2Stateless by design
The conversation is a list you re-send.
The model API is stateless. Each call is independent. The "conversation" is the list of messages you send back every time, growing a turn at a time.
So memory is something you manage in code, using the same list-of-messages shape from tutorial 2. The model sees only what you put in the list.
3A model is not a chatbot
What the application adds around the model.
Your notes make a useful distinction: a language model and a chatbot are not the same thing. The model is a function: tokens go in, the next token comes out, and it forgets everything once it has answered. Training makes each single reply helpful and safe, but nothing carries over to the next turn. The application around it does everything else.
| Layer | Job |
|---|---|
| The model | One helpful, safe reply to one input. No memory before or after |
| Memory | Re-sending the conversation on every call |
| Persistence | A database, so chats survive past one session |
| The product | Interface, sign-in, billing, streaming, moderation, tool calls |
| Context management | Deciding what to keep when the history outgrows the window |
The model is the engine. Everything that makes it feel like a product is the car built around it. That is why most of the work in a chat application is not the model call.
4The context window
The limit on one request.
The context window is the most the model can handle in one request: your system prompt, every message so far, tool definitions and results, and the reply it is writing. Per Anthropic's documentation, recent Claude models have 1 million tokens, with older ones at 200,000.
5Three strategies
Drop, summarise or pick.
Three ways to stay within the limit, from simplest to most careful:
| Strategy | How | Trade-off |
|---|---|---|
| Drop | Remove the oldest messages | Simple, but forgets early details |
| Summarise | Replace old turns with a short summary, written by a model | Keeps the gist, costs a call, may lose detail |
| Pick | Keep only the messages relevant to the new question | Efficient, needs a way to judge relevance |
6Trimming in code
The simplest version, run and checked.
def est_tokens(msgs): # rough: about 4 characters per token
return sum(len(m["content"]) for m in msgs) // 4
def trim(msgs, limit):
msgs = list(msgs)
while est_tokens(msgs) > limit and len(msgs) > 1:
msgs.pop(0) # drop the oldest message
while msgs and msgs[0]["role"] != "user":
msgs.pop(0) # the list must start with a user turn
return msgs
history = [
{"role": "user", "content": "a" * 40},
{"role": "assistant", "content": "b" * 40},
{"role": "user", "content": "c" * 40},
{"role": "assistant", "content": "d" * 40},
{"role": "user", "content": "e" * 40},
]
print("before:", est_tokens(history), "tokens, roles", [m["role"][0] for m in history])
kept = trim(history, 25)
print("after: ", est_tokens(kept), "tokens, roles", [m["role"][0] for m in kept])before: 50 tokens, roles ['u', 'a', 'u', 'a', 'u'] after: 10 tokens, roles ['u']
The 50-token history is over a 25-token limit, so the oldest messages go one by one until it fits. Two things to notice. It stopped dropping while still having two messages, but one was an assistant turn, so that was removed too: the list has to start with a user message. And only the latest message survived, which is why dropping alone loses context fast.
7Prompt caching
Stop paying for the same prefix twice.
Re-sending a long, unchanged prefix, such as a big system prompt or a document, every turn is wasteful. Prompt caching lets the provider reuse the processing for a prefix that has not changed. You mark where the stable part ends, and later requests with an identical prefix are charged much less for that part.
response = client.messages.create(
model="MODEL_ID_FROM_THE_DOCS",
max_tokens=1024,
cache_control={"type": "ephemeral"}, # as shown in Anthropic's documentation
system=LONG_SYSTEM_PROMPT,
messages=history,
)
From the documentation at the time of writing: a cache write costs about 1.25 times the normal input price for the default 5-minute lifetime, and a cache hit about a tenth of it. Prompts below a minimum length, which depends on the model, are silently not cached, and changing the tools or system prompt invalidates the cache.
8Let the platform help
Compaction and context editing.
Some platforms now offer to shorten long conversations for you. Anthropic describes server-side compaction (in beta at the time of writing) that summarises earlier turns automatically, and context editing that clears old tool results. These are the "summarise" strategy done by the platform. They are convenient, but they still lose detail, so test what your application needs to keep.
9Check yourself
Short questions and one line to remember.
Why does a chatbot seem to have no memory?
The API is stateless. Memory is the message list you re-send, so the model only knows what you include.
What counts toward the context window?
Everything: system prompt, all messages, tool definitions and results, and the reply being written.
What invalidates a prompt cache?
Changing an earlier part of the prompt, such as the tools or system prompt. Anything after the change is reprocessed.
Drop or summarise: when do you pick each?
Drop when old turns truly do not matter. Summarise when the gist of earlier turns does.