1Purpose first
Five features beyond a single call.
Tutorial 2 made one model call. Real applications need more: the model using your functions, thinking before answering, replying as it writes, reaching other systems. These are five features that cover most of that. For each: what it is, and the one thing to get right.
2At a glance
What each does and the one thing to get right.
| Feature | What it does | Get this right |
|---|---|---|
| Tool use | Model asks your code to run a function | You run it, you send the result back |
| Extended thinking | Model reasons before it answers | Use the current mode, not the old budget setting |
| Prompt caching | Reuses an unchanged prompt prefix | Stable content first (tutorial 8) |
| Streaming | Reply arrives as it is written | Needed for long replies and responsive UIs |
| MCP | Standard way to plug in tools and data | Only connect servers you trust |
3Tool use
The model asks, your code acts.
You describe functions to the model with a name, a description and a JSON schema for the inputs. The model does not run anything. It replies with a request to call one, and your code does the call. The flow, from Anthropic's documentation:
tools = [{
"name": "get_weather",
"description": "Get the current weather for a given location.",
"input_schema": {
"type": "object",
"properties": {"location": {"type": "string"}},
"required": ["location"],
},
}]
response = client.messages.create(model="MODEL_ID_FROM_THE_DOCS", max_tokens=1024,
tools=tools, messages=messages)
if response.stop_reason == "tool_use":
call = next(b for b in response.content if b.type == "tool_use")
result = run_my_function(call.name, call.input) # your code
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": [
{"type": "tool_result", "tool_use_id": call.id, "content": result}]})
# call the API again with the extended messages
Wrap this in a loop with a step cap and you have the agent from tutorial 10. The description matters more than you would think: it is how the model decides when to use the tool. Some tools, such as web search, are server tools that Anthropic runs for you, so there is no handler to write.
4Extended thinking
Reasoning before answering, and what changed.
Extended thinking lets the model reason before answering, and the thinking tokens are billed as output.
| Setting | Status |
|---|---|
thinking={"type": "enabled", "budget_tokens": N} | Older Deprecated on Claude 4.6 models and not supported from 4.7, per Anthropic's docs |
thinking={"type": "adaptive"} with output_config={"effort": ...} | Current The model decides when and how much to think |
response = client.messages.create(
model="MODEL_ID_FROM_THE_DOCS",
max_tokens=16000,
thinking={"type": "adaptive"},
output_config={"effort": "high"},
messages=messages,
)
for block in response.content:
if block.type == "thinking":
print("Thinking:", block.thinking)
elif block.type == "text":
print("Answer:", block.text)
Use it for multi-step problems where quality matters, and skip it for simple lookups, where it only adds cost and delay. If a course shows budget_tokens, the idea carries over but the setting does not.
5Prompt caching
Do not pay twice for the same prefix.
Covered in tutorial 8. In short: mark the stable prefix with cache_control, keep changing content at the end, and check your prompt is above the model's minimum length or nothing is cached, with no error to tell you. Cache hits cost a fraction of normal input.
6Streaming
Replies as they are written.
A long reply can take many seconds. Streaming returns it as it is generated, so a user sees words appear and your request is less likely to hit a timeout on large outputs.
with client.messages.stream(
model="MODEL_ID_FROM_THE_DOCS",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
Underneath, the stream is a sequence of events: a message starts, content blocks start, receive deltas and stop, and a final event carries the stop reason and usage. If you only want the finished message without handling events, stream.get_final_message() returns it.
7MCP
A standard plug for tools and data.
Each tool you write is code only your app can use. The Model Context Protocol (MCP) is an open standard for connecting AI applications to tools, data and prompts, so a connector built once can work in many apps. The project's own comparison is a USB-C port for AI.
| Piece | Role |
|---|---|
| Host or client | The AI application that connects (an assistant, an IDE) |
| Server | Exposes tools, resources and prompts to the client |
| Tools | Actions the model can ask for |
| Resources | Data the application can read |
8Check yourself
Short questions and one line to remember.
Does the model run my tool?
No. It asks for a call. Your code runs it and sends back a tool_result.
What replaced budget_tokens for extended thinking?
Adaptive thinking, with an effort setting. Older manual budgets are deprecated or unsupported on newer models.
Why stream?
To show output as it is written and to avoid timeouts on very long replies.
What problem does MCP solve?
Each integration being rebuilt for every app. It gives one standard way to expose tools and data.