← Home · Tutorial 09 · Building with APIs

Building with the Claude API

Tool use, extended thinking, caching, streaming and MCP, with what has changed since older material.

Last reviewed 5 October 2026 · Written with AI assistance and reviewed by me

Where I learned this: my notes from Anthropic's course Building with the Claude API, which covers tool use, extended thinking, prompt caching, streaming and the Model Context Protocol. Because these features change quickly, I re-checked each against Anthropic's current documentation in October 2026 and mark below where the API has moved on from older course-era material. Code shapes come from the docs and are not run here, since they need an API key.
Who this is for: developers who have made a first model call and want the main features. You will learn tool use, extended thinking, prompt caching, streaming and MCP, in the order you will need them.

1Purpose first

Five features beyond a single call.

Tutorial 2 made one model call. Real applications need more: the model using your functions, thinking before answering, replying as it writes, reaching other systems. These are five features that cover most of that. For each: what it is, and the one thing to get right.

2At a glance

What each does and the one thing to get right.

FeatureWhat it doesGet this right
Tool useModel asks your code to run a functionYou run it, you send the result back
Extended thinkingModel reasons before it answersUse the current mode, not the old budget setting
Prompt cachingReuses an unchanged prompt prefixStable content first (tutorial 8)
StreamingReply arrives as it is writtenNeeded for long replies and responsive UIs
MCPStandard way to plug in tools and dataOnly connect servers you trust

3Tool use

The model asks, your code acts.

You describe functions to the model with a name, a description and a JSON schema for the inputs. The model does not run anything. It replies with a request to call one, and your code does the call. The flow, from Anthropic's documentation:

Tool use
You sendmessage + tool definitions
→
Model repliesstop_reason: tool_use
→
Your coderuns the function
→
You send backa tool_result
→
Model answers
Shape from the docs, not run here
tools = [{
    "name": "get_weather",
    "description": "Get the current weather for a given location.",
    "input_schema": {
        "type": "object",
        "properties": {"location": {"type": "string"}},
        "required": ["location"],
    },
}]

response = client.messages.create(model="MODEL_ID_FROM_THE_DOCS", max_tokens=1024,
                                  tools=tools, messages=messages)

if response.stop_reason == "tool_use":
    call = next(b for b in response.content if b.type == "tool_use")
    result = run_my_function(call.name, call.input)         # your code
    messages.append({"role": "assistant", "content": response.content})
    messages.append({"role": "user", "content": [
        {"type": "tool_result", "tool_use_id": call.id, "content": result}]})
    # call the API again with the extended messages

Wrap this in a loop with a step cap and you have the agent from tutorial 10. The description matters more than you would think: it is how the model decides when to use the tool. Some tools, such as web search, are server tools that Anthropic runs for you, so there is no handler to write.

4Extended thinking

Reasoning before answering, and what changed.

Extended thinking lets the model reason before answering, and the thinking tokens are billed as output.

SettingStatus
thinking={"type": "enabled", "budget_tokens": N}Older Deprecated on Claude 4.6 models and not supported from 4.7, per Anthropic's docs
thinking={"type": "adaptive"} with output_config={"effort": ...}Current The model decides when and how much to think
Shape from the docs, not run here
response = client.messages.create(
    model="MODEL_ID_FROM_THE_DOCS",
    max_tokens=16000,
    thinking={"type": "adaptive"},
    output_config={"effort": "high"},
    messages=messages,
)
for block in response.content:
    if block.type == "thinking":
        print("Thinking:", block.thinking)
    elif block.type == "text":
        print("Answer:", block.text)

Use it for multi-step problems where quality matters, and skip it for simple lookups, where it only adds cost and delay. If a course shows budget_tokens, the idea carries over but the setting does not.

5Prompt caching

Do not pay twice for the same prefix.

Covered in tutorial 8. In short: mark the stable prefix with cache_control, keep changing content at the end, and check your prompt is above the model's minimum length or nothing is cached, with no error to tell you. Cache hits cost a fraction of normal input.

6Streaming

Replies as they are written.

A long reply can take many seconds. Streaming returns it as it is generated, so a user sees words appear and your request is less likely to hit a timeout on large outputs.

Shape from the docs, not run here
with client.messages.stream(
    model="MODEL_ID_FROM_THE_DOCS",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

Underneath, the stream is a sequence of events: a message starts, content blocks start, receive deltas and stop, and a final event carries the stop reason and usage. If you only want the finished message without handling events, stream.get_final_message() returns it.

7MCP

A standard plug for tools and data.

Each tool you write is code only your app can use. The Model Context Protocol (MCP) is an open standard for connecting AI applications to tools, data and prompts, so a connector built once can work in many apps. The project's own comparison is a USB-C port for AI.

PieceRole
Host or clientThe AI application that connects (an assistant, an IDE)
ServerExposes tools, resources and prompts to the client
ToolsActions the model can ask for
ResourcesData the application can read
Real risk: an MCP server can run actions and its tool descriptions are read by the model. A malicious or careless server can leak data or steer the model. Connect only servers you trust, give them the least access they need, and be careful combining ones that read untrusted content with ones that can act.

8Check yourself

Short questions and one line to remember.

Does the model run my tool?

No. It asks for a call. Your code runs it and sends back a tool_result.

What replaced budget_tokens for extended thinking?

Adaptive thinking, with an effort setting. Older manual budgets are deprecated or unsupported on newer models.

Why stream?

To show output as it is written and to avoid timeouts on very long replies.

What problem does MCP solve?

Each integration being rebuilt for every app. It gives one standard way to expose tools and data.

One line to remember: "The model asks, my code acts, and every feature here is a way of deciding who does what and what it costs."
Notes based on Anthropic’s course Building with the Claude API, checked against tool use, extended thinking, streaming, prompt caching and the MCP introduction in October 2026. Check the docs before shipping. Independent notes, not affiliated with or endorsed by any company or course named here. Examples and wording are my own.