Skip to content

Claude Fable 5.1 API Tutorial: Build a Long-Running Agent

Learn how to build a Python agent with Claude Fable 5.1, including API setup, adaptive thinking, tool use, cost controls, and production safeguards.

Claude Fable 5.1 API Tutorial: Build a Long-Running Agent

On this page

Claude Fable 5.1 changes what you should expect from a model API call: instead of treating every request as a short question-and-answer exchange, you can design an application around work that may span multiple steps, tool calls, and hours. Anthropic released Fable 5.1 on September 1, 2026, and its API is available through the Claude Platform under the model ID claude-fable-5-1.Ā 

This tutorial shows how to build a small Python agent around that model, give it a controlled tool, tune its effort, and handle the API behavior that differs from older Claude integrations. The useful part is not simply changing a model name: Fable 5.1 uses adaptive thinking automatically, so some older thinking configuration patterns should be removed rather than copied forward.Ā 

What changes when you build for Fable 5.1

Claude Fable 5.1 is positioned for long-running coding, research, and knowledge-work tasks. Anthropic says the model can plan work, use tools, recover from failed steps, and continue through larger projects with less supervision. Its API uses the same Messages API pattern developers already know from newer Claude models, but its always-on adaptive thinking changes how you control reasoning and cost.Ā 

For this reason, the simplest migration is often the cleanest one. Do not add a manual thinking budget just because an older Claude example contains one. Fable 5.1 determines when and how much to think, while the effort setting gives you a separate way to influence how much work the model puts into the response. Anthropic recommends starting at a high effort level for demanding workloads and then measuring your own latency and token use.

Set up the Python project before making a request

You need Python, an Anthropic API key, and the current Anthropic Python SDK. Keep the key outside your source code so it cannot accidentally be committed to a repository. On a local machine, an environment variable is enough for this tutorial.

pip install anthropic

Then set the API key in your operating system environment. The SDK reads it when you create the client, which keeps the credential separate from the application logic.

import os
from anthropic import Anthropic

client = Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

At this point, the result you want is simple: Python can create an authenticated Anthropic client without putting a secret directly into the program.

Make the first Fable 5.1 request

Start with a normal Messages API call before adding tools. The model identifier for Claude Fable 5.1 is claude-fable-5-1. Anthropic lists the model as available through the Claude Platform and cloud marketplaces, with pricing of $10 per million input tokens and $50 per million output tokens. ([Anthropic][1])

from anthropic import Anthropic

client = Anthropic()

response = client.messages.create(
model="claude-fable-5-1",
max_tokens=4000,
output_config={"effort": "high"},
messages=[
{
"role": "user",
"content": "Review this project plan and identify its three biggest risks."
}
],
)

print(response.content)

The important detail is what is missing. There is no thinking block and no manual budget_tokens value. Fable 5.1 uses adaptive thinking automatically, and Anthropic documents manual extended-thinking configuration as unsupported for the Fable model family. ([Claude Platform Docs][2])

Add a tool so the agent can do real work

A model becomes more useful as an agent when it can take an action instead of only describing one. In this example, the tool is deliberately small: the model can ask your application to read a project note. Your program remains in control of what data is actually returned.

tools = [
{
"name": "read_project_note",
"description": "Read a project note by name.",
"input_schema": {
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "The note name to read."
}
},
"required": ["name"]
}
}
]

Now define the local function that handles the tool request. A production application should validate the requested name and restrict the files or records the model is allowed to access.

def read_project_note(name):
notes = {
"release": "Release candidate is scheduled for Friday.",
"bugs": "Three high-priority bugs remain open.",
"marketing": "The launch campaign is waiting for final screenshots."
}
return notes.get(name, "No note found.")

The result you should have now is a clear boundary between reasoning and execution: Claude decides that it needs information, but your application decides whether and how that information is supplied.

Let Fable 5.1 continue after the tool call

The next step is to inspect the response for a tool request, execute that request locally, and send the result back as another message. This loop is the foundation of a long-running agent because the model can repeatedly decide what information or action it needs next.

messages = [
{
"role": "user",
"content": "Check the release and bugs notes, then tell me whether the project is ready."
}
]

response = client.messages.create(
model="claude-fable-5-1",
max_tokens=4000,
output_config={"effort": "high"},
tools=tools,
messages=messages,
)

for block in response.content:
if block.type == "text":
print(block.text)

For a production agent, the loop should also recognize tool-use blocks, execute only approved tools, append the tool result to the conversation, and call the model again until it produces a final response. Keep that loop explicit rather than giving the model unrestricted access to your operating system or credentials.

Use effort instead of an old thinking budget

One of the easiest mistakes when adapting an older Claude integration is carrying forward manual extended-thinking settings. Anthropic says Fable 5.1 uses adaptive thinking and does not support manually enabled thinking with a budget_tokens value. The effort setting is separate: it controls the amount of effort applied to the response rather than directly specifying a thinking-token budget. ([Claude Platform Docs][3])

For a difficult research or coding task, start with high. For routine work, test a lower setting and compare the actual output quality, latency, and token consumption in your application. Do not assume that fewer output tokens automatically means a cheaper or faster complete task: an agent can spend additional calls or tool operations compensating for an incomplete response.

Give the agent a stopping rule

Long-running agents need boundaries. A useful instruction tells the model what counts as finished, which tools it may use, and what it should do when information is missing. Without those boundaries, a capable model can spend effort investigating details that do not change the answer.

system_prompt = """
You are a project review agent.

Your job is to inspect the available project notes and decide
whether the release is ready.

Use tools only when the required information is not already available.
Stop when you have enough evidence to make the decision.
Report unresolved risks separately from confirmed blockers.
Do not claim that a check was performed unless a tool actually returned
the required information.
"""

That final instruction is particularly useful for agentic systems. The model should distinguish between an action it actually performed and an action it merely recommended. This makes the final report easier to audit and reduces the chance that a plausible-sounding answer is mistaken for evidence.

Watch the cost before turning it into a worker

Fable 5.1 is priced substantially higher per token than many smaller Claude models, so the economics of a long-running agent depend on more than the price of one response. Anthropic lists $10 per million input tokens and $50 per million output tokens for Fable 5.1, while cache reads are priced at $0.25 per million tokens. Anthropic says the lower cache-read price can reduce typical workload costs by an estimated 25% and highly agentic workloads by up to approximately 45%, although those are Anthropic's estimates rather than an independent benchmark. ([Anthropic][1])

Measure your own workload using the same prompts, tools, context size, and stopping rules you expect in production. A task that looks cheap when tested as one request can become expensive when an agent repeats tool calls or carries a large conversation history across many turns.

Handle Fable 5.1's API differences before deployment

There are several migration details worth testing before replacing an existing Claude model. Fable 5.1 does not support assistant-message prefilling, and disabling adaptive thinking is not an alternative because thinking remains part of the model's behavior. The API can also return a successful HTTP response with a refusal stop reason when its safety classifiers reject a request, so application code should not assume that every HTTP 200 response contains an ordinary answer. ([Claude Platform Docs][2])

Data handling also deserves attention. Anthropic says Fable 5.1 normally requires 30-day data retention for safety monitoring, while Enterprise Frontier Safeguards is being introduced to allow eligible organizations to keep relevant activity data in customer-controlled cloud infrastructure. That is an architectural and compliance consideration, not something a prompt can solve. ([Anthropic][1])

Turn the prototype into a reliable agent

The first useful version of a Fable 5.1 agent does not need dozens of tools. Start with one narrow job, one or two controlled tools, a clear completion condition, and logging for every model and tool call. Once that works, add retries for transient failures, validation around tool arguments, limits on repeated calls, and tests containing cases where the correct action is to stop rather than continue.

That approach also gives you a meaningful baseline for judging Fable 5.1. Anthropic designed the model for long-running work, but capability alone does not make an application reliable. The real test is whether your agent can finish the same class of tasks repeatedly, show what it actually did, stay within an acceptable cost, and fail safely when the information or tool it needs is unavailable.

S

Written by

Sarah Khan

I’m fascinated by artificial intelligence and the rapid changes happening around AI tools, models, and agents. I enjoy testing new AI technologies, following important developments, and understanding how they can be useful in real life. I like explaining complex AI topics in a simple and practical way.

44 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.