How to Master Claude Fable 5.1 (Full Guide)

Anthropic just shipped the strongest model it has ever built. I spent the last 24 hours pushing it until it broke. Here's everything that actually matters, plus the exact settings and prompts I'd give a friend.
What actually changed :
2. The price is still $10 in / $50 out per million tokens. But cached inputs are now 75% cheaper, dropping to $0.25. This is where you can save money.
3. Terminal-Bench-Science 0.1 jumped from 24.7% to 52.6%. Terminal-Bench 4.0 also went up, from 42.0% to 55.8%.
4. There are three important changes you need to know about. Forced tool use now gives a 400 error. Older models can't read Fable 5.1's thinking blocks. And if you edit an earlier message, those thinking blocks become invalid.
5. Effort still matters a lot. There are five levels, and high is the default. But the names don't work the same way they did on your previous model.
number 4 is the one nobody is really talking about in their launch posts, and it's the one that could cause you problems the most. I'll come back to it.
The effort dial
Effort is the highest-leverage setting on this model, and it’s also the one most people will get wrong by just using the default.
It lives inside output_config, not at the top level of your request.
And unlike older models, adaptive thinking is always on.
So thinking: {"type": "disabled"} is no longer an option, and budget_tokens is gone.
The simple version is that you can now control how much effort the model puts into a response, rather than manually managing its thinking budget.
Five levels. Start at high, then sweep against your own evals. Find the lowest setting that still gives you the quality you need.
Here is what each level is actually for.
LOW
For high-volume, latency-sensitive work and subagents you're spawning by the dozen. At low, Fable 5.1 can be competitive on cost per task while scoring higher than smaller models. The catch: it uses search and retrieval tools less often and answers from memory more.
MEDIUM
Roughly matches Fable 5's quality at lower cost. If you were happy with Fable 5, this is the drop-in setting, and most teams should land here for routine work.
HIGH
The API default and where you start. Not because it's the best setting for your workload, but because it's the reference point you measure the others against.
XHIGH
For coding and agentic loops where you've measured that high leaves something on the table. For Fable 5.1, start at high and move up only when your evals show you need it.
MAX
Removes the token ceiling entirely. I wouldn't run it by default, and I'll explain why in the section on traps.
THE RULE
Effort levels don't mean the same amount of thinking across models. A setting tuned on Fable 5 isn't necessarily right for Fable 5.1. Re-run the sweep instead of blindly porting your old config.
NEW IN 5.1
You can change effort mid-conversation without invalidating the prompt cache. Raise it for the hard step, then drop it for routine ones.
Settings that matter even beyond effort
MAX TOKENS
Covers both thinking and the final reply, not just the reply. At xhigh and max, they share one budget, so don't set it to only the length of your expected answer. For long outputs, stream instead of waiting for one huge response.
1M CONTEXT
Useful, but probably less often than you think. Fable 5.1 is better at connecting details across a huge context, but most tasks that seem to need 1M tokens actually need better retrieval. Use it when the connections across the whole corpus matter.
CACHING
Cache reads are much cheaper on Fable 5.1, which changes the usual strategy. Compacting a conversation early to save money may no longer be the best move. Keeping more of the warm cached context can be cheaper than making the model rethink it later.
TASK BUDGET
For agent loops, "task_budget" gives you an advisory token ceiling. Set it. An unattended agent running at $50 per million output tokens can turn into a very expensive surprise.
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-fable-5-1",
max_tokens=64000, # room for thinking AND the reply
output_config={
"effort": "high", # start here, sweep later
"task_budget": {"type": "tokens", "total": 200000},
},
messages=[{"role": "user", "content": "..."}],
)Three things now return a 400 on this model: prefilling the assistant turn, using any non-default temperature / top_p / top_k, and setting thinking to enabled-with-budget or disabled. If your wrapper sets temperature to 0 by reflex, it is already broken.
The three breaking changes
This is the part the launch coverage skipped.
Forced tool use is gone. tool_choice: {"type": "any"} and {"type": "tool", "name": "..."} now return a 400. Thinking is always on, so forcing a tool would bypass it and push the model's reasoning into the tool arguments. If you used forced tools to guarantee valid JSON, switch to strict tool use with tool_choice: {"type": "auto"} or structured outputs. If you need a tool called, just tell the model in the prompt - it follows explicit tool instructions reliably.
Thinking blocks only move forward. Fable 5.1 can read thinking blocks from earlier models, but earlier models can't read Fable 5.1's. So if your router falls back from Fable 5.1 to Opus 5 mid-conversation, the reasoning gets dropped unless you use the thinking-binding-controls-2026-08-01 beta header and monitor input_transformations.
Editing earlier turns now breaks thinking blocks. This applies to accounts created on or after August 31, 2026. Changing the system prompt, injecting a reminder and removing it later, or rewriting old turns can break the conversation. Treat history as append-only. Use turn-scoped system messages for reminders and let server-side compaction handle trimming.
Prompting Fable 5.1
Anthropic published a model-specific prompting guide with canonical versions of several of these. Mine are written to drop straight into a system prompt; theirs are worth reading for the reasoning.
1. Long-running agentic coding
The single highest-value block. This model can run for hours, and its biggest failure mode is stopping to ask for permission to do work you already authorized.
You are operating autonomously. The person who assigned this task is not
watching in real time and cannot answer questions mid-run. Asking "want me
to proceed?" blocks the work.
Proceed without asking on any reversible action that follows from the
original request. Stop only for destructive actions, or for a genuine scope
change the requester has to decide.
Before you end your turn, read your final paragraph. If it is a plan, a
question, a list of next steps, or a promise about work you have not done,
that work is not finished. Do it now, with tool calls. Retry after errors.
Gather missing information yourself. A long session is not a reason to stop.
End your turn only when the task is complete, or when you are blocked on
something only the requester can provide.
Scope discipline: the request sets the deliverable. Do not quietly narrow
it, widen it, or swap it. If you find a pre-existing bug or a performance
problem the task did not mention, do not fix it. Report it as a follow-up
in your summary. Commit tests only where the task asks for them, sized like
the neighboring test files.this works because the opening line about nobody watching does most of the work. The final check turns “here’s what I’d do next” into actually doing it.
2. The tool-batching fix
This fixes the parallel-call regression. Append it to the end of every turn where you send tool results back.
Before calling any tool, privately list everything you need next. Then, in
this single response, request every item that does not depend on another
item's result. Only serialize calls that genuinely have to wait on each
other.Why it works: The regression happens when the next reads are implied by the task instead of explicitly named. This makes the model identify them first. Deliver it as a turn-scoped system message with "clear_at: "next_user_message"" so you append it instead of editing history.
3. Research across the 1M window.
You have the full corpus in context. Do not summarize it document by
document.
Work in this order. First, identify the claims that appear in more than one
source, and note where they agree and where they conflict. Second, identify
claims that appear in exactly one source and flag each as single-sourced.
Third, name what is missing: questions the corpus raises but does not answer.
Cite every claim by document and section. When you use a source's exact
wording, mark it as a quotation. Do not reproduce passages unmarked.
If two sources conflict, say so explicitly and give both figures. Do not
silently pick one.Why it works: Fable 5.1 is more likely than Fable 5 to reproduce source wording without marking it as a quote, so the quotation instruction actually matters.
4. Forcing progress updates back on
The model writes less user-facing text between tool calls, and it gets worse at higher effort. First, check that your client is actually receiving them. Progress updates arrive as thinking blocks, which are empty by default because "thinking.display" is set to ""omitted"". Set "display: "updates"" with the "thinking-display-updates-2026-08-18" beta header.
Then remove any old prompt instruction telling the model to save findings for the final response. Only after that, add this:
Open with one line stating what you are about to do. While you work, post a
short update whenever you finish a meaningful step: what you found, and what
you are doing next. Close with a recap that stands on its own, so someone
who reads only your last message knows what you found, what you changed, and
what remains.5. Root-cause debugging. This is where the model is genuinely a step change.
A bug is described below. Do not propose a fix yet.
First, reproduce it. Confirm you can trigger the failure before theorizing.
If you cannot reproduce it, say so and state what you need.
Second, trace the failure to the line where the invariant actually breaks,
not the line where it surfaces. Read the surrounding code and the call
sites. A stack trace tells you where it died, not where it went wrong.
Third, state the root cause in one paragraph, and separately state the
evidence for it. If the evidence is a pattern match to a known failure
mode, say that, and go find direct evidence.
Only then propose the fix. Explain what the fix does not cover.
If the evidence supports two causes, present both and say what would
distinguish them.This model is especially good at debugging because it tends to reject band-aid fixes. This prompt keeps it in that mode by separating the diagnosis from the fix, and forcing the evidence to be stated separately from the conclusion.
6. Whole-file rewrite suppression
Fable 5.1 rewrites entire files for small edits more often than Fable 5 did. The result is usually correct, but it is more expensive than necessary.
Editing tokens are a cost worth controlling. Where the finished file would
be identical either way, patch the specific lines that change instead of
regenerating the file end to end. Full rewrites are for cases where most of
the file is actually changing.7. Formatting
Fable 5.1 uses bold, headers, and lists less than earlier models. That means anti-formatting rules you wrote for Sonnet or Opus may now suppress structure your content actually needs. Delete them instead of tuning them. If you need a rule, make it conditional:
Use lists and headers when the content is genuinely multifaceted and
structure aids clarity. Use plain prose for conversational, personal, or
emotional exchanges, and whenever the reader has asked for minimal
formatting.When not to use Fable 5.1
Anthropic's own model docs say to start with Opus 5 for most workloads. Reach for Fable 5.1 when you need demanding reasoning, long-horizon agentic work, or when Opus 5 still falls short at higher effort. I'd follow that.
The honest default is Opus 5 at half the token price.
Opus 5 is $5 in / $25 out, exactly half. Artificial Analysis disputes the savings story: in pre-release testing, they found Fable 5.1 costs about 20% more per task than Fable 5 because it uses roughly 1.7x more output tokens. They also scored Opus 5 three points lower on their index, but at a meaningfully lower cost per task. Their numbers, not Anthropic's, so weigh them accordingly.
Sonnet 5 still makes sense for high-volume work where you don't need either model's ceiling. Batch processing is half price at $5 / $25 per million if your workload can tolerate the latency.
Max effort is usually a trap. At xhigh and especially max, Fable 5.1 can spend most of a long deliverable in its thinking, then write it all again as the final reply. You wait longer and pay for the same content twice. Run long deliverables at high and only move up when your evals show a real quality gain.
One more cost note: the tokenizer changed with Opus 4.7. The same text produces roughly 30% more tokens than older models, so comparing spend against a pre-4.7 baseline isn't apples-to-apples.
Three things to change today/now
if you made it this far, follow @mikenevermiss for more detailed AI articles every week.


