Opus 5 Killed the Fable Tax
Claude Opus 5 launched at $5/$25 per million input/output tokens. Its vendor-reported benchmarks make it worth evaluating; migration, effort and optional fallback need workload checks.
Anthropic shipped Claude Opus 5 today (model id claude-opus-5) at $5/M input, $25/M output — the exact price of Opus 4.8 — while, by Anthropic’s own numbers, landing near Fable 5 on the reported CursorBench configuration. The reported comparison is specific to that benchmark, with a dial that trades cost for capability per turn.
If you’ve been paying the Fable premium, today is the day to test the switch.
The Fable Tax argument is about whether a premium model earns its additional cost on a particular workload. A lower-priced alternative changes that comparison; it does not settle every workload in advance.
Anthropic’s July 24 announcement positions Opus 5 near Fable 5 on several of its evaluations, at lower API rates. That is a useful reason to test an alternative, not proof of equivalent performance on every task.
What actually shipped (the confirmed part)
Opus 5 launched at $5 per million input tokens and $25 per million output tokens, matching Opus 4.8. Anthropic made it the default on Claude Max and available on Pro. Availability, fast mode and billing differ by product surface; check the current documentation for the route being used.
The effort setting trades additional reasoning work against latency and token consumption. Treat effort as an evaluation parameter, not a fixed cost multiplier or a guarantee that every task benefits from the highest setting.
Anthropic reports stronger alignment and deliberately constrained offensive-cyber capability. Safety classifier declines can trigger an optional fallback, but ordinary API requests do not automatically route to Opus 4.8 whenever the model refuses.
Server-side fallback is an opt-in beta on the Claude API, enabled through the documented beta header and fallbacks parameter. It applies to eligible classifier declines, not general rate limits or server errors, and some categories have no fallback. Record the serving model and per-attempt usage; supported surfaces and the recommended fallback can change.
What the documentation now confirms
The current Opus 5 documentation confirms a one-million-token context window. It also describes migration changes, including thinking enabled by default and restrictions on disabling it at higher effort. Use the official model and effort documentation instead of early community aliases or launch rumors.
The dial is the actual news
The effort dial adds another routing choice within Opus 5. It lets a team compare settings within one model as well as compare different models. That does not establish a universal replacement for a flagship: evaluate model and effort combinations on representative tasks, including correctness, latency and total retry cost. A narrow benchmark result is a reason to run that comparison, not a substitute for it.
A working policy:
The rule is simple: start lower when a mistake is cheap and obvious, start higher when it's expensive or subtle. If a low-effort answer fails inspection, bump that turn — don't re-run the whole conversation at max because one step needed it. (And ignore anyone selling a clean “effort tier = cost multiplier” formula; Anthropic didn't publish one, so they're decorating guesses with decimal points.)
A testable hypothesis: effort routing can reduce total cost where a lower setting still meets the same acceptance criteria. Include verification, retries and failures in that comparison. Some workloads may justify maximum effort throughout.
The benchmarks (Anthropic's own — trust, but verify)
The following results are attributed to Anthropic’s July 24 announcement. They describe the vendor’s particular evaluations and do not establish equivalent savings or quality on an arbitrary production workload.
Note the shape of the coding claim, too: a model can cost the same per token and still be cheaper per finished task if it needs fewer retries and less babysitting. Opus 5 starts at 4.8's exact price and, on these numbers, gets there with less — which is why this launch is ugly for 4.8, not just for Fable.
The narrow result matters: Anthropic reports a small CursorBench score gap at lower task cost and a strong OSWorld comparison. Neither result proves general intelligence equivalence or guarantees a saving on a different workload.
Opus 5 vs Opus 4.8 vs Fable 5
Where a premium model may still earn its cost
Anthropic continues to recommend Fable for demanding long-horizon agentic work. Whether that recommendation applies depends on the task distribution, failure consequences and measured accepted results.
Opus 5 is a candidate to evaluate alongside the existing default. Keep the model that meets quality, compatibility and cost requirements for each workload; a premium tier should have an evidence-based role rather than an assumed one.
What to do Monday
1. Evaluate before changing defaults. Compare claude-opus-5 on representative tasks and review the migration guide, particularly thinking behavior, effort support, tools and optional fallback. Change defaults only where accepted-result quality, compatibility and total cost meet requirements.
2. Evaluate the effort dial. Compare candidate settings on the same acceptance criteria, including latency, retries and review effort. A cheaper individual response is not necessarily a cheaper completed task.
3. Measure finished tasks, not benchmark theater. Anthropic's numbers are a strong prior; your workload is the actual test. Run Opus 5 against your current 4.8 and Fable flows on the same inputs, and track success rate, retries, human corrections, and total cost per accepted result — not tokens on a leaderboard.
4. Keep a justified model portfolio. Retain another model where it demonstrably improves the result or provides a needed availability path. A benchmark release alone does not justify changing every project or dropping a premium model universally.
The useful change is a new option at familiar API rates. Re-evaluate when capability, pricing or access changes, and keep the acceptance test stable enough to distinguish an actual improvement from a different configuration.