Explanation
Why the waste happens and who it affects.
Per-token prices differ several-fold across tiers: Claude Fable 5.1 lists at $10 / $50 per million input / output tokens, Opus 5.5 at $4 / $20, Sonnet 5.5 at $2 / $10 and Haiku 4.5 at $1 / $5. On simple, checkable tasks that a smaller model completes just as reliably, the difference is pure overspend.
Anthropic's own guidance is to choose Haiku for simple tasks, Sonnet for most production workloads and Opus for the most complex reasoning, and to start efficiency-first for high-volume, straightforward work. It also warns that the comparison must be made on cost per completed task, not per token: a more capable model can finish harder tasks with fewer turns, a failed task on a cheaper model still bills its tokens plus the retry, and in Anthropic's measurements the ranking flips by workload. The inefficiency is defaulting to the top tier, or to maximum effort, for work nobody has shown needs it.
Billing model
The pricing dimensions that drive this cost.
- Per-model token rates
- Input and output tokens are billed per million at rates that rise with model tier
- Thinking and effort
- Thinking tokens are billed as output, and the effort parameter controls how much thinking, tool calling and self-verification a model does per request
- Cost per completed task
- The effective cost of a workload, including retries and failures, which is the basis Anthropic recommends for comparing models
How to detect
4 checks to find it in your estate.
- Break down spend by model and by API key or workspace with the Usage and Cost Admin API, and map keys to calling applications
- Identify high-volume request types with short inputs and short, structured outputs, such as labels, JSON fields or yes/no decisions, that run on Fable or Opus tiers
- Check the effort setting on those calls; high, xhigh or max effort on simple tasks multiplies thinking output, so measure whether it changes outcomes
- Run an offline evaluation on a sample of real traffic with outcome checks, recording cost per completed task for the current model and for smaller tiers at different effort levels
How to fix
5 ways to remove the waste.
- Route simple, checkable tasks to Claude Haiku 4.5 or Sonnet 5.5 where evaluations show equal task success, and keep Opus or Fable for complex reasoning and long agentic loops
- Lower effort before switching models when quality is already fine; Anthropic notes that tuning effort is often a better lever than changing models
- For checkable outputs, run at a low tier or low effort first and re-run only failures at a higher setting
- Use multi-model patterns such as a lower-cost executor with a frontier advisor, or an orchestrator delegating bulk work to cheaper workers, where measurement shows they beat the single-model baseline
- Replace a single global model ID with per-task configuration and re-run the evaluation suite when new models ship
Documentation
Vendor references for pricing and configuration.