Skip to content
Cloud Efficiency Hub

Using High-Cost Claude Models for Low-Complexity Tasks

The short version

Routing, classification, tagging, entity extraction, short summarization and similar high-volume tasks are often sent to the same top-tier Claude model that powers an application's hardest agentic work, simply because one model ID is configured globally.

PointFive Research

Cloud cost research at PointFive

Anthropic service
Claude API
Category
AI
Reference
CER-0506
Type
Oversized Model Selection

Explanation

Why the waste happens and who it affects.

Per-token prices differ several-fold across tiers: Claude Fable 5.1 lists at $10 / $50 per million input / output tokens, Opus 5.5 at $4 / $20, Sonnet 5.5 at $2 / $10 and Haiku 4.5 at $1 / $5. On simple, checkable tasks that a smaller model completes just as reliably, the difference is pure overspend.

Anthropic's own guidance is to choose Haiku for simple tasks, Sonnet for most production workloads and Opus for the most complex reasoning, and to start efficiency-first for high-volume, straightforward work. It also warns that the comparison must be made on cost per completed task, not per token: a more capable model can finish harder tasks with fewer turns, a failed task on a cheaper model still bills its tokens plus the retry, and in Anthropic's measurements the ranking flips by workload. The inefficiency is defaulting to the top tier, or to maximum effort, for work nobody has shown needs it.

Billing model

The pricing dimensions that drive this cost.

Per-model token rates
Input and output tokens are billed per million at rates that rise with model tier
Thinking and effort
Thinking tokens are billed as output, and the effort parameter controls how much thinking, tool calling and self-verification a model does per request
Cost per completed task
The effective cost of a workload, including retries and failures, which is the basis Anthropic recommends for comparing models

How to detect

4 checks to find it in your estate.

  • Break down spend by model and by API key or workspace with the Usage and Cost Admin API, and map keys to calling applications
  • Identify high-volume request types with short inputs and short, structured outputs, such as labels, JSON fields or yes/no decisions, that run on Fable or Opus tiers
  • Check the effort setting on those calls; high, xhigh or max effort on simple tasks multiplies thinking output, so measure whether it changes outcomes
  • Run an offline evaluation on a sample of real traffic with outcome checks, recording cost per completed task for the current model and for smaller tiers at different effort levels

How to fix

5 ways to remove the waste.

  • Route simple, checkable tasks to Claude Haiku 4.5 or Sonnet 5.5 where evaluations show equal task success, and keep Opus or Fable for complex reasoning and long agentic loops
  • Lower effort before switching models when quality is already fine; Anthropic notes that tuning effort is often a better lever than changing models
  • For checkable outputs, run at a low tier or low effort first and re-run only failures at a higher setting
  • Use multi-model patterns such as a lower-cost executor with a frontier advisor, or an orchestrator delegating bulk work to cheaper workers, where measurement shows they beat the single-model baseline
  • Replace a single global model ID with per-task configuration and re-run the evaluation suite when new models ship

Documentation

Vendor references for pricing and configuration.