Skip to content
Cloud Efficiency Hub

Outdated Claude Model Versions on the Claude API

The short version

Claude API requests are billed at the rate of the model ID they name, and applications usually pin an ID in code or configuration.

PointFive Research

Cloud cost research at PointFive

Anthropic service
Claude API
Category
AI
Reference
CER-0505
Type
Outdated Version

Explanation

Why the waste happens and who it affects.

When Anthropic ships a newer model in the same family, pinned workloads keep running on the older one. Several older models that are still active are priced higher per token than their successors: Claude Sonnet 4.5 and 4.6 list at $3 / $15 per million input / output tokens against $2 / $10 for Sonnet 5 and 5.5, and Claude Opus 4.5 through Opus 5 list at $5 / $25 against $4 / $20 for Opus 5.5.

Anthropic's own cost guide calls the model string the cheapest lever for teams a model or two behind, and reports that in its measurements each newer model usually solved at least as many tasks for less per solved task. The saving is not automatic, though. Claude 4.7 and later models use a tokenizer that produces about 30 percent more tokens for the same text, and Sonnet 5.5 and Opus 5.5 run adaptive thinking by default, billed as output tokens. Anthropic also states the direction is not guaranteed: on one of its benchmarks the upgrade cost more per task. The waste is paying a higher price per result on an older model without having measured the alternative.

Billing model

The pricing dimensions that drive this cost.

Per-model token pricing
Input and output tokens are billed per million at the rate of the model ID in each request
Tokenizer change
Claude 4.7 and later models produce about 30 percent more tokens for the same text, so per-token prices are not directly comparable across that boundary
Thinking tokens
Billed as output tokens; newer models that think by default can emit more output per request unless effort is tuned
Retired models
Requests to retired model IDs fail rather than bill, so the cost risk sits with active but superseded versions

How to detect

4 checks to find it in your estate.

  • Group usage by model with the Usage and Cost Admin API usage report, or export the Console Usage page to CSV, which breaks usage down by API key and model
  • Flag traffic on active but superseded IDs such as claude-sonnet-4-5-20250929, claude-sonnet-4-6, claude-opus-4-5-20251101, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8 and claude-opus-5, and on models marked Deprecated
  • Check the model deprecations page for each pinned ID's lifecycle state and tentative retirement date, and prioritize workloads whose model is nearest retirement
  • Search code and configuration for hard-coded model IDs that have no owner or review date

How to fix

5 ways to remove the waste.

  • Evaluate the current model in the same family on a sample of real traffic and compare cost per completed task, not cost per token, as Anthropic recommends
  • Sweep effort on the newer model rather than carrying over old settings; a newer model at lower effort is often the cheapest configuration, and Sonnet 5.5 can run without up-front thinking using the between_tools setting
  • Follow the model's migration guide for breaking changes such as rejected sampling parameters, forced tool choice and prefill, and re-baseline max_tokens because it now covers thinking plus text
  • Roll out behind a shadow or canary slice, then move the pinned ID and record an owner and review date so the next upgrade is not missed
  • Where the measured cost per task on the newer model is higher for a workload, keep the older model until its retirement date and document why

Documentation

Vendor references for pricing and configuration.