# Outdated Claude Model Versions on the Claude API

Canonical: https://www.pointfive.co/efficiency-hub/inefficiencies/outdated-claude-model-versions-on-the-claude-api

Claude API requests are billed at the rate of the model ID they name, and applications usually pin an ID in code or configuration.

By: PointFive

Updated: 2026-09-28

[Cloud Efficiency Hub](https://www.pointfive.co/efficiency-hub) 

The short version

Claude API requests are billed at the rate of the model ID they name, and applications usually pin an ID in code or configuration.

PointFive Research

Cloud cost research at PointFive

Anthropic service

[Claude API](https://www.pointfive.co/efficiency-hub/cloud-services/claude-api)

Category

[AI](https://www.pointfive.co/efficiency-hub/service-category/ai)

Reference

CER-0505

Type

Outdated Version

## Explanation

Why the waste happens and who it affects.

When Anthropic ships a newer model in the same family, pinned workloads keep running on the older one. Several older models that are still active are priced higher per token than their successors: Claude Sonnet 4.5 and 4.6 list at $3 / $15 per million input / output tokens against $2 / $10 for Sonnet 5 and 5.5, and Claude Opus 4.5 through Opus 5 list at $5 / $25 against $4 / $20 for Opus 5.5.

Anthropic's own cost guide calls the model string the cheapest lever for teams a model or two behind, and reports that in its measurements each newer model usually solved at least as many tasks for less per solved task. The saving is not automatic, though. Claude 4.7 and later models use a tokenizer that produces about 30 percent more tokens for the same text, and Sonnet 5.5 and Opus 5.5 run adaptive thinking by default, billed as output tokens. Anthropic also states the direction is not guaranteed: on one of its benchmarks the upgrade cost more per task. The waste is paying a higher price per result on an older model without having measured the alternative.

## Billing model

The pricing dimensions that drive this cost.

Per-model token pricing

Input and output tokens are billed per million at the rate of the model ID in each request

Tokenizer change

Claude 4.7 and later models produce about 30 percent more tokens for the same text, so per-token prices are not directly comparable across that boundary

Thinking tokens

Billed as output tokens; newer models that think by default can emit more output per request unless effort is tuned

Retired models

Requests to retired model IDs fail rather than bill, so the cost risk sits with active but superseded versions

## How to detect

4 checks to find it in your estate.

- Group usage by model with the Usage and Cost Admin API usage report, or export the Console Usage page to CSV, which breaks usage down by API key and model

- Flag traffic on active but superseded IDs such as claude-sonnet-4-5-20250929, claude-sonnet-4-6, claude-opus-4-5-20251101, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8 and claude-opus-5, and on models marked Deprecated

- Check the model deprecations page for each pinned ID's lifecycle state and tentative retirement date, and prioritize workloads whose model is nearest retirement

- Search code and configuration for hard-coded model IDs that have no owner or review date

## How to fix

5 ways to remove the waste.

- Evaluate the current model in the same family on a sample of real traffic and compare cost per completed task, not cost per token, as Anthropic recommends

- Sweep effort on the newer model rather than carrying over old settings; a newer model at lower effort is often the cheapest configuration, and Sonnet 5.5 can run without up-front thinking using the between\_tools setting

- Follow the model's migration guide for breaking changes such as rejected sampling parameters, forced tool choice and prefill, and re-baseline max\_tokens because it now covers thinking plus text

- Roll out behind a shadow or canary slice, then move the pinned ID and record an owner and review date so the next upgrade is not missed

- Where the measured cost per task on the newer model is higher for a workload, keep the older model until its retirement date and document why

## Documentation

Vendor references for pricing and configuration.

- [Optimizing for cost and intelligence - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence)

- [Pricing - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/about-claude/pricing)

- [Model deprecations - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/about-claude/model-deprecations)

- [Migrating to Claude Opus 5.5 - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/models/opus-5-5/migration-guide)

- [Migrating to Claude Sonnet 5.5 - Claude Platform Docs  platform.claude.com](https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide)

## Related inefficiencies

[Browse the library](https://www.pointfive.co/efficiency-hub)

- Claude API  CER-0503

### [Missing Prompt Caching on the Claude API](https://www.pointfive.co/efficiency-hub/inefficiencies/missing-prompt-caching-on-the-claude-api)

Agents, coding assistants, RAG applications and multi-turn chat on the Claude API resend the same long prefix on every request: tool definitions, a system prompt, reference documents and the conversation so far. Without prompt caching,...

AI

- Claude API  CER-0504

### [Latency-Tolerant Claude API Workloads Not Using the Message Batches API](https://www.pointfive.co/efficiency-hub/inefficiencies/latency-tolerant-claude-api-workloads-not-using-the-message-batches-api)

Many Claude API workloads have no user waiting on the answer: evaluation suites, content moderation backlogs, bulk classification and tagging, data extraction, summarization of document sets and generation of product descriptions. When...

AI

- Claude API  CER-0506

### [Using High-Cost Claude Models for Low-Complexity Tasks](https://www.pointfive.co/efficiency-hub/inefficiencies/using-high-cost-claude-models-for-low-complexity-tasks)

Routing, classification, tagging, entity extraction, short summarization and similar high-volume tasks are often sent to the same top-tier Claude model that powers an application's hardest agentic work, simply because one model ID is...

AI

---
Source: the public page above. Product screenshots and illustrative interfaces are examples, not live customer data.

