Skip to content
Cloud Efficiency Hub

Latency-Tolerant Claude API Workloads Not Using the Message Batches API

The short version

Many Claude API workloads have no user waiting on the answer: evaluation suites, content moderation backlogs, bulk classification and tagging, data extraction, summarization of document sets and generation of product descriptions.

PointFive Research

Cloud cost research at PointFive

Anthropic service
Claude API
Category
AI
Reference
CER-0504
Type
Suboptimal Pricing Model

Explanation

Why the waste happens and who it affects.

When these run as ordinary synchronous Messages API calls, every token is billed at the standard rate. The Message Batches API accepts the same requests, processes them asynchronously, and charges 50 percent of standard prices for both input and output tokens.

The synchronous API is the default path, so offline jobs written as loops over it pay the standard rate unless someone deliberately moves them. Teams may also assume batching means long delays, although Anthropic states most batches finish in under an hour, with a 24-hour limit. Anthropic's own cost guidance calls batching the second-largest free lever after caching for unattended work such as evaluation runs, backfills and scheduled jobs, and the discount applies to every token of a request, including cached ones.

Billing model

The pricing dimensions that drive this cost.

Standard Messages API tokens
Input and output tokens billed per million at each model's standard rate
Batch tokens
All Message Batches usage is charged at 50 percent of standard API prices, for input and output
Unbilled batch results
Requests that error, are canceled before being sent to the model, or expire after 24 hours are not billed
Stacking with caching
Prompt caching multipliers apply on top of the batch discount

How to detect

4 checks to find it in your estate.

  • Use the Usage and Cost Admin API usage report grouped by service_tier, model and API key or workspace to see how much volume runs outside the batch tier
  • Identify callers triggered by schedulers, pipelines, CI evaluation jobs or backfills rather than by interactive users, and compare their volume with batches listed in each workspace
  • Flag workloads that send thousands of similar requests in loops with client-side concurrency or retry logic, a common sign of batch work done synchronously
  • Check whether latency requirements for these jobs are measured in hours rather than seconds

How to fix

4 ways to remove the waste.

  • Submit offline work through the Message Batches API, with a unique custom_id per request, and poll or list batches to collect results; each batch can hold up to 100,000 requests or 256 MB
  • Design for the limits: results arrive when all requests finish or after 24 hours, unfinished requests expire, results can return in any order and stay downloadable for 29 days
  • Combine batching with prompt caching for shared context, and consider the 1-hour cache duration because batches can take longer than 5 minutes to process
  • Keep synchronous calls for interactive paths, and for features batches do not support, such as streaming, fast mode and Claude Managed Agents sessions

Documentation

Vendor references for pricing and configuration.