Skip to content
Cloud Efficiency Hub

Using High-Cost OpenAI Models for Low-Complexity Tasks

The short version

Classification, routing, triage, simple data extraction and small scoped edits are often sent to OpenAI's most capable model because a single model name is configured for the whole application.

PointFive Research

Cloud cost research at PointFive

OpenAI service
OpenAI API
Category
AI
Reference
CER-0509
Type
Oversized Model Selection

Explanation

Why the waste happens and who it affects.

The per-token gap between tiers is large: at Standard short-context rates, gpt-6-astra lists at $10.00 input and $50.00 output per million tokens, gpt-6-sol at $2.00 and $10.00, and gpt-6-luna at $0.10 and $0.50. A high-volume task that the smaller model handles equally well pays many times more on the flagship.

OpenAI's own model selection guide positions Luna for scoped tasks, triage and frequent automations, Sol as the everyday model, and Astra for ambiguous problems and deep analysis, and advises keeping the lightest setting that meets the quality bar. Its cost optimization guide lists selecting a smaller model as a primary lever. Reasoning effort matters too: running a capable model at a high effort level on simple work adds billed reasoning tokens, and OpenAI recommends Luna at low effort for simple data extraction.

Billing model

The pricing dimensions that drive this cost.

Per-model token rates
Input, cached input, cache writes and output are billed per million tokens at model-specific rates
Processing tier multiplier
Batch and Flex halve, and Fast mode doubles, the Standard rate for the chosen model
Reasoning effort
Higher effort settings let reasoning models spend more tokens per request, which are billed at the model's rates

How to detect

4 checks to find it in your estate.

  • Group spend by model, project and API key with the Admin Usage API completions endpoint or the usage dashboard, and map keys to applications
  • Identify high-volume request types with short, structured outputs, such as labels, routing decisions or extracted fields, that run on gpt-6-astra or at high reasoning effort
  • Check for a single global model setting shared by interactive agents and background automations
  • Run evaluations on sampled production inputs comparing the current model with smaller tiers and lower effort, recording task success and cost per request

How to fix

4 ways to remove the waste.

  • Route scoped, frequent tasks such as triage, simple extraction and small edits to gpt-6-luna or gpt-6-sol where evaluations show equal quality, keeping gpt-6-astra for ambiguous and demanding work
  • Lower reasoning effort on simple tasks before or alongside switching models, following OpenAI's model and effort guidance
  • Use a cascade that sends requests to the smaller model first and escalates only on low confidence or failed checks
  • Configure the model per task instead of per application, and repeat the evaluation when new models are released

Documentation

Vendor references for pricing and configuration.