Explanation
Why the waste happens and who it affects.
The per-token gap between tiers is large: at Standard short-context rates, gpt-6-astra lists at $10.00 input and $50.00 output per million tokens, gpt-6-sol at $2.00 and $10.00, and gpt-6-luna at $0.10 and $0.50. A high-volume task that the smaller model handles equally well pays many times more on the flagship.
OpenAI's own model selection guide positions Luna for scoped tasks, triage and frequent automations, Sol as the everyday model, and Astra for ambiguous problems and deep analysis, and advises keeping the lightest setting that meets the quality bar. Its cost optimization guide lists selecting a smaller model as a primary lever. Reasoning effort matters too: running a capable model at a high effort level on simple work adds billed reasoning tokens, and OpenAI recommends Luna at low effort for simple data extraction.
Billing model
The pricing dimensions that drive this cost.
- Per-model token rates
- Input, cached input, cache writes and output are billed per million tokens at model-specific rates
- Processing tier multiplier
- Batch and Flex halve, and Fast mode doubles, the Standard rate for the chosen model
- Reasoning effort
- Higher effort settings let reasoning models spend more tokens per request, which are billed at the model's rates
How to detect
4 checks to find it in your estate.
- Group spend by model, project and API key with the Admin Usage API completions endpoint or the usage dashboard, and map keys to applications
- Identify high-volume request types with short, structured outputs, such as labels, routing decisions or extracted fields, that run on gpt-6-astra or at high reasoning effort
- Check for a single global model setting shared by interactive agents and background automations
- Run evaluations on sampled production inputs comparing the current model with smaller tiers and lower effort, recording task success and cost per request
How to fix
4 ways to remove the waste.
- Route scoped, frequent tasks such as triage, simple extraction and small edits to gpt-6-luna or gpt-6-sol where evaluations show equal quality, keeping gpt-6-astra for ambiguous and demanding work
- Lower reasoning effort on simple tasks before or alongside switching models, following OpenAI's model and effort guidance
- Use a cascade that sends requests to the smaller model first and escalates only on low confidence or failed checks
- Configure the model per task instead of per application, and repeat the evaluation when new models are released
Documentation
Vendor references for pricing and configuration.