Explanation
Why the waste happens and who it affects.
Teams commonly deploy several candidate fine-tunes to compare them, keep last quarter's version deployed alongside the new one, or leave a fine-tune serving an internal tool that is used a few times a week. Each of these deployments bills around the clock.
Azure deletes a fine-tuned deployment automatically after 15 consecutive days with no chat completions or Responses calls, which caps the cost of a completely abandoned deployment at roughly two weeks of hosting. The larger and more persistent waste is deployments that receive occasional calls, such as a weekly job, a health check or a rarely used feature, because any call resets the inactivity clock and the deployment is never removed. For those, the hosting fee can dwarf the token spend. Stored fine-tuned models cost nothing, so a deployment can be deleted and recreated when needed.
Billing model
The pricing dimensions that drive this cost.
Fine-tuned models are charged for training, hosting and inference.
- Hosting fee
- Hourly charge per deployed fine-tuned model on Standard and Global Standard ($1.70 per hour in East US, USD, on the pricing page), applied even if the model is unused
- Inference
- Per-token charge at the same rate as the base model's deployment type
- Stored model
- A fine-tuned model that is not deployed is stored at no cost and can be redeployed at any time
- Developer deployment
- Per-token only with no hourly hosting fee and no SLA; deleted automatically after 24 hours
How to detect
5 checks to find it in your estate.
- List deployments of fine-tuned models (model names containing .ft-) on each Azure OpenAI or Foundry resource with az cognitiveservices account deployment list, including their sku.name
- For each, sum the Azure OpenAI Requests (AzureOpenAIRequests) metric split by ModelDeploymentName over the last 30 days; flag deployments with only a handful of requests per day or week
- Compare each deployment's monthly hosting charge in Cost analysis (group by meter and filter to the fine-tuned hosting meters) with its token charges; hosting that exceeds inference by a wide margin indicates a low-traffic deployment
- Look for several fine-tuned deployments of the same base model and use case, which usually indicates leftover evaluation candidates or superseded versions
- Check non-production resources for Standard or Global Standard fine-tuned deployments that are used only during testing
How to fix
5 ways to remove the waste.
- Delete fine-tuned deployments that are not serving production traffic; the underlying model stays stored at no cost and can be redeployed from the portal, REST API or CLI when needed
- Evaluate candidate fine-tunes on Developer deployments, which have no hourly hosting fee, accepting that they have no SLA and are removed after 24 hours
- Consolidate traffic onto one production fine-tune per use case and retire superseded versions promptly after a cutover
- For infrequent batch-style use, deploy the fine-tune just before the job and delete it afterward through automation, rather than keeping it hosted continuously
- Weigh the hosting fee against the token savings the fine-tune provides; for low-volume use cases a base model with prompt engineering may cost less overall
Documentation
Vendor references for pricing and configuration.