Skip to content
Cloud Efficiency Hub

Low-Traffic Fine-Tuned Azure OpenAI Deployments Accruing Hosting Fees

The short version

A fine-tuned Azure OpenAI model deployed as Standard or Global Standard carries an hourly hosting fee on top of per-token inference, and Microsoft states that the fee applies whether or not any chat completions or Responses API calls are made.

PointFive Research

Cloud cost research at PointFive

Category
AI
Reference
CER-0561
Type
Idle or Unused Resource

Explanation

Why the waste happens and who it affects.

Teams commonly deploy several candidate fine-tunes to compare them, keep last quarter's version deployed alongside the new one, or leave a fine-tune serving an internal tool that is used a few times a week. Each of these deployments bills around the clock.

Azure deletes a fine-tuned deployment automatically after 15 consecutive days with no chat completions or Responses calls, which caps the cost of a completely abandoned deployment at roughly two weeks of hosting. The larger and more persistent waste is deployments that receive occasional calls, such as a weekly job, a health check or a rarely used feature, because any call resets the inactivity clock and the deployment is never removed. For those, the hosting fee can dwarf the token spend. Stored fine-tuned models cost nothing, so a deployment can be deleted and recreated when needed.

Billing model

The pricing dimensions that drive this cost.

Fine-tuned models are charged for training, hosting and inference.

Hosting fee
Hourly charge per deployed fine-tuned model on Standard and Global Standard ($1.70 per hour in East US, USD, on the pricing page), applied even if the model is unused
Inference
Per-token charge at the same rate as the base model's deployment type
Stored model
A fine-tuned model that is not deployed is stored at no cost and can be redeployed at any time
Developer deployment
Per-token only with no hourly hosting fee and no SLA; deleted automatically after 24 hours

How to detect

5 checks to find it in your estate.

  • List deployments of fine-tuned models (model names containing .ft-) on each Azure OpenAI or Foundry resource with az cognitiveservices account deployment list, including their sku.name
  • For each, sum the Azure OpenAI Requests (AzureOpenAIRequests) metric split by ModelDeploymentName over the last 30 days; flag deployments with only a handful of requests per day or week
  • Compare each deployment's monthly hosting charge in Cost analysis (group by meter and filter to the fine-tuned hosting meters) with its token charges; hosting that exceeds inference by a wide margin indicates a low-traffic deployment
  • Look for several fine-tuned deployments of the same base model and use case, which usually indicates leftover evaluation candidates or superseded versions
  • Check non-production resources for Standard or Global Standard fine-tuned deployments that are used only during testing

How to fix

5 ways to remove the waste.

  • Delete fine-tuned deployments that are not serving production traffic; the underlying model stays stored at no cost and can be redeployed from the portal, REST API or CLI when needed
  • Evaluate candidate fine-tunes on Developer deployments, which have no hourly hosting fee, accepting that they have no SLA and are removed after 24 hours
  • Consolidate traffic onto one production fine-tune per use case and retire superseded versions promptly after a cutover
  • For infrequent batch-style use, deploy the fine-tune just before the job and delete it afterward through automation, rather than keeping it hosted continuously
  • Weigh the hosting fee against the token savings the fine-tune provides; for low-volume use cases a base model with prompt engineering may cost less overall

Documentation

Vendor references for pricing and configuration.