Skip to content
PointFive
All articles

Deployed and Forgotten: LLM Cost Optimization on Bedrock, Microsoft Foundry, and Google Cloud

Gal Ben-DavidLinkedInCo-founder & CPO, PointFive6 min read

Most teams treat model choice as a one-time decision. Someone picks a model when the feature is built, the feature ships, and the model never changes again.

Meanwhile, the catalog moves. Amazon Bedrock, Microsoft Foundry, and Gemini Enterprise Agent Platform (formerly Vertex AI) each offer a large and growing set of models. Some are current. Some were the best option a year ago and have since been replaced by models that are cheaper, more capable, or both. A workload that was well chosen at launch can quietly become one of the most expensive ways to do its job.

This is the single most common pattern behind LLM overspend on the hyperscalers: deployed and forgotten.

Three ways LLM workloads waste money

1. Running on last year's model

New models arrive constantly, and many of them change the economics of a task overnight.

A concrete example: OpenAI's GPT-6 Luna became generally available on Amazon Bedrock on September 22, 2026. OpenAI lists it at $0.10 per million input tokens and $0.50 per million output tokens. Claude Haiku 4.5, a common choice for high-volume tasks, lists at $1 and $5. That is a tenfold difference in price per token.

Price per token is not the whole story. Third-party benchmarks generally report lower latency for Haiku 4.5 than for Luna at its higher reasoning settings, so for latency-sensitive work the cheaper model may not be the better one. Which model actually wins for a given task depends on that task. The point is not that every team should switch. It is that a team that has never tested the newer option does not know what it is paying for.

Older models also do not stay available forever. Amazon Bedrock, for example, moves models through Active, Legacy, and End-of-Life states. Once a model enters Legacy, new Provisioned Throughput can no longer be created for it, and after its end-of-life date requests to it fail. In Bedrock's own words, migration to a newer model "will not happen automatically." Microsoft and Google retire older model versions on their platforms too.

Deployed and forgotten is not just expensive. Eventually it breaks.

2. A big model for a small task

The second pattern is using a large, general model for a task that a smaller model handles just as well: classification, extraction, routing, short summaries, format conversion. The large model is often chosen during prototyping, when quality matters more than cost and nobody knows yet what the task will look like at volume. Then it ships as is.

At production volume, that choice multiplies. A task that runs millions of times a month on a model several times more expensive than it needs to be is one of the largest line items an AI team can have, and one of the easiest to miss, because the output looks fine.

3. Measuring the wrong thing

The third pattern is the one that makes the first two invisible. Most teams measure raw metrics: tokens, requests, spend per model, spend per day. Those numbers tell you how much you used. They do not tell you whether the usage worked.

A cheaper model that fails more often is not cheaper. Failed tasks get retried, escalated to a larger model, or fixed by a person. Their cost does not disappear. It moves somewhere harder to see.

The metric that matters: cost per successful task

The right unit of LLM cost is cost per successful task: the total cost of a workload divided by the number of tasks it completed successfully.

What counts as success depends on the task, and you should define it yourself. It might be an evaluation pass, a user accepting the answer, a ticket resolved without escalation, a record extracted correctly, or a downstream step completing. What matters is that success is measured, not assumed.

An illustration of why it matters: imagine two models for the same task. Model A costs a quarter as much per request as Model B, but succeeds only half as often. Every failure triggers a retry on Model B. Per request, Model A looks like an obvious saving. Per successful task, it can easily cost more, because it is paying for the failures and then paying Model B anyway.

Only cost per successful task makes that visible.

The fix: evaluate every task on several models, continuously

The catalogs on Bedrock, Microsoft Foundry, and Gemini Enterprise Agent Platform are a feature, not a complication. They give you a large set of models from several providers, at very different price points, through a single platform. Use them.

  1. Inventory your tasks, not your models. List what your LLM workloads actually do: each distinct task, its volume, and the model it runs on today.
  2. Define success for each task. Pick the signal that tells you a task was done right.
  3. Evaluate each task on several candidate models. Include current models from more than one provider, and include smaller ones. Compare cost per successful task, along with latency where it matters.
  4. Pick the best model per task, not one model for everything.
  5. Re-run the evaluation when the catalog changes. New models arrive, prices change, and older models enter their retirement windows. The right answer last quarter may not be the right answer now.

Then pull the other levers

Model choice comes first because it changes the price of everything else. After that, the familiar levers still matter:

For the full picture across infrastructure, AI services, data platforms, and agents, see our guide to AI cost optimization. When the model calls run inside agents, see how to find the most expensive bottleneck in production AI agents.

A checklist for LLM workloads on the hyperscalers

  • Do you know which model every production task runs on, and when that choice was last reviewed?
  • Are any of those models in a Legacy or retirement window?
  • Is any task running on a large model that a smaller one could handle?
  • Do you measure cost per successful task, with success defined per task?
  • Have you evaluated each task on several current models, from more than one provider?
  • Do you re-run evaluations when new models arrive or prices change?
  • Is work that can wait running in batch?
  • Are repeated prompt prefixes cached?
  • Is provisioned capacity used only where load is steady?

How PointFive helps

PointFive sees how LLM workloads actually run across Amazon Bedrock, Microsoft Foundry, and Gemini Enterprise Agent Platform: usage and cost metrics, optional invocation logs, and the activity of AI agents through TokenShift. Invocation logs and agent activity are analyzed in the customer's environment; the raw data is not pulled out of it.

From there, PointFive detects the patterns described above, including outdated or suboptimal model choices on Bedrock and Azure OpenAI, outdated Claude model versions, missing batch and caching, and mismatched provisioned capacity. Cost per successful task is measured in several ways, with each customer defining what success means for its own tasks.

Read more about AI cost optimization with PointFive.

The bottom line

The hyperscalers give you more models than ever, and they keep adding better and cheaper ones. That only helps if you keep choosing. Stop treating model selection as a launch decision. Measure cost per successful task, evaluate every task on several models, and run the evaluation again whenever the catalog changes.