An AI agent in production is a strange kind of software. It decides, at run time, how much work to do. It chooses which tools to call, how many steps to take, how much context to carry, and which model to send each step to. Every one of those decisions has a price, and none of them is fixed.
That is why agent costs surprise teams that are used to ordinary API costs. A single user request can become dozens of model calls. The same request can cost very different amounts on different days. And the bill grows in places nobody is looking.
There are many ways production agents go wrong. Most of them come down to the same root cause: agents that are built loosely and then left alone.
Where production agents waste money
1. Built for a vague job, not a precise task
The most common mistake happens before the first line of code. An agent is scoped as "handle customer questions" or "help with data requests" instead of a specific, well-defined task. A broad agent has to work out what the job is on every run. That costs tokens, steps, and wrong turns, and it makes the agent harder to measure, because there is no clear definition of done.
Agents that are built precisely for a task are cheaper, faster, and easier to improve.
2. Tools it does not need
Every tool you give an agent has a cost, even when it is never called. Tool definitions are typically sent to the model with each request, so a long tool list adds input tokens to every step. Worse, every extra tool is another option the model can choose wrongly, which leads to detours, retries, and longer runs.
Give an agent the tools its task requires, and nothing else.
3. AI for problems that have a deterministic answer
Not every step needs a model. Parsing a known format, applying a fixed business rule, looking up a value, routing by a simple condition: these have deterministic answers, and code will give the right one every time, for a fraction of the cost, in a fraction of the time.
When an agent uses a model for work that code could do, you pay for reasoning you did not need, and you add uncertainty where there was none. Use the agent for judgment. Use code for everything else.
4. One model for every step
Agents often send every step to the same large model, including the trivial ones: classifying an intent, extracting a field, formatting an output. Different steps need different capabilities. Routing simple steps to smaller models is one of the largest savings available in most agents. For how to choose, see cost per successful task.
5. Running blind
The most expensive mistake is the one that hides the others: not measuring. Many teams know what an agent costs in total, but not what each step costs, which steps fail, where retries happen, or which part of the run takes the longest. Without that profile, there is no way to know what to fix first.
Agents do not stay the same
Even a well-built agent will not behave the same way every time, and it will not behave the same way next quarter as it does today.
The underlying models are updated or replaced. The data the agent works on changes. Prompts get edited, tools get added, and users find new ways to ask for things. Each change can shift how many steps the agent takes, which tools it uses, and how often it fails. An agent that was efficient at launch can drift into an expensive one without a single line of its code changing.
That is why a one-time optimization pass is not enough.
The practice: monitor continuously, fix the worst bottleneck, repeat
Production agents need continuous monitoring and a simple, repeated loop:
- Profile. Measure cost, latency, failures, and retries per run and per step, not just in total.
- Find the most painful bottleneck. The step, tool, or path that wastes the most money or time.
- Fix it. Narrow the task, remove an unused tool, replace a step with deterministic code, route the step to a smaller model, or cache repeated context.
- Measure the result. Confirm the fix improved cost per successful task without hurting quality.
- Repeat. There is always a next bottleneck, and drift will create new ones.
The goal is not a perfect agent. It is an agent that gets a little better every cycle, and never quietly gets worse.
Guardrails that keep agents in budget
Monitoring tells you where the money goes. Guardrails keep a single bad run from becoming a large bill:
- Budgets and limits per task: cap the spend, steps, or tokens a single run can use.
- Model policies: decide which models each agent, team, or step is allowed to use.
- Routing per step: send each step to the model that fits it.
- Tool allowlists: block tools and frameworks an agent should not use.
- Kill switches: stop runaway loops before they burn through a budget.
- Approval for expensive actions: require a human to approve the actions that cost the most.
- Spend caps by team: keep each team's agents within its allocation.
For allocating agent costs to teams, products, and customers, see our framework for FinOps for AI agents.
A checklist for production agents
- Is each agent built for one precise task, with a clear definition of done?
- Does it have only the tools that task needs?
- Are deterministic steps handled by code instead of a model?
- Are simple steps routed to smaller models?
- Do you measure cost, latency, failures, and retries per step?
- Do you know the agent's most painful bottleneck right now?
- Do you track cost per successful task over time, to catch drift?
- Are there budgets, step limits, and a kill switch for runaway runs?
- Do expensive actions require approval?
How PointFive helps
TokenShift is the unified control plane for every AI agent. It provides visibility, governance, and optimization for AI usage across workstations, production workloads, and hosted agent environments.
TokenShift attributes AI consumption to applications, tasks, and types of work, so you can profile an agent run by run and find its bottlenecks. It covers the guardrails above: model policies, budgets and spend caps, limits, tool and framework blocking, kill switches, and approval for expensive actions. Smart Routing, an optional module, can select a different model for a task based on the task and your organization's policies. Agent activity is analyzed locally, in your environment. Only derived metadata and analysis results are sent to PointFive, and prompt content is not.
PointFive OS covers the cloud, data, and AI services those agents run on, so agent costs and the infrastructure behind them can be seen together.
The bottom line
Production agents are not set-and-forget software. Build them precisely for the task, give them only the tools they need, let code handle what code can, and route each step to the right model. Then keep watching: profile every run, fix the most painful bottleneck, measure, and repeat. And expect the agent to change, because it will.