Skip to content
Back to Blog
AI

Why Companies Overspend on AI: The 2026 Cost Visibility Gap

Dave AndersonJuly 29, 20267 min read

Companies overspend on AI because AI spend moved to consumption pricing faster than anyone built the instrumentation to read it. Worldwide AI spending reaches $2.59 trillion in 2026, up 47% year over year. Full, real-time visibility into what that AI costs to operate sits at 26%. The gap between those two numbers is the whole problem.

Four numbers that frame it:

NumberWhat it measuresSource
79%Enterprises with AI cost overruns in the past 12 monthsDoiT / Sapio Research, Feb 2026, n=500
26%Organizations with full, real-time visibility into what their AI costs to operateKPMG AI Pulse Q2 2026, n=204
98%FinOps practitioners now managing AI costs — up from 31% in 2024FinOps Foundation, 2026, n=1,192
36%Organizations with direct token or usage controls in placeKPMG AI Pulse Q2 2026, n=204

The 36-second version: three contestants, one question about what their AI actually costs, no correct answers.

How Fast Did AI Spending Actually Grow?

Faster than the tooling around it. AI spending grew 47% in a single year to $2.59 trillion, and the model layer grew faster than the total.

FigureWhat it coversSource
$2.59TWorldwide AI spending in 2026, up 47% year over yearGartner
GenAI model spend more than doubled in 2026 — growth above 110%Gartner
$453BAI software in 2026, up from $282.8B in 2025Gartner
24×Forecast token consumption growth, to 120 quadrillion tokens per month by 2030Goldman Sachs
40%CAGR in daily LLM queries, reaching 11 billion by 2030Goldman Sachs

Cost per token is falling — 60–70% per year for inference, per Goldman Sachs. Unit price down, total bill up. Volume is winning by a wide margin, and volume is the variable nobody is metering.

Agents are why. An agentic workflow triggers 10–20 model calls per user task where a chatbot triggered one. The same request costs an order of magnitude more to serve, and the bill arrives shaped like usage, not like a licence.

How Many Companies Know Their Full AI Costs?

Roughly a quarter, at best. Every credible 2026 survey lands in the same band.

ShareFindingSource
26%Have full, real-time visibility into what their AI systems cost to operateKPMG AI Pulse Q2 2026, n=204
42%Report only partial visibility into AI spendingKPMG Global AI Pulse Q2 2026, n=2,145
33%Cite limited understanding of AI cost structures, tokens includedKPMG Global AI Pulse Q2 2026
15%Forecast AI costs within 10% of actual — nearly one in four miss by more than 50%Mavvrik / Benchmarkit, n=372
15%Of finance leaders can calculate AI ROI without significant bottlenecksDoiT / Sapio
36%Name lack of clear financial attribution as a core barrierDoiT / Sapio

The sharpest number in the KPMG data is the distance between watching and controlling. 66% maintain AI cost monitoring dashboards. 61% have cost reviews in their approval process. Only 36% have direct token or usage controls.

That is a dashboard problem, not an awareness problem. Two-thirds of organizations are looking at AI spend. A third can act on it. Effort went up; the ability to intervene did not.

Attribution is the specific failure. 55% of organizations place AI spend accountability with technology. 53% place it with finance. The numbers exceed 100% because both groups think it is theirs, which in practice means it is nobody's.

Are Companies Really Blowing Their AI Budgets?

Yes — and the honest number is between 68% and 79%, not the 93% figure circulating in feeds. Two well-attributed surveys:

  • 79% of enterprises had AI cost overruns in the past 12 months — DoiT, surveyed by Sapio Research, February 2026, 500 finance leaders at US and UK organizations with 1,000+ employees, ±4.4pp
  • 68% say at least some AI initiatives ran over budget; 33% say overruns happen mostly or always; only 9% say more than three-quarters of AI initiatives delivered measurable financial return — WitnessAI, July 2026, n=300 executives

The counter-intuitive finding sits inside the DoiT data. Overruns get worse with FinOps maturity.

FinOps maturityShare with overrunsMean overrun
Very mature / leading edge89%30.9%
Early stage69%16.1%

Mature teams are not worse at cost control. They are better at measuring, so they can see the overrun they are having. Early-stage teams are running the same overrun with the lights off.

The Receipts

Named, dated, public.

Uber exhausted its entire 2026 AI coding budget by April — four months in. Coding-agent adoption went from 32% to 84% of a roughly 5,000-engineer org in about a month, and 95% of engineers now use AI tools monthly. Average spend ran $150–250 per engineer per month, with heavy users reported near $2,000 a month and one two-hour demo session at $1,200.

Microsoft ended its coding-agent licences six months after starting the pilot.

Neither is a story about a badly run engineering org. Both are stories about a pricing model that changed under the floor: a flat seat licence made token spend invisible because the price did not move with usage. Consumption pricing moves with usage by definition, and almost nothing in the stack was built to watch it.

Why Have FinOps Teams Pivoted to AI Costs?

Because their scope was redefined for them in 24 months. The FinOps Foundation's 2026 State of FinOps report — 1,192 practitioners, $83B+ in annual cloud spend represented — shows the fastest scope shift in the discipline's history.

YearPractitioners managing AI costs
202431%
202563%
202698%

Everything else in the report follows from that line:

  • "FinOps for AI" is the #1 forward-looking priority, ahead of scope expansion and organizational alignment
  • AI cost management is the #1 skillset teams plan to add in the next 12 months, with 58% prioritising it
  • The #1 tooling request is granular monitoring of AI spend — tokens, LLM requests, GPU utilisation

Read that last item as what it is. The people whose job is cost visibility are naming this as the capability they most lack.

There is a structural reason it landed on FinOps. GPU consumption, token billing, model retraining cycles and hybrid placement decisions all introduce financial volatility with an engineering root cause. Only one function was already fluent in both.

Where Does AI Spend Actually Hide?

Below the line your dashboard draws. Your cost and usage report is accurate and complete about charges, by service, by account — and says nothing about workloads. Seven places, in rough order of how often they go unmetered:

  1. Coding agents. Usage-based, per-engineer, spiky. An agent working for hours against a large monorepo consumes orders of magnitude more than an autocomplete accept. Rarely attributed to a team or a repo.
  2. Model API calls outside the platform team. Direct provider keys, expensed on cards. Invisible to the cloud bill entirely.
  3. Provisioned throughput bought for peak, run at trough. A flat committed charge with nothing wrong with it — and no signal in billing data that it is 7× the tokens actually consumed.
  4. Inference routing. Cross-region instead of in-region serving costs materially more per call for identical work. There is no resource to inspect; the waste is in a routing config.
  5. GPU fleets idling between training runs. Utilisation looks fine at the instance level. Nobody owns the gap between jobs.
  6. Retrieval and context inflation. RAG architectures expand context windows 3–5×. The cost lands on every call, forever, and nobody re-reads the retrieval config after launch.
  7. Shadow AI. The average enterprise runs around 14 distinct AI tools; IT knows about 4–5. Roughly half of employees use AI tools their employer never approved.

The common property: none of these are visible in a cost and usage report. A CUR describes charges. It does not describe workloads. No amount of skill applied to billing data will surface a bad inference route, because the information is not in the file.

What Should Companies Do About AI Costs?

Get one view across every coding agent and every AI service before optimising anything. Five steps — and the order matters more than the individual steps.

1. Inventory Every AI Surface, Including the Ones Not on the Cloud Bill

Coding agents, model APIs, direct provider keys, GPU fleets, vector stores, AI features inside SaaS you already pay for. If it consumes tokens or GPU-hours, it goes on the list. Most organizations find the list is two to three times longer than expected — the 14-tools-versus-4-known gap is the reason.

2. Instrument at the Unit, Not the Invoice

Tokens, requests, GPU-hours, agent-runs. Invoice granularity cannot answer which repo, which agent, which engineer, which model — and those are the only questions that lead to an action. This is the FinOps Foundation's #1 tooling request for a reason.

3. Attribute to an Owner Before a Cost Centre

55% say technology owns AI spend; 53% say finance. Pick one owner per surface, name them, and put the number in front of that name weekly. Shared accountability at 108% is what produced the overruns.

4. Read Configuration, Not Just Consumption

The highest-yield AI findings are configuration errors, not volume problems: provisioned throughput at 7× actual token use, cross-region inference routing, non-production clusters on on-demand pricing. Each is a single config change. None appear in a bill.

5. Close the Loop With a Merged Fix

A finding without a fix is gossip. Route each one to the owning team as a pull request with the evidence attached — the line, the commit, the delta — then verify the saving against billing after the merge. The window for this is closing: the share of organizations orchestrating multiple AI agents across workflows doubled from 9% to 18% in a single quarter, while token and usage controls sit at 36%. Every quarter without remediation adds surface area faster than it adds control.

Steps 1 and 2 are visibility. Steps 3 through 5 are what visibility is for. Most 2026 AI cost programmes stopped at step 2, produced a dashboard, and reported the overrun more precisely.

AI Cost Questions, Answered

What Percentage of Companies Overspend on AI?

79% of enterprises experienced AI cost overruns in the past 12 months (DoiT / Sapio Research, February 2026, 500 finance leaders). A separate July 2026 WitnessAI survey of 300 executives found 68% had at least some AI initiatives run over budget, with 33% reporting overruns mostly or always.

How Many Companies Have Full Visibility Into AI Costs?

26%. KPMG's Q2 2026 AI Pulse survey of 204 US business leaders at $1B+ revenue organizations found only 26% have full, real-time visibility into what their AI systems cost to operate. KPMG's parallel global survey of 2,145 leaders across 20 markets found 42% have only partial visibility into AI spending.

How Fast Is AI Spending Growing?

Worldwide AI spending reaches $2.59 trillion in 2026, up 47% year over year, per Gartner. GenAI model spend more than doubled. Token consumption is forecast to grow 24× to 120 quadrillion tokens per month between 2026 and 2030, per Goldman Sachs.

Why Are FinOps Teams Now Responsible for AI Costs?

Because AI spend behaves like engineering decisions priced by consumption, and FinOps was the only function already fluent in both. Practitioners managing AI costs went from 31% in 2024 to 63% in 2025 to 98% in 2026, per the FinOps Foundation. "FinOps for AI" is now the discipline's #1 forward-looking priority.

Why Can't Cloud Cost Tools See AI Waste?

Because cost and usage reports describe charges, not workloads. A provisioned-throughput deployment running at one-seventh of its capacity produces a flat, unremarkable charge. Cross-region inference routing costs more per call for identical work but creates no anomalous resource. The waste lives in configuration, which is not in the bill.

What Is the First Step to Controlling AI Costs?

Inventory every AI surface, including the ones not on the cloud bill — coding agents, direct provider API keys, GPU fleets, vector stores, AI features in existing SaaS. Instrumentation before optimisation: you cannot reduce a number you cannot attribute.

Do You See What I See?

Connect one account, read-only. PointFive reads configuration as well as consumption — across cloud, AI services, and coding agents — and returns each finding as a fix with the evidence attached.

Sources

All figures accurate as of July 2026. Survey figures are reported with sample size and fielding date so they can be checked.

About PointFive

PointFive is the AI Efficiency OS. By combining a real-time cloud and infrastructure data fabric with AI-driven detection and guided remediation, PointFive transforms efficiency from a reporting exercise into an operational discipline. Customers achieve sustained improvements in cost, performance, reliability, and engineering accountability, at scale.

To learn more, book a demo.