Skip to content
Back to Press Releases
Press Release

PointFive Research Finds Cutting AI Tokens Can Actually Increase Costs

PointFive TeamAugust 6, 2026

Benchmarking across nearly 3,000 coding sessions, report finds ~80% of the AI bill is spent re-sending cached context, not generating answers

NEW YORK – August 6, 2026 – New research from AI Efficiency OS company PointFive challenges a common assumption about AI cost management: that using fewer tokens automatically means spending less. An empirical study titled Token Reduction Is Not Cost Reduction charts the path toward more strategic AI budget allocation and use at scale. The study reviewed 2,908 paid coding sessions and found that reducing tool-output tokens by 38.4% actually increased billed costs by 6.8%, highlighting the disconnect between token consumption and the true cost of AI workloads.

Managing the cost of AI is the top forward-looking priority in the FinOps Foundation's 2026 State of FinOps survey. According to McKinsey's July 2026 Enterprise AI FinOps Survey, 93% of companies have exceeded their AI budgets while only 26% of companies have real-time visibility into the cost of running AI (KPMG). PointFive's report sheds light on what is happening between AI usage visibility and budget expenditure and why the most logical attempt at reducing AI costs – using fewer tokens – is failing.

"AI efficiency is a brand-new field, and most teams are making cost decisions without evidence," said Alon Arvatz, CEO of PointFive. "We did this research to understand how AI workloads actually behave, whether in the cloud or in the coding agent, and how we can bring those efficiencies to our customers. This is the first of many research reports we will publish. Research-backed evidence is how we build our product, and we look forward to sharing more discoveries."

Proper AI Measurement

PointFive's report spanned across 2,908 paid coding sessions, and incorporated 103 specific software tasks, using seven code repositories and tested against three models using Claude Code. The study produced surprising findings:

  • Nearly 94% of agent costs are tied to harness system instructions and concealed reasoning. Conventional compression tools can only influence the user-visible input tokens that represent the remaining 6% of costs.
  • RTK, an open-source tool, improved cost per successful task by 2.9%. However, the results weren't statistically significant.
  • An experimental aggressive technique developed by PointFive researchers (RTK-ML) cut 38.4% of tool-output tokens but resulted in a 6.8% increase in billed costs.
  • Headroom, a widely used open-source third-party tool, increased cost per successful task by 46.4% while showing no measurable improvement in task success.

Why Cutting Prompts Raises – Not Saves – the AI Bill

  1. Aggressively removing context can force an agent to spend additional steps finding information it previously had available and is already stored in the cache. The saving comes back as a larger bill.
  2. Our hypothesis is that compressed context makes the model "less confident" in the results, due to the drift it creates from what it was trained on. This leads to multiple turns of thinking and tool calling to justify its assumption.

The Path to Lower AI Costs: Visibility and Governance

"The research findings propose that there is more value in governance controls, selecting the right coding agent for the task, setting budgets, limiting personal use and model routing. Prompt compression still holds significant potential with cost reduction if directed to influence hidden thinking that accounts for nearly 20% of the bill. We will continue to focus our research on where there are optimal efficiencies for AI usage." Alon Arvatz, CEO of PointFive

The Benchmark Is Open Source

The full paper, Token Reduction Is Not Cost Reduction, is a free download on arXiv, and the AI Efficiency Benchmark behind it is open source at github.com/PointFiveLabs/ai-efficiency-benchmark, so any savings claim, including PointFive's own, can be run through it. For more information, read the paper, download the research summary, or read the launch blog. Further research in the series will follow, published the same way.

Methodology

The study was authored at PointFive and evaluates the company's own experimental build, an unmodified open-source compressor, alongside Headroom, a third-party tool. It is not independent research. Method, per-session data and limitations are published in full. Every cost figure was read from the actual provider bill rather than estimated from a token counter.

About PointFive

PointFive is the AI Efficiency OS, from the cloud to coding agents. It continuously improves efficiency across cloud infrastructure, data platforms, AI workloads, and coding agents, helping teams understand spend in plain language and ship optimizations at scale. Organizations like Fanatics, H&M, Hertz, Nubank, Citizens Bank and other Fortune 500 companies trust PointFive to create greater transparency into AI spend to uncover opportunities to maximize AI ROI. Learn more at https://pointfive.co.

Media contact: [email protected] +1 (339) 242-0393