Skip to content
Back to Press Releases

PointFive Research Finds Cutting AI Tokens Can Actually Increase Costs

PointFive Team

Updated September 17, 2026 to the version 5 paper: the primary campaign executed 2,908 runs; 2,848 runs were analyzed. Pooled billed-cost changes are distinct from cost per successful execution and from holdout-only estimates. See the versioned methodology and results.

A PointFive experiment compares billed costs across Claude Code configurations, with results scoped to the tested workload

NEW YORK - August 6, 2026 - New research from AI Efficiency OS company PointFive challenges a common assumption about AI cost management: that using fewer tokens automatically means spending less. An empirical study titled Token Reduction Is Not Cost Reduction charts the path toward more strategic AI budget allocation and use at scale. The study ran 2,908 paid coding sessions (2,848 analyzed) and found that reducing tool-output tokens by 38.4% actually increased billed costs by 6.8%, highlighting the disconnect between token consumption and the true cost of AI workloads.

Managing the cost of AI is the top forward-looking priority in the FinOps Foundation's 2026 State of FinOps survey. In McKinsey's May 2026 Enterprise AI FinOps Survey, 93% of respondents reported exceeding their AI budgets. The survey had 75 qualified respondents from 120 enterprise participants; the results were published in July. PointFive's report sheds light on what is happening between AI usage visibility and budget expenditure and why the most logical attempt at reducing AI costs - using fewer tokens - is failing.

"AI efficiency is a brand-new field, and most teams are making cost decisions without evidence," said Alon Arvatz, CEO of PointFive. "We did this research to understand how AI workloads actually behave, whether in the cloud or in the coding agent, and how we can bring those efficiencies to our customers. This is the first of many research reports we will publish. Research-backed evidence is how we build our product, and we look forward to sharing more discoveries."

Proper AI Measurement

PointFive's study ran 2,908 paid coding sessions (2,848 analyzed) across 103 software tasks, seven repositories and three models in Claude Code. The study reported the following findings for its tested configurations, not every coding-agent workload:

  • Results are specific to the tested workload and configurations.
  • RTK, an open-source tool, improved billed cost by 2.7%. However, the holdout-only interval crossed zero.
  • An experimental aggressive technique developed by PointFive researchers (RTK-ML) cut 38.4% of tool-output tokens but resulted in a 6.8% increase in billed costs.
  • Headroom, a widely used open-source third-party tool, increased billed cost by 48.4% in the tested setup. See the paper for task-success results and confidence intervals.

Why context changes can affect the bill

  1. Aggressively removing context can force an agent to spend additional steps finding information it previously had available and is already stored in the cache. The saving comes back as a larger bill.
  2. Our hypothesis is that compressed context makes the model "less confident" in the results, due to the drift it creates from what it was trained on. This leads to multiple turns of thinking and tool calling to justify its assumption.

Scope of the results

The billed-cost results above come from the study's primary campaign. The paper's component-attribution estimates use a separate corpus and modeling assumptions, so they are not a measured, universal breakdown of AI spending.

The Benchmark Is Open Source

The full paper, Token Reduction Is Not Cost Reduction, is a free download on arXiv, and the AI Efficiency Benchmark behind it is open source at github.com/PointFiveLabs/ai-efficiency-benchmark, so any savings claim, including PointFive's own, can be run through it. For more information, read the paper, download the research summary, or read the launch blog. Further research in the series will follow, published the same way.

Methodology

The study was authored at PointFive and evaluates the company's own experimental build, an unmodified open-source compressor, alongside Headroom, a third-party tool. It is not independent research. Method, per-session data and limitations are published in full. Every cost figure was read from the actual provider bill rather than estimated from a token counter.

About PointFive

PointFive offers two separate products: PointFive OS for infrastructure efficiency and TokenShift for AI-agent visibility, governance, and optimization. Organizations like Fanatics, H&M, Hertz, Nubank, Citizens Bank and other Fortune 500 companies trust PointFive to create greater transparency into AI spend to uncover opportunities to maximize AI ROI. Learn more at https://pointfive.co.

Media contact: [email protected] / +1 339 242 0393