Updated September 17, 2026. Based on version 5 of the paper.
Token reduction is not cost reduction
The tested compression configurations changed billed cost differently. Compare the bill, task outcomes, and the scope of the experiment.
What the experiment measured
PointFive ran four Claude Code configurations across 103 tasks, seven repositories, and three models. The campaign executed 2,908 runs; 2,848 were included in the primary analysis.
| Configuration | Delivered tool-output token change | Pooled billed-cost change |
|---|---|---|
| Unmodified Claude Code | Baseline | Baseline |
| RTK v0.44.1 | -1.3% | -2.7% |
| RTK-ML, PointFive experimental build | -38.4% | +6.8% |
| Headroom v0.27.0 | Not measured | +48.4% |
These are paired billed-cost changes against the baseline, not cost per successful task. RTK's pooled estimate was below zero, but its holdout-only confidence interval crossed zero. The paper reports intervals, exclusions, and success metrics separately.
What this means for an evaluation
Token reduction alone is insufficient evidence of a lower bill. Compare provider charges and task outcomes across representative workloads. Include retries, caching, additional model calls, and service fees. Hold the task and acceptance criteria constant and check whether the work still succeeds.
Scope and disclosure
The results apply to the tested configurations. They do not establish a universal savings ceiling or describe every compression tool. RTK-ML was an experimental PointFive build, not unmodified RTK. This is PointFive-authored research, not an independent evaluation of PointFive's products.