Updated September 17, 2026 against version 5 of the paper.
What the experiment measured
PointFive ran four Claude Code configurations across 103 tasks, seven repositories, and three models. The campaign executed 2,908 runs; 2,848 were included in the primary analysis.
| Configuration | Delivered tool-output token change | Pooled billed-cost change |
|---|---|---|
| Unmodified Claude Code | Baseline | Baseline |
| RTK v0.44.1 | -1.3% | -2.7% |
| RTK-ML, PointFive experimental build | -38.4% | +6.8% |
| Headroom v0.27.0 | Not measured | +48.4% |
These are paired billed-cost changes against the baseline, not cost per successful task. RTK's pooled estimate was below zero, but its holdout-only confidence interval crossed zero. The paper reports the intervals, exclusions, and success metrics separately.
What this means for an evaluation
Token reduction alone is insufficient evidence of a lower bill. Compare provider charges and task outcomes across representative workloads. Include retries, caching, additional model calls, and any service fees. A useful experiment holds the task and acceptance criteria constant and checks whether the work still succeeds.
Scope and disclosure
The results apply to the tested configurations. They do not establish a universal savings ceiling or describe every compression tool. RTK-ML was an experimental PointFive build; it was not unmodified RTK. This is PointFive-authored research, not an independent evaluation of PointFive's products.