Updated September 17, 2026. Based on version 5 of the paper.

Token reduction is not cost reduction

The tested compression configurations changed billed cost differently. Compare the bill, task outcomes, and the scope of the experiment.

What the experiment measured

PointFive ran four Claude Code configurations across 103 tasks, seven repositories, and three models. The campaign executed 2,908 runs; 2,848 were included in the primary analysis.

ConfigurationDelivered tool-output token changePooled billed-cost change
Unmodified Claude CodeBaselineBaseline
RTK v0.44.1-1.3%-2.7%
RTK-ML, PointFive experimental build-38.4%+6.8%
Headroom v0.27.0Not measured+48.4%

These are paired billed-cost changes against the baseline, not cost per successful task. RTK's pooled estimate was below zero, but its holdout-only confidence interval crossed zero. The paper reports intervals, exclusions, and success metrics separately.

What this means for an evaluation

Token reduction alone is insufficient evidence of a lower bill. Compare provider charges and task outcomes across representative workloads. Include retries, caching, additional model calls, and service fees. Hold the task and acceptance criteria constant and check whether the work still succeeds.

Scope and disclosure

The results apply to the tested configurations. They do not establish a universal savings ceiling or describe every compression tool. RTK-ML was an experimental PointFive build, not unmodified RTK. This is PointFive-authored research, not an independent evaluation of PointFive's products.

Read the evidence