Skip to content

Introducing PointFive Labs

Open, reproducible research on how AI infrastructure spend is actually incurred.

Something has gone wrong in how our industry talks about AI cost.

Two years ago the questions were about cloud. They were hard, but they were answerable, and there was a decade of accumulated practice behind the answers. The questions companies are asking now are different. What does a coding agent actually cost to run. Does compressing context save money. Which model should handle which task. When is a cache worth paying for. Whether the bill is going up because usage grew or because something is being done inefficiently.

These are empirical questions. Almost nobody is treating them empirically.

What exists instead is a great deal of confident assertion, most of it from vendors, most of it unverifiable, and a fair amount of it repeated so often it now reads as established fact. Buyers are making eight and nine figure decisions on top of claims nobody has tested. We have been on the receiving end of that too, and we have made assumptions internally that turned out to be wrong when we finally measured them.

So we are formalizing something we have already been doing. Two papers went out over the summer, both with the benchmarks released alongside them. What changes today is that the work has a name, a permanent home, and two people leading it.

PointFive Labs

PointFive Labs is the research arm of PointFive. It publishes open, reproducible research on how AI infrastructure spend is actually incurred, and releases the benchmarks and data behind its findings.

Three commitments, which matter more than the name:

  1. We publish in the open.Not gated, not summarized into a landing page, not held back for a launch moment.
  2. We publish what can be reproduced.Where a finding can be checked, we release what is needed to check it. A cost claim nobody can re-derive is marketing, whatever it is labeled.
  3. We publish findings that complicate our own story.This is the one that will be tested over time, and it is the one we care most about keeping. Research that only ever confirms the product roadmap is not research.

Who runs it

Amir Hozez

Amir Hozez

Co-Head of PointFive Labs

Amir is co-founder and CTO of PointFive. Before starting the company he was VP of R&D at IntSights, which was acquired by Rapid7. He is a co-author of both Labs papers.

Sarel Weinberger

Sarel Weinberger

Co-Head of PointFive Labs

Sarel has founded several AI companies and led AI at PwC. He lectures at Bar-Ilan University. He is lead author of both Labs papers and maintains the open benchmarks behind them.

I picked the two of them for a specific reason. Labs needed to be run by the people doing the measuring, not by people describing the measuring afterwards. Amir and Sarel are in the data daily. That means Labs will sometimes publish things that are inconvenient for us commercially, and it means what it publishes will be right.

The research so far

Labs opened in July 2026 with Token Reduction Is Not Cost Reduction: An Empirical Study of End-to-End Efficiency in API-Based Coding Agents. It examines whether the standard technique for lowering AI cost actually lowers AI cost.

Token Reduction Is Not Cost Reduction

July 2026

Delivered tool-output tokens, at the most aggressive compression setting

38.4% lower

Billed cost, same setting

6.8% higher

Token reduction turned out to be close to uncorrelated with cost reduction. The paper argues for measuring cost per successful task instead. The open benchmark was released with it so the result can be checked rather than believed.

The second paper followed in August. Same Task, Different Work: Prompt-Induced Waste in Coding Agents asks how much the wording of a prompt changes what a coding agent costs to run.

Same Task, Different Work

August 2026

Reasoning, when asked to consider multiple approaches, across six models

2.4x to 7.4x

Change in success rate from that instruction

None

Demanding certainty produced verification loops costing up to 18 times the baseline. Across 4,644 runs, the same task done the same way cost wildly different amounts depending on how it was asked.

Neither finding helps anyone sell anything, including us, which is precisely the point.

What happens next

Labs publishes new research on a routine basis. Some of it will be large, some of it will be a single measurement worth writing down. It will always come with the data.

If you are working on these questions inside your own organization and seeing something we are not, Amir and Sarel would like to hear from you.