Guardrails that cost more than inference
Per-text-unit assessment costs can exceed the inference cost they protect.
What Labs investigates
Labs compares guardrail charges with inference costs to surface the imbalance.
Research into the infrastructure inefficiencies that billing data alone cannot explain.
Four detection examples from the Q1 2026 update. Each starts with a specific behavior and investigates its cost.
Per-text-unit assessment costs can exceed the inference cost they protect.
Labs compares guardrail charges with inference costs to surface the imbalance.
Intermediate tables can be rebuilt, but still incur Fail-safe and extended Time Travel charges.
Labs traces table lineage to find data written exclusively from upstream sources.
Content-safety rejections can leave you paying for input processing with no useful output.
Labs examines rejection patterns and their associated token costs.
Burstable instances can cost more than a comparable non-burstable fleet.
Labs aggregates CPU-credit overages across auto-scaling groups.
Configuration, usage, and cost analysis across cloud, data, and AI.
AWS EC2, EBS, EFS, S3, Lambda, OpenSearch; Azure Storage, Site Recovery, App Configuration; GCP Dataflow, BigQuery, Compute
ElastiCache provisioned and serverless; SQS standard and FIFO patterns; DynamoDB and RDS Multi-AZ behavior
Snowflake warehouses, table lineage, and ingestion; Databricks and BigQuery coverage in active expansion
Bedrock and SageMaker; Azure OpenAI deployments and PTU economics; Routing, caching, guardrails, model selection, and GPU utilization
PointFive Labs is the in-house research team behind DeepWaste detections.
Research begins with real configurations, usage patterns, and workload behavior.
Researchers apply an investigative approach drawn from cybersecurity, examining configuration, utilization, and lineage.
Detections include a cost model to quantify potential impact. Teams validate savings after implementation.
Labs also publishes open, reproducible research on how AI infrastructure spend is incurred, with the benchmarks and data behind each finding.
July 2026
An empirical study of end-to-end efficiency in API-based coding agents.
August 2026
How prompt wording changes reasoning, effort, and end-to-end cost for coding agents.
Benchmarks
The benchmark and data behind both papers, released so the results can be checked.
The Q1 2026 report on new detections across cloud, data, and AI.
Read the full reportReported for the Q1 2026 customer cohort, including ~$8M in AI opportunities. Addressable savings are estimates, not realized savings or a forecast for every environment.