Skip to content
PointFive
All articles

Safe Automated Remediation: Responsibility Without the Burden

Gal Ben-DavidLinkedInCo-founder & CPO, PointFive7 min read

Every FinOps team knows the backlog. Hundreds of validated savings opportunities, each one sitting in a ticket, waiting for an engineer who has better things to do. The waste is not hidden. It is just not fixed.

The obvious answer is automation, and the obvious objection is safety. One bad change in production can cost more than a year of the waste it was meant to remove. So most teams stay manual, and most of the savings stay on paper.

That trade-off is false. Automated remediation can be safe, and the way to get there is not to give agents more freedom. It is to change who carries what.

What "safe" actually means

Safe automated remediation means a person takes the responsibility, but not the burden.

Someone still owns the decision. Someone still says yes. But everything around that decision, the investigation, the validation, the plan, the paperwork, the follow-up, is done by AI agents. The owner gets a short explanation, the evidence, and a single button. The agent takes care of the rest.

Most remediation programs get this backwards. They hand the owner the burden (a ticket with a link to a dashboard and a vague recommendation) and keep the responsibility diffuse (nobody is quite sure who should approve it). The result is the backlog.

The real risk: agents are not deterministic

The fear of automated remediation is reasonable. AI agents are not deterministic by nature. Ask the same agent to fix the same problem twice and you may get two different approaches. Give an agent broad autonomy over production infrastructure and you are betting your environment on its judgment in that moment.

The answer is not to avoid agents. It is to change what you ask them to produce.

Instead of asking an agent to fix a problem, ask it to produce a deterministic playbook for that particular problem: an explicit, step-by-step plan that does the same thing every time it runs. Then judge that plan in several different ways before a human ever sees it, and let a human approve it. The agent's creativity goes into understanding the problem and designing the fix. The execution is predictable.

Done this way, you will not see an agent get out of control, because the agent is never improvising against your infrastructure. It is executing a plan that was checked and approved.

The safe remediation pipeline

Here is what a safe automated remediation flow looks like, end to end. AI agents do every step except the decision.

1. Validate the inefficiency

Before anything else, confirm the finding is real. Use current usage, configuration, and cost data, not the snapshot the recommendation was generated from. Resources change. A finding that was true last month may not be true today.

2. Investigate the root cause

An idle resource is a symptom. Find out why it is idle: a forgotten experiment, an over-provisioned default in a template, a workload that moved. The root cause decides what the right fix is, and whether the same waste will come back.

3. Plan a deterministic playbook

Turn the fix into an explicit plan: which resources, which changes, in what order, under what conditions. The plan should be specific enough that running it twice produces the same result.

4. Judge the plan in multiple ways

A single check is not enough. Judge the plan from several independent angles:

  • Policy and permissions: is this change allowed, and does it stay within the access the agent was granted?
  • Live validation: does the plan still hold against current usage, dependencies, and configuration?
  • Dry run: what would the change actually do if executed?
  • LLM as a judge: have a separate model review the plan for errors, missing steps, and unintended consequences.

Only a plan that passes every check moves forward.

5. Find the owner and reach them where they are

The right approver is the person who owns the resource, not the FinOps team. The agent identifies that owner and tracks them down wherever they work: a Slack message, a GitHub pull request, a task in the platform. Nobody should have to log in to a separate console to do the right thing.

6. Communicate briefly

The message to the owner should be short and complete: what the problem is, what it costs, what the fix is, and what the resolution process looks like. Evidence attached, ROI stated, one decision requested. If the owner has to investigate before they can approve, the agent has not done its job.

7. One-click approval, then execute

The owner approves with one click. The agent executes the approved playbook using the write access the customer authorized, scoped to what the plan needs.

8. Monitor, report, and roll back if needed

The job is not done when the change is applied. The agent monitors the result, tracks the realized ROI against the estimate, and reports on it. If something goes wrong, a rollback option is ready. And every step, from the first validation to the final report, is logged.

Approval is a setting, not a constraint

How much human approval a process needs is a choice, and it should be made per process, not once for the whole program.

Many teams start with per-change approval: every execution waits for the owner's click. That is the right place to start while trust is being built.

Once a team has seen the same playbook run safely again and again, many choose standing approval: they approve the playbook once, as a policy, and let agents apply it whenever the conditions match. That is still human control. It is just exercised once, at the level of the playbook, instead of a hundred times at the level of the resource.

And for processes the team fully trusts, approval can be dropped altogether. The agent validates, plans, judges, executes, monitors, and reports on its own, and people read the results.

Some changes should always keep a human in the loop, for example where compliance requires an explicit approval. That does not mean those changes have to be slow. The goal is to reduce the friction of taking action to almost zero: provide the evidence, find the owner, show it concisely, and get the approval.

How to roll it out

  1. Start with clear, reversible waste. Idle resources, unattached storage, non-production environments that run around the clock. See our guide to automatically shutting down non-production environments.
  2. Require per-change approval at first. Let owners see the evidence, the plan, and the result for themselves.
  3. Measure realized savings, not estimates. Monitoring and reporting are what turn a fix into proof.
  4. Graduate trusted playbooks to standing approval. When a playbook has a clean track record, approve it once, and drop approval entirely for the processes you trust most.
  5. Keep compliance-bound changes behind explicit approval, and keep making that approval easier.

A checklist for evaluating remediation automation

If you are assessing a platform or building your own flow, ask:

  • Does it validate the finding against live data before acting?
  • Does it find the root cause, not just the symptom?
  • Does the agent produce a deterministic plan, or does it improvise at execution time?
  • How is the plan judged before a human sees it, and by how many independent checks?
  • Does it identify the owner and reach them in Slack, GitHub, or wherever they work?
  • Is approval one click, and can you choose per process between per-change approval, standing approval, and none?
  • Does it monitor the result, report realized ROI, and offer rollback?
  • Is every step logged?

How PointFive does it

In PointFive OS, Coworkers run this pipeline, and they are flexible enough to run any process a team defines. AI agents validate inefficiencies, investigate the root cause, find the relevant owner, and reach out to them with the problem, the ROI, and the resolution process. They execute the fix using customer-authorized write access, then monitor the results, report on them, and provide a rollback option if something goes wrong. Everything is logged.

Approval is optional and set per process: one click per change, a standing approval for a trusted playbook, or none at all. When approval is required, it happens wherever the owner is: in Slack, on GitHub, or in the platform.

For teams that prefer to keep changes in code, prompt remediation packages the finding and its context for an AI coding agent, so engineers generate and review the fix in their own IDE and deploy it through their own workflow. Read more about both on the Agentic Remediation page.

The approach works in production today. At least 20% of the waste PointFive detects is now handled completely autonomously by AI agents, under the approval policies each customer sets.

This is part of a broader shift in how FinOps teams work. For the full picture, see what agentic FinOps means.

The bottom line

Automated remediation is not a question of trusting agents with your infrastructure. It is a question of designing the process so that you do not have to. Let agents carry the burden: validate, investigate, plan, judge, find the owner, explain, execute, monitor, and report. Let people carry the responsibility: one clear decision, made with the evidence in front of them.

That is how the backlog finally gets fixed.