This website uses cookies

Read our Privacy policy and Terms of use for more information.

You are going to be shown a settings screen this quarter, and the entire argument will be sitting in a dropdown.

The screen belongs to an AI agent that watches your production estate. The dropdown decides what it is allowed to do without asking first. One option investigates and writes up what it found. Another proposes a fix and waits. Another simply gets on with it.

Whoever picks from that list has made a governance decision, whether or not anybody in the room describes it that way. In most firms I have looked at, nobody has worked out who that person is supposed to be.

Nobody is actually selling you an unsupervised agent

I went looking for the pitch everyone complains about, the one that promises to remove the human entirely, and it is not really there.

Dynatrace's Autonomous Operations release of 27 July 2026 leads with the phrase "human oversight and governance". Datadog's Bits AI SRE investigates and recommends, and does not remediate at all. PagerDuty wrapped its March 2026 agentic release in the language of agent governance. Resolve AI's own copy says engineers step in to direct and to act. Even DataAgent, the newest and loudest entrant, sells actions the user defines.

So the story is not vendors overselling autonomy while practitioners push back. It is more awkward than that, and considerably more expensive.

Autonomy shipped as a setting. Settings have owners.

What Microsoft put in the documentation

Microsoft's own documentation for Azure SRE Agent is unusually plain. Whether the agent applies a mitigation on its own or waits for approval depends, in their words, on your configured run mode. The same documentation says your team defines the boundaries even for fully automated workflows. Individual tools carry Allow, Ask and Deny.

Read that as a buyer and it is a feature list. Read it as a regulated firm and it is a transfer of liability, documented and agreed in advance.

Autonomy shipped as a setting. Settings have owners.

Which matters more than it did last year. On 10 July 2026 HM Treasury designated Microsoft Ireland Operations Ltd a Critical Third Party, with Bank of England, PRA and FCA oversight following on 13 July. The hyperscaler is now inside the UK regulatory perimeter. And it ships a production agent with an autonomous run mode whose switch belongs to the customer.

If you are under SM&CR, that switch has a name attached to it. Not a supplier's name.

Somebody in your building.

I suspect the first firm to work this out will do it in a supervisory conversation rather than in a design review, which is the expensive order to learn things in.

Write the AI agent autonomy ladder before you watch the demo

The word autonomous has been stretched across four genuinely different products, and the demo will not tell you which one you are looking at. So bring your own vocabulary.

Read-only. The agent watches, correlates and explains. It will write you the incident timeline. It touches nothing.

Advised. It recommends. Here is what I think broke, here is what I would do about it, here is who I would escalate to. A person still does every part of it.

Approved. It can act, but each action waits for someone to say yes.

Autonomous. It acts alone, inside defined limits, with guardrails.

Four rungs. Read-only, Advised, Approved, Autonomous.

The four rungs of AI agent autonomy, rising left to right as a staircase. Rung 1 Read-only watches and explains and touches nothing. Rung 2 Advised recommends while a person does every part. Rung 3 Approved acts only when someone says yes. Rung 4 Autonomous, marked as the subject, acts alone inside limits and is the setting in the dropdown. Arrows between the rungs mark the evidence needed to climb. A band beneath reads: accountability does not climb, it is the same named human at rung 1 and rung 4.

A rung is earned on evidence, not bought in a demo. The accountability sits in the same place on all four.

Choosing a rung is the easy half of the work. The harder half is writing down what a system has to demonstrate before it is allowed to climb to the next one, and doing that before a vendor is in the room shaping the criteria for you.

Nobody gives a new starter production write access on day one

I want to borrow something from an unlikely place. The clearest staged-rollout discipline I have read this year was not in an enterprise architecture document at all. It was a productivity walkthrough for building a personal agent to handle somebody's email and calendar. Start read-only. Prove the output over repeated runs. Add one permission at a time, and only after the last one has earned its keep.

That is not sophisticated. It is what any decent manager already does with a new hire.

Nobody gives a new starter write access to production on their first morning, and it has nothing to do with their CV. They may well be better than half the team. You still make them earn it, because trust in an operational setting is a record rather than an assessment, and the record takes time to build.

Give an agent less than that and you have decided a probabilistic system deserves more benefit of the doubt than a person with references. I have yet to hear anyone defend that out loud.

Enterprise pilots skip the ramp constantly, and the reason is almost never technical. Something has to show value before the quarter closes, so the permissions land on day one and the evidence is meant to catch up later. It rarely catches up.

Which is where I have to update my own advice. For years the rule I gave people was to automate everything they could, and that rule has a failure mode I have watched land more than once. You can automate the action. You cannot automate the accountability. When it goes wrong, and one day it does, a machine cannot sit in the review and own it. So the rule now reads: automate the action, never the accountability, and the gate on anything with real blast radius stays exactly where accountability already lives.

Going slowly is also a decision

I do not want any of this read as a case for permanent distrust, because that is its own quiet failure.

You can automate the action. You cannot automate the accountability.

A team that keeps every agent on read-only for two years has not managed risk. It has refused to gather the evidence that would let it move, and it will end up slower than its competitors for reasons it will describe internally as prudence. The ramp is the point.

Not the brake.

But the evidence for climbing carefully is stronger than the evidence for climbing fast. Traversal publishes 82% root cause accuracy at a Fortune 100 financial services customer, on a platform that converts diagnosis into action. That is a good number. It is also one diagnosis in five that is wrong, and if a wrong one acts, the outcome is yours. In April 2026 PocketOS described a Cursor agent that was remediating a credential mismatch, deleted a production database and its backups in nine seconds, and then wrote a confession enumerating the rules it had broken. Gartner put a figure on the pattern on 26 May 2026: 40% of enterprises demoting or decommissioning autonomous agents by 2027 over governance gaps found only after a production incident.

The most instructive case is not an agent at all. When AWS us-east-1 failed on 20 October 2025, AWS's own post-event summary records that recovery required manual operator intervention, and that it disabled the DNS Planner and Enactor automation worldwide. Automation built to protect availability took it away, humans put it back, and the remedy was to switch the automation off.

None of that argues against automation. It argues for knowing exactly where the switch is, and who has their hand on it.

What to take into the next vendor meeting

Write the ladder down. Four rungs, and beside each one the evidence a system must produce before it climbs.

Read-only for a quarter, with an accuracy rate you actually tracked against the timelines your own engineers wrote. Advised until its recommendations match what those engineers would have done, most of the time. Approved before it ever acts alone, and only on actions that are bounded and reversible.

Then put a human name beside every rung.

That last line is the whole exercise. Build the ladder once and it works for every agent any vendor ever sells you, because you have stopped evaluating products and started evaluating whether a thing has earned a rung in your estate, with your systems, on your evidence.

And treat a change of run mode as a change. Record it, name it, date it, and write the reversal down. I have sat in enough change advisory boards to know what happens to a control that is easier to alter than the thing it controls. It gets altered.

If moving an agent from Approved to Autonomous is quicker in your organisation than shipping a config change to a payments service, you have not made a governance decision at all. You have made a settings change that quietly behaves like one.

Ask who signed for it. If the honest answer is that nobody did, that is your finding, and you found it before the incident did.

 

Work with me

An honest read on what your observability is actually doing.

If you lead observability in a regulated enterprise, I run a fixed-scope Observability Assessment for senior IT and engineering leaders. It ends in a written roadmap and a readout, not a sales deck.

See how it works →

Not ready to talk? Start with a free chapter of Metrics & Mayhem.