The last change I signed off without fully understanding it went through on a Thursday evening. The room had thinned to four people and a speakerphone. I read the change record, I asked the two questions I always ask, and then I put my name against it, because someone had to, and because the name on the record is most of what the record is for.
I have thought about that evening a lot since reading what Dynatrace has built.
What Dynatrace actually shipped
In July, Dynatrace rolled out a sub-brand called Bluebox.ai to early-adopter customers. Founder and CTO Bernd Greifeneder described it to Beth Pariseau at TechTarget, in an interview published on 29 July 2026, in about as plain a sentence as a vendor ever offers. Bluebox takes "the human completely out of the loop of the entire SRE and DevOps cycle."
The design is coherent. Observability and SRE agents feed context, prompts and skills to coding agents. The coding agents instrument, build, test and deploy. When production misbehaves, the investigation and the remediation happen together, autonomously, and the fix goes back out. There is no approval step. There is no change record with a person's name on it.
That is a real, named, dated counter-position to something I have argued on this blog and on the podcast for a year: that the gate stays human. So I went looking for the point where it falls over.
It does not fall over. It is more interesting than that.
|
Taking the human out of the loop does not remove the gate. It relocates it. |
The same vendor, the same week, two vocabularies
Two days before that interview ran, on 27 July, Dynatrace put out a press release about a different part of the same platform. The language is the exact opposite.
It describes agents that resolve incidents "while maintaining the human oversight and governance enterprises require." Every action is "transparent, auditable, and governed." A new Cloud SRE Agent centralises findings into "a single auditable record." The release's own FAQ says trust comes from "transparent action logging, auditable records, and built-in human oversight capabilities that allow for human approval and intervention where required."
Same company. Same week. One product line sells governance and human approval as the reason to trust it. The other sells the removal of the human as the reason to buy it.
Greifeneder names the split himself, and it is the most useful thing in the story. His customers, he says, have gone bimodal. Mode one is human-led, with AI assisting people inside the processes they already run. Mode two is agent-led, and Bluebox is the sub-brand built for it.
So this is not a vendor caught contradicting itself. It is a vendor that has drawn the line most enterprises have not drawn yet, and put a product on each side of it.
Where the gate went
Ask what replaces a person's signature and Greifeneder answers directly. Pressed on how the Bluebox agents are kept honest, he names four things:
Deterministic algorithms wherever they can be used, with the probabilistic agents reasoning on top rather than underneath.
An AI lakehouse that turns "petabytes of production data into kilobytes," small volumes of accurate context, specifically to hold hallucination risk down.
Results captured in test and staging before anything ships, as an early feedback loop.
Feature flags used "far more heavily than before," with A/B tests and the ability to switch a feature off quickly.
Read that list again, but read it as a control rather than as a feature list. Curated evidence. A deterministic check. A pre-production gate. A reversal mechanism.
Those are the four things a competent change advisory board does on a good day. Dynatrace has not removed them. It has rebuilt every one of them in software and taken the meeting out.
Taking the human out of the loop does not remove the gate. It relocates it.
Four questions for a relocated gate
This is the part that matters if you lead any of this, and none of the questions are about the model.
Can you point at it? A human gate has a name and a timestamp, and a regulator or an incident review can be walked straight to both. A relocated gate is a set of confidence thresholds, a feature-flag policy, a staging environment and a context-curation pipeline. In most estates I have worked in, those four things live in different repositories, owned by three different teams, and nobody has ever been asked to describe them as a single control. If you cannot draw your gate on one page, you do not know where it is.
Does it hold when it is the thing that is wrong? The whole design assumes the deterministic layer is correct and the curated context is representative. That is a reasonable assumption right up until it is not. When a person is wrong, you get one bad change. When the gate is wrong, you get every change, at machine speed, until something downstream is unhappy enough to tell you.
What does reversal actually cost you? Feature flags and fast rollback carry more weight in this design than anything else, and they are the piece most organisations quietly do not have. Turning a flag off is easy. Reversing a schema migration, a changed cache format, or anything that has already written to a downstream system is not, and no amount of agent sophistication changes that. Before this model helps you, your pipeline has to be genuinely reversible, and that is a platform-engineering programme, not a purchase.
Who answers for it? Nothing in the four safeguards touches this. I have made the fuller argument in When the Agent Acts, Who Owns the Decision? and the model in the Signal Drop episode on the same theme, and I have not found a reason to soften it. Accountability does not relocate. The gate can move into tooling. The person who has to explain the outage stays exactly where they were.
The honest ledger
The cost of this is not the licence, and it does not land where you would expect. It lands on the platform team that has to build the reversal machinery, and it lands well before anyone sees the benefit.
What is genuinely good here, and I want to be fair about it: putting deterministic analysis underneath the probabilistic layer is the right instinct, and it is not what most of this market is doing. Reducing petabytes to a small, accurate context window is a real engineering answer to hallucination rather than a slogan about it. Leaning on feature flags and reversibility as the primary safety mechanism is more honest than leaning on model confidence, and Alois Reitbauer, Dynatrace's chief technology strategist, has pushed the open OpenFeature standard rather than a proprietary one, which is worth something.
What it costs: money, and the shape of the bill is unfamiliar. Reitbauer was unusually blunt about the economics of running agents at this depth. "If I can spend $25 on every single interaction, that is fine, but nobody has the luxury." Note what that admits. The quality is achievable; the unit cost is the constraint. On the public review aggregators, the pattern in user reporting about Dynatrace has been consistent for years: strong marks for depth of capability, recurring complaints about cost, and the cost complaints get louder the smaller the estate. That is aggregated user sentiment rather than a measurement, but it is the same sentiment every year, and agent-led delivery makes the meter run faster, not slower.
And here is the part nobody is selling: the prerequisite. Bluebox assumes staging parity, feature-flag discipline, and a deployment pipeline that can reverse itself under load. If you have those, this is a serious proposition. If you do not, buying it does not give them to you, it just moves your existing gaps somewhere harder to see.
The tell
Two things sit either side of that interview, and they are worth putting next to each other.
The first is in the interview itself. Reitbauer, describing how Dynatrace judges whether its SRE agent is any good, said the team tracks "what the actual output quality is, like how often the output is actually approved and used in the further process."
Approved.
|
Their own measure of a good agent is how often a human said yes. |
Their own measure of a good agent is how often a human said yes. The human is not outside the loop in mode one. The human is the metric.
The second is what happened two weeks later. On 13 August 2026, Dynatrace signed an agreement to acquire Arize for $915 million, roughly $815 million of it in cash. Arize's own description of the problem it solves, in Dynatrace's announcement: teams need frameworks "to detect hallucinations, measure output quality, and continuously validate AI behavior." Arize's chief executive Jason Lopatecki put the founding motivation more plainly still: "AI teams needed a way to know their agents were actually working correctly, not just running."
Seventeen days after saying the human comes out of the loop, the same company committed $915 million to the machinery that checks whether the agent was right.
That is a gate. It is a very expensive one, and I think buying it was the correct decision.
The vendor agrees with you, for now
The most honest line in the whole interview is Greifeneder's answer to the obvious question, which is how a customer actually gets from mode one to mode two.
"To be honest, no one has done that yet. The core applications are too large and too critical for you to move all of this right away."
His advice is to pick a smaller, newer application, one you want to move fast on, and start there, because the process will take a few more months to mature and you largely begin from scratch.
I could be wrong about how long that holds. The direction of travel is not seriously in doubt, and anyone who tells you agent-led delivery will never reach core systems is guessing as confidently as I am. But right now, the man building the thing is telling you the same thing I would tell you: not on the systems that matter, not yet.
What to do on Monday
The argument was never human against no human. It was always where the gate lives, and what backs it up when it is wrong.
So take your three most automated remediation paths and give each one three columns. First, the specific mechanism that stops a bad change, named precisely, not "the pipeline". Second, the person who owns that mechanism and would be asked to explain it. Third, what it costs you in time and money to reverse the change once it has run.
Fill in all three and your gate has moved and you know exactly where it went. That is a good position, and it is the position Bluebox is designed for.
Leave the middle column blank and you have not relocated a gate. You have removed one. The name that would have gone in that column is still on an org chart somewhere, and it will still be the name in the incident review.
|
Work with me An honest read on what your observability is actually doing. | |
|
If you lead observability in a regulated enterprise, I run a fixed-scope Observability Assessment for senior IT and engineering leaders. It ends in a written roadmap and a readout, not a sales deck.
Not ready to talk? Start with a free chapter of Metrics & Mayhem. |

