This website uses cookies

Read our Privacy policy and Terms of use for more information.

I heard a story on a podcast recently, and I have kept it anonymous because the shape is what matters, not the company. A data centre had its cooling units on manual. Someone on the other side of the world woke up, saw spare capacity, and turned servers on. Nobody had told them maintenance was running. The temperature did what temperature does.

Now change one thing. Make that person an agent that never stops asking whether it was supposed to act. That question sits under every 'agentic operations' pitch this year, and the honest answer is uncomfortable: we do not yet have an agreed way to run these things safely.

We do not yet have an agreed way to run these things safely.

We write the protocol fast, and the practice slowly

Our industry is good at one thing and bad at another. When something new lands, we write a protocol for it quickly. We are much slower to write the practices because practices come from watching the thing behave in production over time, and agents have not been in production long enough for anyone to have earned that knowledge. So there is a gap: the capability is here, the operating pattern is not. That gap is normal for new technology. It is also exactly where the trouble starts.

The capability is here; the agreed way to run it safely is not. Vendors fill that gap with a best practice that quietly describes their product. Until the practice earns its place, keep a human on the decision.

A vendor best practice is their product with a halo on it

Into that gap step the vendors, with 'best practices' for running agents. Read them closely, and they tend to describe one thing very well: the vendor's own platform. That is not dishonesty, it is gravity. When no shared pattern exists, the only pattern anyone can point to is the one their product already implements. Adopting it feels like adopting a standard. It is adopting a roadmap, and the roadmap is theirs, not yours. Before a real pattern has emerged from the field, a vendor's best practice is just their product with a halo on it.

Where this breaks

This is not an argument to freeze and adopt nothing. Waiting forever is its own failure, and some of these tools are genuinely worth running now. The point is narrower: do not mistake a vendor's framing for an agreed pattern, and do not hand an autonomous process the authority to change production state at two in the morning because a slide said it was best practice. The same outage shape keeps repeating across my career, a system or a person acting without checking whether they were meant to. An agent is that failure with the brakes off and nobody obviously holding the responsibility.

An agent is that failure with the brakes off and nobody obviously holding the responsibility.

What to do on Monday

Keep a human on the decision until the pattern earns its place. Concretely, for one agent or automation you already run, write down three things: the action it is allowed to take unsupervised, the blast radius at which a human must approve first, and the name of the person who gets the call when it acts wrongly. If the honest answer to that last one is 'the team', your gate is not human yet, it is just unowned. Fix that before you add the next agent, not after. The fuller argument is in When The Agent Acts, Who Owns The Decision?, and if the word itself is still doing too much work, start with What Is an AI Agent? When there is no agreed pattern yet, keep the human on the decision until the pattern earns the seat.

 

Work with me

An honest read on what your observability is actually doing.

If you lead observability in a regulated enterprise, I run a fixed-scope Observability Assessment for senior IT and engineering leaders. It ends in a written roadmap and a readout, not a sales deck.

See how it works →

Not ready to talk? Start with a free chapter of Metrics & Mayhem.

Reply

Avatar

or to participate

Keep Reading