Observability you can act on, whether you build it or lead it, plus the few signals worth your time.
July kept circling one question: when something acts on your behalf, whether that is an AI agent or a person running an incident, who actually owns what happens next? Five full posts and this month's Signal Drop all arrive at some version of the same answer.
You can automate the action. You cannot automate the accountability. |
Set your stream: reply TECH for the hands-on, practitioner cut, reply LEADERSHIP for the strategy cut, or do nothing to keep the full issue. One word, change it whenever.
This month's deep dive
The Gate Stays Human: Who's Accountable When AI-Ops Acts. A governance model for agent actions that fits in your head: decide by blast radius, not confidence. Low risk, open the gate and log it. High risk, a named human approves before it acts. You can automate the action. You cannot automate the accountability.

METRICS & MAYHEM
The Gate Stays Human: Who's Accountable When AI-Ops Acts
Read the companion piece: The Gate Stays Human.
On the blog this month
Blogs publish web-only and are not emailed individually, so here are July's full posts:
When The Agent Acts, Who Owns The Decision? The accountability gap agents open the moment they move from advising to acting, and how to close it with a blast-radius test.
You Are Doing Financial Engineering When You Should Be Doing Real Engineering. An industry now exists to shave the observability invoice. That is financial engineering. The real fix is deciding what each signal is for.
There Is No Agreed Pattern For Running AI Agents Yet. A vendor best practice, before a real pattern exists, is just their product with a halo on it. Keep a human on the decision until the pattern earns its place.
What's trending this month
Two arguments kept coming back, all month, everywhere we look: Reddit, LinkedIn, our own inbox.
Who actually owns the agent when it acts? Not "can it act", everyone has moved past that. The industry still has no good answer for who's accountable when it does, and that gap is where I keep landing too.
Everyone wants to be the cheaper alternative, or wants to stop competing altogether. The cost pressure on observability spend is producing two completely different responses this month: undercut on price, or get bought before you have to.
Neither is settled. I'll pick this back up next month and tell you where each one actually moved, not just where it was loudest.
Byte-size drops this month
Three quick explainers worth a mention, each paired with its Tech Tuesday video:
What Is MCP? The open standard letting AI agents use your tools and data, and its honest limits.
What Is Agent Observability? Not "is the model up", but "why did the agent do that": model calls, tool calls and reasoning as one trace.
What Is OpenTelemetry Profiling? Continuous profiling is becoming observability's fourth signal, and the honest cost of always-on.
Full Tech Tuesday videos: on the podcast page.
Good posts I saw this month
A few things from other voices in the space, worth your time and credited to them, not me:
Christian Posta (Field CTO, Solo.io): not everything has to be an AI agent. Counter-hype from a respected architecture voice: CRUD is still CRUD, and known workflows don't need a non-deterministic wrapper.
Dan Gomez-Blanco (OpenTelemetry maintainer): "OpenTelemetry is too complex" gets the same reality check "Kubernetes is too complex" got in 2016. Owned complexity you can see beats a simple box that quietly lies to you.
Charity Majors (CTO, Honeycomb): the problem was never "AI SRE", it's that on-call sucks. A model laid on top of a broken rotation just automates the mess faster.
Worth your attention
Four listens and two reads that earned it this month:
[Podcast] Christine Yen (CTO, Honeycomb) on the Dev Interrupted podcast. Observability in the AI-agent era needs richer telemetry than the old three pillars, and she reframes it as a revenue driver rather than a cost centre.
[Podcast] Krishna Nakoda (senior SRE, Microsoft) on SRE Pod. Treat alert logic like production code and use chaos engineering to surface the failure modes nobody has named yet. A sharp companion to the alert-ownership problem I am recording next.
[Podcast] Judd Naus (founder, Prepper) on an Observability Podcast interview. Real-time pattern detection and compression cutting observability data volume 90 to 95 per cent, and what that does to your high-cardinality problem.
[Podcast] Alex Zendla (CTO, Adira) on the Cloud Native Security Podcast. A node-isolation approach to Kubernetes security, from someone who has had to defend the design in front of a board.
[Article] Apple quietly acquired SigScalr, the team behind the open-source observability tool SigLens, disclosed via an EU Commission filing rather than an announcement. Site offline, GitHub archived read-only.
[Article] Most enterprises will hand root cause analysis to AI agents within two years. Elastic's Thaddeus Walsh on agent-led investigation and where the adoption numbers already sit.
Next month I want to go one layer down from accountability: what it actually costs a team, in hours and in trust, when nobody has answered the ownership question before the agent or the incident forces the answer.
Allan
Everything in one place New here? Start with the free chapter. | ||||
More from Metrics & Mayhem
|
What's changing in global hiring?
Global hiring is changing fast. From AI's impact on HR to evolving compliance requirements and international expansion strategies, the rules are constantly shifting.
Oyster's events bring together HR leaders, founders, operators, and global employment experts to discuss what's working now—and what's coming next.
Whether you're hiring internationally today or planning for tomorrow, you'll walk away with practical insights you can actually use.

