This website uses cookies

Read our Privacy policy and Terms of use for more information.

In partnership with

Observability you can act on, whether you build it or lead it, plus the few signals worth your time.

This week: the bill nobody itemises, and what your AI SRE spends while it thinks.

Last Friday I said I would show you how to put a number on the thing you cannot fix. Same discipline, different bill. Here it is.

An agent that fires ten queries where a human fires two is not cheaper. It is billed later.

The lead

What your AI SRE spends while it thinks. Every autonomous triage demo shows you the same thing. The reasoning trace, the tidy summary, the root cause found in ninety seconds. Nobody shows you the query log. That is the part I would want to see first, because it is the part that lands on your bill.

A human engineer investigating an incident narrows. They carry context from the last six months, so they discard four hypotheses without running a single query. An agent does not narrow, it fans out. It checks the things an experienced engineer would not bother checking, because it has no reason not to. An agent that fires ten queries where a human fires two is not cheaper. It is billed later.

Reality check: the cost is not the model licence, and your vendor will happily talk about the licence all day. The cost is what the fan-out does to your data layer. Scanned bytes. Query concurrency. The rate limit you share with your own dashboards. The failure mode worth worrying about is not a large invoice at month end. It is the agent exhausting the query budget at 02:00 and the humans losing their tools at the exact moment they get the call.

None of this is an argument against autonomous triage. It is an argument for pricing it before you buy it, and the sum is smaller than the vendors make it sound. Take your last ten incidents. Count the queries an agent would have fired, multiply by whatever your platform actually bills you for, then divide by the ones it would have closed without a human stepping in. That is a cost per resolved incident, with a window and a scope attached, and it is either defensible or it is not. Better that the number is yours.

The full piece carries the two questions that make a vendor sweat: what a single investigation costs to run at your data volume, and what the agent does when it hits the rate limit mid-incident. Ask to watch that second one live. Most cannot show you, and that is the answer.

Signal check

Two that came in this week.

  • "What has actually changed about OpenTelemetry lock-in since 2024? What has the standard freed us from, and where are we still stuck?" It freed you from the agent, which was always the expensive part. Rip and replace used to mean touching every service. Now it means changing an exporter endpoint. What it has not freed you from is everything downstream of ingestion: your dashboards, your alert definitions, your saved queries, the semantic conventions your team invented before the spec caught up, and the four years of muscle memory in the people who use the tool every day. That is where the switching cost sits now, and it is a people cost more than a plumbing one. Worth knowing, because it changes what you negotiate on.

  • "When we finally produce an actionable insight, who looks after it at the other end? Who do you actually shout at, and what happens if no team owns the response?" Then it is not an insight; it is a fact. That is not a semantic point. I have lost count of the dashboards I have seen that proved something was wrong for months while nothing changed, because the finding had no route to a person with a budget. So before you build the detection, name the receiver. If you cannot name them, build the governance route first and the detection second. It feels like the wrong order. It is not.

Latest on the blog

  • Own the Signal, Rent the Platform. Split the stack in two and changing vendors becomes a reconfiguration rather than a re-migration. Read it.

  • What Are Wide Events? Observability 2.0, Without the Hype. Three pillars or one wide row, and what actually changes when you switch. The most-opened link in last week's issue, so here it is again. Read it.

What you missed

Three from the week, and what was worth the time.

  • The big one. Security Boulevard, 24 August, on the ownership vacuum underneath AI observability. Their figure and not mine: in LangChain's survey of more than 1,300 practitioners, 94 per cent running agents in production have some form of observability, and 32 per cent still name quality as the thing blocking them from shipping. Buying the telemetry turned out to be the easy half. Read it.

  • From the field. A thread on r/Observability this week opens with the line that asking your observability vendor how to cut your bill is asking the fox to guard the henhouse. Harsh on a few decent account teams, but the structural point stands. Nobody whose number goes down is the right person to design your reduction plan. Read it.

  • The counterweight. Also this week, on r/sre: change fail rate is not rising because of AI, it is rising because batch sizes doubled. Useful if your board has already decided the coding assistant is the problem. Read it.

That is the week. Next Friday: what actually slows an observability rollout in a large enterprise, and why it is almost never the tooling.

Allan

PS: If you have not read the book yet, start with Chapter 4 below. It is the CrowdStrike chapter: why one airline recovered in a day and another took five, from the identical fault.

 

Everything in one place

New here? Start with the free chapter.

Read the free chapter →

More from Metrics & Mayhem

The book  ·  Get Metrics & Mayhem →
Podcast  ·  Listen on Spotify & Apple →
Connect  ·  Allan on LinkedIn →

Cut Lead Review From Hours To Minutes

Sign up for a free trial of Attio, the agentic CRM.

Ask Attio to build a daily workflow that surfaces the deals that need your attention today, like anything with a stage change, a recent reply, or a new signal in the last 24 hours.

Review your pipeline in Claude, synced live from Attio via MCP.

That's it.