This website uses cookies

Read our Privacy policy and Terms of use for more information.

In partnership with

Observability you can act on, whether you build it or lead it, plus the few signals worth your time.

This week: the four arguments that decide an observability rollout, and why not one of them is technical.

Last Friday I said this issue would take on what actually slows a rollout in a large enterprise. Here it is.

Nobody is obstructing. Everybody is being reasonable. And the programme still cannot say what it is watching.

The lead

The blocker is never the tooling. Somebody in the programme meeting eventually asks what looks like the easy question. What is a business journey? Watch the room when that lands. Everybody knows the answer and no two people give the same one. The application team describes a sequence of technical calls. The service owner describes a customer outcome. Risk describes a control boundary. Somebody from the platform side describes a set of components, which is not a journey at all, it is an inventory with ambition.

Here is what makes that expensive. Nobody is obstructing. Everybody is being reasonable. And the programme still cannot say what it is watching, so it instruments components instead, reports on components, and then wonders why the executive dashboard is green while a customer cannot complete a payment.

The part we miss: this is not a technology decision, so nobody ever puts it in the plan.

Say you get the definition agreed. Somebody still has to go first, and in a large estate a single journey is owned by nine teams, each watching the other eight. Going first means absorbing the effort, the schema arguments, and the awkward discovery that your own service is noisier than you claimed, while the other eight watch. Nobody moves. In a steering meeting that reads as resistance. It is not resistance; it is rational behaviour under uncertainty, and calling it a culture problem makes it worse. What unsticks it is removing the risk of going first rather than selling the value of it: somebody senior enough to reallocate a sprint names the team, names what they are allowed to drop in exchange, and names what counts as done.

Two more arguments sit in the full piece, and the one I did not expect to be writing about is your own governance. Most regulated estates run a control that is quietly destroying the audit trail it was written to protect, and the fix is a single word. Worth ten minutes if anybody has ever asked you to prove what your team believed at 03:20.

Signal check

Two came in this week.

  • "A runtime trace shows what happened. What would make it strong enough to serve as assurance, not just after-the-fact debugging?" Most traces are nowhere near, and the reason is who the trace was written for. A debugging trace only has to convince you, this afternoon, with the context already in your head. An assurance trace has to convince somebody who was not there, months later, with none of it. Three things move it across: identity on every span, so you can say who or what acted rather than only that something did; retention that outlives the question, because your audit window is almost always longer than your trace window and nobody checks that until they need to; and the reasoning, not just the outcome. That last one is the real gap. Your trace records what the system did. It rarely records what anybody believed at the time. Evidence is a trace somebody else can read.

  • "Who actually owns observability tooling, the BAU team keeping the platform healthy or the dev team writing the tests and parsing the results? And who funds the gap?" That is not a daft way to frame it, but the split you are describing is usually the symptom rather than the boundary. Both teams are funded to do their own job well. Neither is funded to own the space between them, which is where observability actually lives. So stop drawing the line by team and draw it by artefact. Somebody owns the instrumentation standard. Somebody owns the platform's health. Somebody owns whether the output ever reaches a person who can act on it. Name those three and the funding conversation gets easier, because you are no longer arguing about who is more responsible, you are pointing at which of three named things has no owner. It is nearly always the third.

Latest on the blog

  • Autonomous SRE Agents: Where Does the Gate Actually Live? Dynatrace shipped a mode with no human in the loop. Read its own safeguards and the gate has not gone anywhere. It moved into software, and somebody still owns it. The four safeguards, and who holds the gate now.

  • byte-size: Observability vs Monitoring. Three minutes on a distinction most vendors are happy to leave blurred. Built to forward to whoever keeps using the two words interchangeably. Why three dashboards are not observability.

What you missed

Three published this week, and what was worth the time.

  • The big one. Sachin Bansal of Schellman, published by the Cloud Security Alliance on 31 August, on the gap between what CISOs say about AI governance and what their organisations do. His numbers come from a private CISO summit rather than a random sample, so treat them as directional: 54 per cent cannot say who owns AI governance. His line for it is worth stealing. It is hard to pass an audit nobody owns. He also gives a four-question test for telling a real governance programme from paperwork. The four questions.

  • The one that will argue with you. InfoQ published Yao Yue's talk this week. Her case, after years on cache and performance engineering at Twitter, is that the line chart became observability's default by accident, not by fit. Joining two readings with a line invents everything between them, and it misleads hardest when the data is noisy, which is exactly when something is wrong. The talk was given at QCon last year and published only now, so it is new to read rather than newly said. Why a latency chart with one line is not a latency chart.

  • The practical one. Honeycomb is donating its adaptive tail sampling processor to the OpenTelemetry Collector. It is a vendor writing about its own work, so read it accordingly, but the mechanism stands up. Instead of one fixed rule, it groups traces by shape and spreads a single sampling budget across the groups, so a quiet tenant still shows up and a spike on one endpoint stops swallowing everything else. How one endpoint stops eating everyone else's traces.

That is the week. Coming up: the stack I would stand up first, before asking anybody for a budget.

Before you go

Reply and tell me the observability problem on your desk this week. I read every reply. The best ones become a Signal check, and if it is a bigger piece of work, that is exactly the kind of thing I help teams with.

How useful was today's Signal?

One click helps shape next week. Add a line if you want to tell me why.

Login or Subscribe to participate

Allan

PS: Strip this issue back and it is one question. Who owns the outcome when the outcome has no owner? That is Chapter 4, and it is the one I give away free, just below. Three layers of ownership, why most organisations build only the platform layer, and one move for Monday: name the person who owns your most critical outcome, not the team.

 

Everything in one place

New here? Start with the free chapter.

Read the free chapter →

More from Metrics & Mayhem

The book  ·  Get Metrics & Mayhem →
Podcast  ·  Listen on Spotify & Apple →
Connect  ·  Allan on LinkedIn →

The Future of AI in Marketing. Your Shortcut to Smarter, Faster Marketing.

Unlock a focused set of AI strategies built to streamline your work and maximize impact. This guide delivers the practical tactics and tools marketers need to start seeing results right away:

  • 7 high-impact AI strategies to accelerate your marketing performance

  • Practical use cases for content creation, lead gen, and personalization

  • Expert insights into how top marketers are using AI today

  • A framework to evaluate and implement AI tools efficiently

Stay ahead of the curve with these top strategies AI helped develop for marketers, built for real-world results.