Observability you can act on, whether you build it or lead it, plus the few signals worth your time.
This week: the $915m question about who owns your telemetry, and what to sort out before someone answers it for you.
A quick note first. This is The Signal, and from now on it lands every Friday, whether or not there is a podcast episode that week.
The lock-in was never the dashboard. It was the agent you buried in every service. |
The lead
Own the Signal, Rent the Platform. Most observability consolidation programmes open with the same question. Which vendor do we standardise on? It is the wrong question, and it buys you the next migration as well as this one.
Split the stack in two. The signal layer is your instrumentation, your semantic conventions, your collector configuration and your ownership tags. That layer is yours, and it should be portable. The platform layer is storage, query and dashboards. That layer is rented, and it should be swappable.
Reality check: the platform is the part you replace, and the signal is the part you keep. The lock-in was never the dashboard. It was the agent you buried in every service. Standardise the signal on an open standard, and changing platforms becomes a reconfiguration rather than a re-migration. You also get to open renewal talks from somewhere other than "we cannot afford to leave".
Where this breaks: you now own a component you used to outsource. Someone has to run the collector fleet, patch it and keep up with the spec. That is a real cost. It is a known, upfront engineering cost instead of an open-ended recurring one, and that is a trade most finance directors will take once you put both numbers in front of them.
Read the full piece, including the pilot-slice approach for proving it on one service before you commit the estate.
Signal check
Two that came in this week.
"We rolled out observability, found latency is terrible across a whole region, and we cannot fix the infrastructure. Are we just measuring for the sake of it?" No, but you are one step short. Data nobody can act on is a report. Data with an owner and a decision attached is observability. If the fix sits outside your control, the finding still changes something: it moves the conversation from "the app feels slow" to "this region costs us this much, here is the evidence, here is who signs off the change". Take the number, the window and the scope to whoever owns the constraint. If nobody owns it, you have found a bigger problem than latency.
"I am an observability and SRE engineer. How should I actually upskill in AI?" Learn to observe it, not just use it. The scarce skill right now is not prompting. It is being the person who can say how often the agent was right, how often a human overruled it, and what it cost per resolved incident. Almost nobody is measuring that yet. Start by instrumenting one AI-assisted workflow you already run, and treat the model like any other dependency: latency, error rate, cost, and a human who signs off.
Latest on the blog
What Are Wide Events? Observability 2.0, Without the Hype. Three pillars or one wide row, and what genuinely changes when you switch. Read it.
Lead with the promise, not the plumbing. Why platform pitches that open with architecture lose the room before slide four. Read it.
The Observability Maturity Assessment: Where Does Your Team Actually Stand? Five levels, used as a map of what to fix next rather than a scorecard to feel good about. Read it.
What you missed
A few signals from the week, and what was worth the time.
The big one. On 13 August, Dynatrace agreed to buy Arize in a cash and stock deal valued at $915 million, aimed at AI observability across the full lifecycle from evaluation through to production. The announcement.
Then it happened again. On 19 August, Dash0 announced it is acquiring Polar Signals, the continuous-profiling company, with founder Frederic Branczyk going on Code RED to explain the thinking. Two acquisitions in six days is not a coincidence, and between them they make this week's lead argument for me. Consolidation is accelerating, which is exactly when owning your instrumentation stops being a philosophical preference. Listen.
Worth a listen. Code RED again, this time with Rachael Wonnacott of Fidelity International, on platform engineering inside a regulated enterprise. Her point that observability has to be a first-class citizen of the platform rather than an add-on later is the whole argument for treating the signal layer as a product. Listen.
Worth a read. Kearney on managing the cost curve of AI at enterprise scale. Useful if you are the one being asked to forecast what agents will do to next year's telemetry bill. Read it.
That is the week. Next Friday: what a region you cannot fix actually costs you, and how to put a number on it before someone else puts a story on it.
Before you go
Reply and tell me the observability problem on your desk this week. I read every reply. The best ones become a Signal check, and if it is a bigger piece of work, that is exactly the kind of thing I help teams with.
How useful was today's Signal?
Allan
PS: If you have not read the book yet, start with Chapter 4 below. It is the CrowdStrike chapter: why one airline recovered in a day and another took five, from the identical fault.
Everything in one place New here? Start with the free chapter. | ||||
More from Metrics & Mayhem
|
Winning Startups Aren't Bigger. They're Leaner.
Founders deploying AI agents across sales, marketing, and customer service are closing more with smaller teams. 65% more sales leads. Less headcount.
Download the free Practical Guide to Agentic GTM for Startups and start building the stack your competitors dream of.

