Observability you can act on, whether you build it or lead it, plus the few signals worth your time.
Who signed for it? That question sat under every piece in September. An agent's run mode, the CPU your dependencies burn, the redeploys you do just to answer one question: all of it already costs you something, and almost none of it carries an owner's name.
Most of what you pay for in observability was never signed for by anyone. |
This month's deep dive
The Fourth Signal Was Hiding In Your CPU Bill. Metrics tell you the host is busy. Continuous profiling tells you which code is spending the cycles, including code you did not write. OpenTelemetry has taken profiling on as its fourth signal, and it reads a bill you already pay and have never seen itemised.
No new Signal Drop this month. The podcast feed has been quiet since 11 August while it is rebuilt around Tech Tuesday, and it comes back in October. The episode that pairs with this deep dive is from July, and it is the explainer the new piece builds on.

METRICS & MAYHEM
OpenTelemetry Profiles: The Fourth Signal, Explained
Read the full piece: The Fourth Signal Was Hiding In Your CPU Bill.
On the blog this month
Five full pieces went up in September:
WebMCP Breaks Synthetic Monitoring. Nobody Has Noticed Yet. Websites can now hand an AI agent a list of tools instead of a page to read. Your scripted checks still assume a user with eyes.
What I'd Spin Up First (Before Asking Anyone for a Budget). Most observability initiatives die in a procurement queue, not on a technical problem. Build the proof yourself on a free open source stack and let it do the arguing.
Somebody Turned It On. Nobody is selling you an unsupervised agent. They are selling you a dropdown. Autonomy shipped as a setting, and settings have owners.
The Fourth Signal Was Hiding In Your CPU Bill. This month's deep dive, above.
Conway's Law Came for Your Agents. Every team wants its own agent, because owning one is the price of being a team. Nobody owns the total. Ask for the register before you approve another.
Where each of those went, and why you may not have seen them
This month the blog stopped going to everybody. A technical piece goes to the technical stream, a leadership piece goes to the leadership stream, and The Signal still reaches every one of you each Friday. Sending everything to everyone fills inboxes with things people never asked for, so we stopped doing it.
Here is where September's five landed.
Technical stream, 2 pieces:
Leadership stream, 2 pieces:
No inbox at all, 1 piece:
Somebody Turned It On. This one went out on the web only, so it reached nobody's inbox, the leadership stream's included. That was my miss, not yours.
So find yourself in that list. On the technical stream, two of these reached you. On the leadership stream, two did. If you never picked a stream, none of these reached your inbox, and that is the only reason you missed them. Every one of them is on the web right now, at the links above.
Which do you want more of?
What's trending this month
Last month I said I would come back to two arguments and tell you where they moved. The first was that the constraint had shifted from the model to your data. In September the vendors started saying it for me. At Splunk's conference, Karun Subramanian, who leads AI for infrastructure and operations at UnitedHealth Group, described a first AI SRE agent built on fragmented data sources that started at 10 to 15 per cent accuracy on root cause. His answer was not a better model. It was a context layer: telemetry joined to the CMDB, the service desk, the source code and the runbooks.
The second was ownership, of your tools and of your bill. That one moved sideways. Pricing is being rebuilt around what you search rather than what you ingest, and agent token spend now gets its own dashboards. Useful. But a survey of 356 enterprise IT professionals published this month found three in four organisations running up to 12 observability tools, and not one that had consolidated down to a single pane. Consolidation is being bought. Ownership is not.
The new argument sat underneath both: who owns the thing once the person who knew why has gone. It showed up as knowledge that lived in one engineer's head, as vendor renewals kept in somebody senior's memory, and as the question of who decides during an incident when an agent is in the loop. The uncomfortable truth: most of what you pay for in observability was never signed for by anyone. I will pick this one up again in October and tell you what actually changed, not what got loudest.
Byte-size drops this month
None. No byte-size explainer went up in September; the month's explainer work went into the two Tech Tuesday pieces above. The earlier explainers and every Tech Tuesday episode are on the podcast page.
Good posts I saw this month
A few things from other voices in the space, worth your time and credited to them, not me:
Prabhakar D (VP, Site Reliability Engineering, JPMorgan): the problem is rarely observability, it is the gap between signals and decisions. He sets out five questions a mature SRE organisation should answer fast. The fifth, who has the authority to make the call, is the one no platform can answer for you.
Dan Gomez Blanco (Principal Observability Architect, New Relic; OpenTelemetry maintainer): "vendor neutral" undersells why OpenTelemetry won. His case is that the standard and platform engineering grew up feeding each other. If you are the team choosing what to build on, that is the more useful explanation.
Christian Posta (Field CTO, Solo.io): "AI agents behave much more like function executions in a FaaS than a server workload." One line. If he is right, the unit you trace is the invocation, not the process, and the replies under it go straight to identity and authority.
Worth your attention
Three listens and three reads that earned it this month, no two from the same source:
[Podcast] Sir Steve Hansen on the High Performance podcast. The former All Blacks head coach on keeping a list of the inconvenient facts his team was prepared to admit out loud rather than paper over. Swap the team for your estate and you have the post-incident review most organisations never hold.
[Podcast] Andrew Ziegler and Adam Noble on Dev Interrupted's Friday Deploy news segment. The twilight factory: agents that do the work but are built to know when to bring a human back in. The design question is not whether a human is in the loop. It is who decides when.
[Podcast] Nathaniel Whittemore on The AI Daily Brief, on OpenClaw 2.0. He reads out a maintainer's account of moving a team onto one shared agent workspace, where the session itself became the handoff document. If you have ever lost a week to a handover, listen for that part.
[Article] IT orgs grapple with observability tool sprawl, CIO Dive. Makenzie Holland on the EMA research behind the survey figures above, and the researcher's advice to get the data straight before adding AI, because AI is one more tool you then have to observe.
[Article] Seismic shift in Splunk pricing spotlights AI data management, TechTarget. Beth Pariseau on activity-based pricing, which holds off indexing charges until data is searched. Read it for the bill and stay for the UnitedHealth panel.
[Article] OpenTelemetry has graduated… now what?, CNCF blog. Adriana Villela and Reese Lee, the project's community managers, on what graduation means for you and where OpenTelemetry goes next: agentic workloads, browser and mobile, and help for teams running it at scale.
Next month: what your telemetry knows once the person who knew why has gone. Before then, pick one system you run and write down who signs for its bill. If nobody can, you have found October's work.
No Signal this Friday. This Digest is this week's email from me, and I would rather send you one than jam your inbox with two. The Signal is back on Friday 9 October.
Allan
📕 The book is out. Metrics & Mayhem: A CTO's Guide to Observability That Actually Works. Kindle is live now; paperback and hardback launched on 1 June.
Everything in one place New here? Start with the free chapter. | ||||
More from Metrics & Mayhem
|
