Observability you can act on, whether you build it or lead it, plus the few signals worth your time.
This week: what your monitoring shows you when the thing finishing your customer journey is not a person.
Last Friday I promised the stack I would stand up before asking anyone for a budget. It went up on Sunday and it is in Latest on the blog. This week is about something further out.
Every synthetic check you own asserts a fixed route, and an agent picks its own. |
The lead
When your users stop being people. Greg Isenberg closed a recent episode of The Startup Ideas Podcast pitching a business he called the Agent Mystery Shopper. Point an AI agent at the journeys that matter, then tell the site owner where it got stuck. It is a good idea. It is also synthetic monitoring, which our industry has been selling since the late 1990s and which most people reading this already have running.
The temptation is to enjoy that. Resist it. When a marketing audience independently reinvents scripted transaction checks, the useful signal is not that they reinvented our wheel, it is that somebody should check whether ours still turns. WebMCP is the reason to ask. It is a proposed browser standard that lets a site hand an agent an explicit list of things it can do, instead of making it read the page and guess. Google's Chrome team opened it to early preview in February 2026, and in the demo store that episode is built on, the agent sees sixteen tools signed in and three signed out, because the browser session is the authorisation layer.
The uncomfortable truth: every synthetic check you own asserts a fixed route, and an agent picks its own. Same intent, a different path on every run. When a run fails you inherit a triage question that does not exist today. Is the site broken, or did the model simply choose badly? Nobody has written that runbook.
Where this breaks is the timing, and I would rather say so than oversell it. WebMCP sits behind a browser flag, adoption outside demos is close to nothing, and anyone telling you to rebuild your monitoring strategy this quarter is selling something. The part that ages first is not your synthetics anyway, it is your real user monitoring. An agent has no viewport, perceives no layout shift and never rage clicks. Your Core Web Vitals will keep looking healthy while the population they describe quietly shrinks.
So the move this week costs an afternoon. Take one journey you already monitor and answer a single question honestly: if an agent completed it instead of a person, what would you currently see? For most teams the answer is a load-time metric and a script that no longer represents anybody. The full piece walks the four things that break, in the order they break, including the one that turns a badly worded sentence into a production defect with no review gate in front of it.
Signal check
Two came in this week, and they turned out to be the same question from opposite ends.
"How can we detect if Claude in Chrome or other browser agents are accessing our systems, when the traffic we did not provision looks human in the logs?" Start by accepting that you probably cannot identify them, then decide whether you need to. A browser agent drives a real session, so it carries a real user agent, a real cookie and a real address. What gives it away is not identity, it is shape. A person reads, hesitates, scrolls back, and generates incidental traffic on the way to the thing they wanted. An agent arrives with the decision already made: straight to the endpoint, no dwell, intervals that are too even. So look at inter-request timing, and at what is missing rather than what is present. Then be honest about why you are asking. If the real answer is that you want to block them, that is a product decision wearing a monitoring costume. The better number is what share of completions on a revenue journey came from a session that behaved nothing like a person.
"When an AI trace tells you something is degrading, who actually decides what happens next?" Whoever you named before it started degrading. If you did not name anybody, the honest answer is nobody, and the finding will sit there being correct. Most AI observability stops at detect, trace, alert, and that is not a tooling gap, it is an ownership one. I have watched teams build genuinely good detection and then find it had nowhere to go: it landed in a channel, somebody reacted, nothing moved. So write the decisions down before you switch the detector on. Who is allowed to roll back. Who is allowed to take the model out of the path. What threshold makes it their call rather than a discussion. If none of those has a name against it, you have not built a control, you have built a notification.
Latest on the blog
What I would spin up first, before asking anyone for a budget. The one I promised you last Friday. Four boxes, all free, all open source, running on hardware you already control. The stack is the easy part. What matters is the short list a proof has to establish before a budget conversation stops being a deferral, and the point at which you should stop. The three things your proof has to prove.
What actually slows down an observability rollout in a large enterprise. Worth resurfacing for the governance argument. Most regulated estates run a control that is quietly destroying the audit trail it was written to protect, and the fix is one word. The one word that fixes it.
What you missed
Three from the week, and what was worth the time.
The big one. GitLab published a security analysis of an internal evaluation in which an AI coding agent escaped its sandbox by exploiting a vulnerable package proxy sitting on the sandbox's own allowlist. Craig Risi wrote it up for InfoQ on 8 September. A network allowlist is a control, not a trust boundary: every registry, API and internal dev service you permit becomes part of the agent's reach. His recommendation is the observability argument stated plainly by a security team, and the behavioural signals he names all show green on the infrastructure dashboards you already own. The signals your sandbox cannot show you.
The one with a number attached. Netflix is moving more than 30,000 Flink streaming jobs off the cluster-level autoscaler it built around 2019 and onto the open source Apache Flink autoscaler, reported by Leela Kumili for InfoQ on 7 September. The old one did not fail. The scaling unit was wrong: it reasoned about the cluster, so every operator in a job shared one decision, which falls apart on stateful pipelines with branches and joins. Their figures, not mine, and scoped to their estate. Your automation can only ever be as fine grained as the signal underneath it. What changed on one team's bill.
The other bill. Brian Martin of IOP Systems on the runtime cost of instrumentation, published by InfoQ on 3 September, though the talk was given at QCon San Francisco in 2025, so it is new to read rather than newly said. Counter design, per-CPU sharding, lock-free histograms, then eBPF for the system telemetry in-process instrumentation cannot reach. The framing is the part to steal. One observability bill arrives from your vendor and everybody argues about it. The other is paid in latency on every request. How to stop paying the second bill.
That is the week. Next Friday I want to take that second bill seriously: what under-instrumenting actually costs you, and how to find out what yours is before you decide to keep paying it.
Before you go
Reply and tell me the observability problem on your desk this week. I read every reply. The best ones become a Signal check, and if it is a bigger piece of work, that is exactly the kind of thing I help teams with.
How useful was today's Signal?
Allan
PS: Strip this issue back and it is one question. When the thing using your service cannot tell you it is stuck, who notices? An agent will not raise a ticket, will not complain on LinkedIn and will not churn loudly. It just stops completing. That is an ownership problem before it is a monitoring one, and it is what Chapter 4 is about. It is the chapter I give away free, just below: three layers of ownership, why most organisations build only the platform layer, and one move for Monday, which is to take your most critical journey and name the person who owns the outcome, not the team.
Everything in one place New here? Start with the free chapter. | ||||
More from Metrics & Mayhem
|
Hire Ava, the AI BDR built for enterprise
Ava is the first AI BDR to run outbound end to end, and you decide whether she runs autonomously or on copilot.
She finds leads or ingests accounts from your CRM, enriches them, sends personalized emails on behalf of your reps, follows up, handles replies and books meetings. Website visitor de-anonymization, intent signals, and a parallel dialer come built in.
A small team can manage her centrally for thousands of reps who never log in. Everything syncs two-way with Salesforce and HubSpot.
Ava runs outbound for companies like DoorDash and Grammarly, and one customer deploys her across 1,000+ reps. Ava is SOC 2 Type II audited, SSO and GDPR ready. Ava is how revenue teams grow pipeline without growing headcount.


