This website uses cookies

Read our Privacy policy and Terms of use for more information.

Sponsored by

Most observability initiatives do not die on a technical problem. They die in a queue.

I have watched a genuinely good case for observability lose its momentum somewhere between the third architecture review and the second procurement gate, while everyone involved nodded along and agreed it was important. Nobody argued against it. It just never became anyone's this-quarter problem, because the first ask was large, the first ask was abstract, and the first ask arrived as a slide deck rather than as something you could click on.

So here is the move I would make instead, and the one I would tell any engineer sitting on a stalled case to make: do not open with the budget request. Open with a working proof you built yourself, for nothing, on hardware you already have. Let the proof do the arguing.

The bottleneck is permission, not capability

The uncomfortable thing about a modern observability stack is that the hard parts are mostly solved. The collectors are mature. The stores are mature. The open standards are real and widely adopted. If you gave a competent engineer a laptop and an afternoon, they could get real signal flowing out of a real service and onto a dashboard. That is not the constraint.

The constraint is everything wrapped around that afternoon. A licence to sign. A vendor to onboard. A security review of a SaaS agent that wants to ship your telemetry off-premises. A data-residency question nobody wants to own. A cost model priced by ingest volume that finance cannot forecast, so finance says not yet. Each of these is reasonable on its own. Stacked together, they turn a one-afternoon technical task into a two-quarter approval marathon, and approval marathons are where good ideas go quiet.

You cannot procure your way past that quickly. But you can sidestep it, because the tools that let you prove the idea do not need any of those approvals. They are free, they are open source, and they run on a box you already control.

The open source observability stack is four boxes

Strip an observability platform down to its job and it is four things in a line.

Something in your services to emit the signal. A collector to receive it and shape it into a consistent form. A store to keep it. A visualisation layer to read it. Most commercial platforms on the market are a polished, integrated, supported version of those same four boxes. You are not going to reinvent them. You are going to assemble the free versions for the length of the proof.

For the first two boxes, OpenTelemetry is the honest default now. It is an open standard rather than a product, which is exactly what you want here, because the instrumentation you write against it is the instrumentation you keep when you eventually buy the grown-up platform. You instrument your services with the OpenTelemetry SDKs, and the OpenTelemetry Collector receives what they emit, then processes and reshapes it before it forwards. If you want the detail on that middle box, I have written a full guide to configuring a collector properly.

For the store, split by signal and keep it boring. Prometheus for metrics. Loki for logs. Tempo for traces, if you are proving traces at all. None of these will hold your entire production estate, and that is fine, because you are not asking them to.

For the visualisation layer, Grafana reads all three and is what most of the industry already looks at anyway.

Four boxes, and you already know all four. Swap any box for the one your team knows.

The point of naming them is not that this is the one true stack. Swap any box for the one your team knows. The point is that each box has a swappable enterprise equivalent, so nothing you build in the proof is wasted when you scale up. You are building the shape, not the destination.

What you are actually proving

Here is where most proof-of-concept builds quietly overreach and hurt their own case.

You are not proving that this scrappy stack can run production. It cannot, and claiming it can is how you lose the room the moment someone technical pokes it. You are proving three narrower things, and they are enough.

First, that the signal exists: that the thing leadership wants visibility into does actually emit something you can catch. Second, that the signal is useful: that when you put it on a dashboard, it answers a question somebody was previously guessing at. Third, that the shape of the pipeline works end to end, from a real service to a readable view, without a vendor in the middle.

Three narrow proofs are enough. Everything outside them is a sizing conversation.

Prove those three and you have retired the only questions that actually block funding. Everything left, scale, retention, high availability, support, the whole enterprise wishlist, becomes a sizing conversation. And sizing conversations get budgets. It is the "will this even work here" conversations that do not.

Keep the scope that tight and you can stand the whole thing up quickly, which matters, because a proof that takes a quarter to build is just the original programme wearing a smaller badge.

A running demo beats a slide, every time

Now the part that changes the budget conversation.

When you walk into the room with slides, you are asking people to fund a promise. Every question is a reason to wait: what if the data is not there, what if it is too noisy, what if our services are too old. You cannot answer any of them with certainty, because you have not tried, so the safe institutional answer is next quarter.

When you walk in with a laptop and a live dashboard fed by one of your own real services, the questions invert. Now it works, visibly, and the questions become how do we make this real, who runs it, how much to do it properly. You have moved the discussion from whether to how much, and how much is a discussion that ends in a number rather than a deferral.

I would rather demo one real journey badly than present ten hypothetical ones beautifully. A dashboard that is ugly but live outargues a diagram that is elegant but theoretical, because the live one has already answered the question the diagram is only asking.

There is a quieter benefit too. A working proof gives the people who have been hedging something to rally around without personal risk. It is a great deal easier to back a thing that already exists than to sponsor a thing that might.

When to stop doing it this way

This approach has a real cost, and I would rather name it than pretend it does not.

The cost is that everything in the free stack is yours to run. Nobody is on call for it but you. There is no support contract, no hardened multi-tenancy, no compliance attestation, and no one to blame at three in the afternoon when the store falls over because it was never sized for the load you quietly kept adding to it. That is a fair trade for a proof. It is a terrible trade for production.

So the discipline is knowing when the proof has done its job and stopping. Stop when you are tempted to point real production traffic at it. Stop when the answer to a security or data-residency question becomes we will sort that later rather than that does not apply to a throwaway demo. Stop when a second team wants to depend on it, because the moment something is depended on, it is production whether you called it that or not.

At that line, the self-hosted proof has already won the argument it was built to win. What comes next is a funded platform with an owner, a support model and a budget, and the instrumentation you wrote against open standards carries straight over. The scrappy stack was never meant to be the destination. It was meant to get you permission to build the real one.

That is the whole trick. Stop asking for the money to prove it will work. Go and prove it will work, for free, and let finance discover that the only remaining question is how big.

GET THE NEXT ONE

One signal a week. No noise.

If this was useful, Metrics & Mayhem sends one short, practical piece like it to IT operations leaders most weeks. No fluff, no vendor noise.

Prefer to start with the book? Read a free chapter.

Smarter CRM. Less Busywork.

Disconnected data and tools make it harder to understand your customers. HubSpot's Agentic Customer Platform brings your data, teams, and tech stack together with AI built in to help your business work faster and create more personalized customer experiences.

Why HubSpot and what's new

  • Use AI powered tools to take action faster

  • Unify your data, teams, and tech stack in one place

  • Create one shared view of customer data

  • Connect teams around the same customer context

  • Bring your business tools into one place

Connect more of your business in one place and give every team a smarter way to work. Get set up quickly and start checking off your hardest tasks.