This website uses cookies

Read our Privacy policy and Terms of use for more information.

Free chapter · Metrics & Mayhem

What Separates the Teams That Recover Fast from Those That Don't

Chapter 4 of Metrics & Mayhem by Allan Mann. The chapter that answers the question every engineering leader is afraid to ask out loud: when it breaks, who actually owns getting it back?

  • Why one airline recovered from the CrowdStrike outage in a day and another took five, from the identical software fault.
  • The three-layer ownership model: outcome owner, indicator owner, platform owner.
  • Why most organisations build only the platform layer, then wonder why their P1s take an hour.
  • One concrete Monday action you can take before you close the page.
Get the free chapter →
Metrics and Mayhem book cover
The story behind the chapter

One fault. Two recovery times: a day, and five.

In July 2024 a single faulty software update took out machines across the world in hours. Every organisation hit ran the same systems and faced the same broken fault. Yet the recovery times were not the same. Some were trading again inside a day. Others were still cancelling operations most of a week later.

The gap was not the technology. Everyone had the same tooling, the same vendors, the same dashboards. What separated the fast from the slow was something quieter and much harder to buy: a clear answer to the question of who owned getting each critical outcome back.

You can buy every tool on the market and still not know who owns the recovery.

Chapter 4 lays out the model underneath that difference. Recovery speed tracks ownership clarity, and ownership sits in three layers. The outcome owner is the one person accountable for a business result, like "customers can check in". The indicator owner owns the signals that tell you whether that outcome is healthy. The platform owner runs the tools underneath. Most organisations staff only the third layer, then act surprised when a P1 drifts for an hour while everyone waits for someone to own the call.

The chapter walks through how the layers fit together, why the outcome layer is the one almost nobody names, and how the teams that recovered in a day had already answered the ownership question long before the incident forced it. It closes with the one action worth taking this week: pick your single most critical business outcome, and name the one person who owns it end to end. Not the team. The person.

It reads on its own, and it is free.

Read Chapter 4 free →

What you'll take away

Four ideas from Chapter 4

The CrowdStrike recovery gap

Why one airline recovered in a day and another took five, from the identical fault. The answer is not the technology.

The three-layer ownership model

Outcome owner, indicator owner, platform owner: a clear way to assign accountability in complex systems.

Why P1s take an hour

Most organisations build only the platform layer, then wonder why incident response is so slow.

A concrete Monday action

Pick your most critical business outcome and name the one person who owns it end to end. Not the team. The person.

The full book

Chapter 4 is one of eleven

Chapter 1
Why This Is a Board-Level Problem
More tools, more data, more dashboards, and recovery times keep getting worse. Why this stopped being a technology problem.
Chapter 2
What Observability Actually Means for a CTO
Monitoring tells you the system is healthy. Observability tells you whether the customer can pay you.
Chapter 3
What It's Really Costing You
The licence fee is the part you can see. The engineer time spent firefighting is the part you can't.
Chapter 5
The People Problem Nobody's Solving
Lack of team knowledge, not tooling, is the top barrier most organisations name. Training is not capability.
Chapter 6
Why Your Best Engineers Aren't Telling You the Truth
The first five minutes of your incident bridge set the culture for the next five years.
Chapter 7
The Machines Are Getting Better. Are Your Teams?
Human-on-the-loop is the new model. Skill atrophy is the hidden cost.
Chapter 8
The CTO Nobody Trained You to Be
The board does not want forty-seven panels. It wants three numbers.
Chapter 9
What the Organisations That Get This Right Actually Did
Three real responses, three different choices, three lessons in what recovery actually costs.
Chapter 10
How to Build This Without Burning It Down
Outcomes first, tools second. Nine months from announcement to operating model, with evidence at every stage.
Chapter 11
What to Do Monday Morning
Three conversations this week. A 90-day scorecard. The model finally in your hands.
Allan Mann headshot
About the author

Allan Mann

Allan Mann is an IT operations leader and consultant with two decades of experience running large-scale technology environments. He publishes the Mastering Observability newsletter and hosts the Metrics & Mayhem podcast.

Spending on observability has never been higher. Recovery times have never been worse. Metrics & Mayhem is his answer to why, and what to do about it.

Get Chapter 4 free

Read it on Monday. Name your outcome owner by Wednesday. Metrics & Mayhem by Allan Mann, chapter 4, free download.

Get the free chapter →

No spam. One short, practical newsletter most weeks, and you can leave whenever you like.