This website uses cookies

Read our Privacy policy and Terms of use for more information.

It is two in the morning and something important is broken. The dashboards are lighting up, the bridge call has a dozen people on it, and half of them are typing in the chat at once. In that moment, the single biggest variable in how the next hour goes is not your tooling. It is whoever is leading the room.

We pour effort into instrumentation and almost none into the thing that actually decides the outcome: how the incident is led. The good news is that incident leadership is not a personality trait or a title. It is a discipline with known roles, a repeatable structure, and a set of habits you build long before the page fires. This guide walks through all three, in plain English, for the IT and engineering leaders who carry the pager and the responsibility that comes with it.

What incident leadership actually is

The best incident leaders I have worked with were rarely the deepest technical expert on the call. Their job is different: hold the room, keep one clear owner on the problem, and keep everyone asking the next useful question instead of talking over each other. When an incident drags, it is almost never because the fix was hard. It is because nobody owned it, and an outcome owned by everyone is owned by no one.

An outcome owned by everyone is owned by no one.

You do not have to invent the structure. The model most modern teams borrow is the Incident Command System, developed after the 1970 Southern California wildfires, when thousands of responders arrived and could not work together for lack of a common framework. It gave emergency response a clear chain of command, common terminology, and a structure that scales from a small event to a huge one. Technology teams later adapted it: Google's SRE practice and vendors such as PagerDuty run their incident response on the same bones. You are not being asked to be a hero. You are being asked to run a known play.

The roles on the bridge

The core move of incident command is to separate the jobs that one stressed person usually tries to do at once. Google's incident management guide names three, and they are worth adopting almost verbatim.

  • Incident Commander. Holds the high-level picture, makes the calls, and assigns the work. The commander does not fix the problem; the commander runs the response, and owns the decision, so there is never a question of who decides.

  • Operations Lead. The person, or small group, actually applying the fixes: querying, rolling back, failing over. They work the problem so the commander does not have to, and keep the wider view intact.

  • Communications Lead. Owns the updates to stakeholders and the status page, and fields incoming questions so they do not land on the people fixing the issue. On anything customer-facing, this is a real job, not an afterthought.

A longer incident adds a planning role for handoffs, for tracking what has changed so it can be reverted, and for the unglamorous logistics. On a small incident one person may wear two hats, and that is fine. What matters is that the roles are named out loud at the start, so nobody assumes someone else has it.

Split the jobs one stressed person tries to do at once. Name the three roles out loud: an outcome owned by everyone is owned by no one.

The anatomy of a well-led incident

Every incident moves through the same rough phases, and the leader's job changes at each one.

  • Detect and declare. Someone notices, and someone says the word: this is an incident. Declaring early and naming a commander is half the battle; most slow incidents were slow because nobody took charge for the first twenty minutes.

  • Mobilise. Get the right, small set of people on, name the roles, and open one clear channel. A crowded bridge is slower, not faster.

  • Coordinate. The commander keeps one owner on the current theory, times the next check, and stops the call splintering into three side-investigations nobody is tracking.

  • Resolve. Mitigate first, understand fully later. Getting the customer working again is not the same as knowing why it broke, and the leader keeps that distinction clear so the team does not over-investigate while users are still down.

  • Learn. A blameless review that produces owned actions. Blameless does not mean ownerless; the point is to fix the system, not to find a person to carry the blame.

The habits you build before the night

Here is the uncomfortable truth: most of incident leadership happens when there is no incident. The team that recovers fast on the bad night is the team that did the quiet work in the daylight. Four habits matter most.

  • Map the ownership. For every critical system, and every seam between two teams, write down the one person who gets the call. The seams, the parts both teams assume the other owns, are where the long incidents live.

  • Rehearse the failure that has not happened. Run pre-mortems and game days on realistic faults. Reps bought cheaply in a drill are the ones you draw on when it is real, and they are exactly the reps automation quietly removes, the slow decline in The Six-Week Decay.

  • Position before the page. Starting your thinking at the alert is playing on hard mode. Knowing your dependencies, your rollback paths and your roles in advance is the habit in Position Before the Page.

  • Keep the runbooks honest. A runbook nobody has opened in a year is a liability, not an asset. The best time to test it is a calm afternoon, not a live incident.

Where this breaks

Two failure modes undo good incident leadership more than any technical gap. The first is a culture that rewards heroics. If the person who pulls the all-nighter and saves the day gets the applause, you are quietly training the team to let things break so they can be seen to fix them. Heroism is the polite word for a system that was never positioned. The second is the belief that the next tool will fix it. A better dashboard shortens detection; it does nothing for coordination, ownership, or the twenty minutes a cross-team incident loses while everyone works out whose problem it is. And ownership that reads as 'the team' is the version that hurts most, because it feels like an answer and behaves like a gap, the argument in Accountability Is the Job.

Heroism is the polite word for a system that was never positioned.

The metrics that actually matter

Most teams measure mean time to resolve and stop there. It is useful and, on its own, misleading, because it averages the routine incidents your automation already handles with the rare, novel ones that actually test the team. Watch three things instead. First, how you handle the incidents you could not automate, because that is where leadership shows. Second, time to declare and time to the right owner, the two intervals where slow incidents are usually lost. Third, the share of review actions that get owned and closed, because an incident you did not learn from will visit you again. Treat the numbers as a prompt for the honest conversation, not a substitute for it.

What to do on Monday

You do not fix this with a project. You fix it with a handful of habits, starting this week.

  • Adopt the roles. On the next incident, name an Incident Commander, an Operations Lead and a Communications Lead out loud, even if one person covers two. Naming them is most of the value.

  • Name the owners before the night. Write down who gets the call for each critical system and each team seam. If the answer is 'the team', it is unowned.

  • Rehearse one bad night. Run a thirty-minute pre-mortem or drill on something that has not gone wrong yet.

  • Watch your own signal. The team mirrors the leader under pressure. Calm creates clarity; chaos creates noise. Your composure is part of the tooling.

Frequently asked questions

What is the difference between an incident commander and an incident manager? The incident commander is a role for the duration of a single incident: they run that response and stand down when it closes. An incident manager is usually a standing job that owns the process, the tooling and the reviews across many incidents. On a given night, you want a commander, not a committee.

Does blameless mean nobody is accountable? No. Blameless means you do not punish people for honest mistakes, so they tell you the truth. It does not mean an action has no owner. Blameless and ownerless are different words, and confusing them is how a review produces a list of actions nobody does.

We are a small team. Do we really need all these roles? The roles are jobs, not headcount. One person can be commander and operations lead on a minor incident. The discipline is to name the jobs so none of them silently goes undone, not to fill a rota.

How do we start if we have none of this today? Pick one habit. Declare, and name a commander, on the next real incident. That single change, taking charge in the first five minutes, will do more than any tool you could buy this quarter.

 

Work with me

An honest read on what your observability is actually doing.

If you lead observability in a regulated enterprise, I run a fixed-scope Observability Assessment for senior IT and engineering leaders. It ends in a written roadmap and a readout, not a sales deck.

See how it works →

Not ready to talk? Start with a free chapter of Metrics & Mayhem.

Reply

Avatar

or to participate

Keep Reading