You are three updates into a bad incident. You told the team the plan on the bridge, you posted it in the channel, and you moved on. Twenty minutes later two people are doing the same thing, one is doing the opposite, and someone senior is asking a question you already answered. You communicated. They did not receive it. Under pressure, those are two different events, and the gap between them is where incidents quietly go wrong.
Most advice on operational communication is about writing tidier status updates. The real problem is harder and more human: under pressure the message you send is not the message people receive, and the difference grows the more stress everyone is under. This guide covers what actually closes that gap, the structure, the audiences, the cadence, and the failure modes, for the IT and engineering leaders who have to keep a room and a business informed at the same time.
The message you send is not the message people receive. |
Why operational communication is its own discipline
Communicating in normal times is a skill. Communicating in an incident is a different one, because three things work against you at once. People are stressed, so they hear through their own assumptions and whatever they were doing thirty seconds ago. Time is short, so nobody reads to the end. And you are speaking to several audiences at once, each of whom needs something different from the same event. Treating all of that as 'send a clear update' is why clear updates still fail.
The three audiences you are speaking to
Almost every operational message is really three messages, and the mistake is writing one for all of them.
The responders. They need precision: what is confirmed, what is being tried, what you need from them next. Detail is welcome here; ambiguity costs minutes.
The leadership. They need the headline and the shape of the risk: how bad, who is affected, what it means for the business, and when the next update lands. Not the stack trace.
The customers. They need honesty about impact in plain language, and a promise of when they will hear more. The status page is their source of truth, and it should describe what they cannot do, not your internal theory of why.
The same incident, three registers. A leader who blurs them either drowns executives in detail or leaves responders guessing, and both slow the response.
The structure that makes a message land
Under pressure, structure beats volume. A simple, repeatable shape stops the drift: context (what is happening), intent (what we are doing about it), headline (the one thing you need from the reader). Lead with the headline for the people who will only read the first line, and give the context underneath for the ones who need the why. Said the same way every time, it turns a broadcast into something people can act on, the discipline in Did Your Team Actually Hear You? The companion habit is the readback: do not assume the message landed, ask one person to say back what they are doing next. The gap you find is the gap that would have cost you an hour.
Tone is the loudest signal you send
Here is the part leaders underestimate. In an incident, the team reads your tone before they read your words. A calm, clear leader produces a calm, clear room; a leader who is visibly rattled produces noise, second-guessing and slower decisions, whatever the words actually say. The team mirrors you under pressure, the uncomfortable and useful idea in Your Team Mirrors You. Your composure is not a personality trait on the night. It is a piece of operational equipment, and it is the one only you can bring.
Cadence: the rhythm that beats silence
The other half of communication is timing, and the established practice is simpler than most teams run. Acknowledge fast: a first update within fifteen to thirty minutes of declaring, even when you have nothing but 'we know, we are on it'. Then hold a steady cadence, a progress update roughly every thirty to sixty minutes internally, and tighter, every fifteen to thirty minutes, on anything customer-facing. Keep the cadence even when there is no news; 'still working, next update by X' is itself information, and it stops the vacuum that people fill with rumour. And the person writing the updates should not be the person debugging the problem: that is exactly why the communications lead is a named role, so the wording gets someone's full attention while the fixing gets someone else's.
Where this breaks
Operational communication breaks when it becomes performance: the hourly update that exists to look in control rather than to inform, the status note written for the audience instead of the responder, the meeting that generates the feeling of progress without any. If a message does not improve someone's clarity, their next decision, or the reliability of the outcome, it is theatre, and theatre in an incident is worse than silence because it spends attention people do not have to spare, the argument in Skip the Theatre. The opposite failure is just as common: so many channels, updates and side-threads that the signal drowns in its own noise. More communication is not better communication. Clearer, on a rhythm, to the right audience, is.
More communication is not better communication. |
After the incident: the write-up is communication too
The last message is the review, and it is where trust is won or lost. A blameless write-up that is honest about what happened, what the impact was, and what will change, told without hunting for a person to blame, is what turns an outage into credibility. The audiences persist: responders want the timeline and the actions, leadership wants the risk and the fix, customers want to know you understood and are closing the gap. Write it for them, not for the file.
What to do on Monday
Communication is a habit, not a talent. Build these this week.
Use one structure for every update. Context, intent, headline. The same shape every time, so people stop guessing where the important part is.
Write for three audiences on purpose. Before you send, ask which of responders, leadership and customers this is for, and whether it says the right thing for them.
Set a cadence and keep it. Acknowledge inside thirty minutes, then update on a fixed rhythm, even when the update is 'no change, next by X'.
Name a communications lead. On the next real incident, give the updates to someone who is not fixing the problem.
Frequently asked questions
How often should we send updates during an incident? A useful default is a first acknowledgement within fifteen to thirty minutes, then internal updates every thirty to sixty minutes and customer-facing updates every fifteen to thirty, holding the cadence even when there is no change. Adjust to severity, but keep the rhythm predictable.
How much detail do executives actually want? The headline and the shape of the risk: how bad, who is affected, business impact, and when the next update lands. Keep the diagnostics for the responder channel; an executive buried in stack traces cannot help you and will start asking questions that pull focus.
Who should write incident communications? Not the person debugging. A named communications lead writes the updates so the wording has someone's full attention, while the incident commander owns the decision of what gets said and when.
What is the single highest-leverage change? Pick one update structure and use it every time. Consistency does more for comprehension under pressure than any individual well-crafted message.
Did Your Team Actually Hear You? Context, intent, headline: make the message land.
Your Team Mirrors You Why your tone sets the team's under pressure.
Skip the Theatre If it does not improve clarity or reliability, it is theatre.
|
Work with me An honest read on what your observability is actually doing. | |
|
If you lead observability in a regulated enterprise, I run a fixed-scope Observability Assessment for senior IT and engineering leaders. It ends in a written roadmap and a readout, not a sales deck.
Not ready to talk? Start with a free chapter of Metrics & Mayhem. |
