Clouds
DevOps·3 min read

Incident Communication: What to Say While Everything Is Broken

The status page, the Slack thread, and the executive update are three different documents. Writing them after the incident starts is how silence gets misread as an outage.


The worst incident I was part of was not the worst technically. A dependency returned errors for nineteen minutes, we had a fix ready in four. What made it memorable was that for eleven of those minutes nobody outside the on-call channel knew anything was happening, and by minute six the support inbox had decided we were hiding it.

Technical recovery was fine. Communication failed, and communication is what people remember.

Silence is not neutral

The instinct during an incident is to go quiet until you know something useful. It feels responsible — why announce uncertainty?

Because the audience is not only watching your output. They are watching the absence of it. Customers see a broken product and no acknowledgment. Sales sees a prospect asking a question with no reply. Your own engineers see a silent leadership and start guessing, which produces its own thread of speculation that then has to be corrected later.

Saying "we are investigating, next update in ten minutes" costs nothing and removes the inference that you do not know.

Three documents, three audiences

The mistake is trying to write one message that serves everyone. It will serve no one.

text
status page   customers, prospects, public   factual, no blame, no ETA theatre
slack         responders, support            operational, links, asks for help
exec update   leadership                     impact in business terms, not HTTP codes

A customer wants to know whether their data is safe and when service returns. A responder wants the current hypothesis and who owns what. An executive wants revenue exposure and whether to tell anyone else. Sending the responder version to customers produces "we believe the connection pool is exhausted," which is true and useless.

The skeleton that works under pressure

Four fields, filled in the first two minutes, updated on a stated cadence:

text
STATUS     Investigating | Identified | Monitoring | Resolved
IMPACT     What a user experiences, in their words
TIMELINE   Started <time>, next update by <time>
ACTION     What we are doing right now, in one line

The fourth field is the one that earns trust, because it is the only one that demonstrates activity rather than awareness. "We are failing over the primary database" tells a reader something is happening. "We are aware of the issue" tells them you read the same alert they did.

Commit to a cadence and keep it

Pick an interval — ten minutes is a good default for an active incident — and publish on it even when the update is "no change, still investigating." A steady heartbeat with empty content is more reassuring than a burst of detail followed by ninety minutes of nothing, because silence after detail is read as new bad news.

When you genuinely have nothing, say the true version: "We have applied the fix and are waiting on the connection pool to drain. No change expected for a few minutes." People can handle a slow wait. They cannot handle an unexplained one.

Do not schedule an ETA you have not tested

"We expect this to be resolved by 3pm" is a promise made with incomplete information under time pressure. It will be quoted back to you at 3:05.

Give a time for the next update — that you control — and describe the state of the fix rather than its completion. "Fix deployed, observing" is honest and does not expire.

After the incident, the communication is evidence too

The timeline you published becomes the raw material for the writeup, which is why incident comms belongs in the postmortem rather than in a separate log. If you cannot reconstruct what you told people and when, you also cannot tell whether the eleven silent minutes were a process gap or a tooling gap.

Summary

Communicate before you are certain. Keep separate artifacts for customers, responders, and leadership. Fill four fields — status, impact, timeline, action — commit to an update cadence you can actually keep, and promise times for updates rather than for resolution. Technical recovery closes the incident; communication is what determines whether anyone trusted you during it.

#incident-response#communication#on-call#postmortems#sre

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles