Incident communication and running a status page
How to stop answering the same question a hundred times during an outage: the first-post threshold, update cadence, severity levels and where the page should live.
At 9:12 on a Monday the dealer ordering portal stops loading. Within nine minutes fourteen messages hit the support inbox, all asking the same question. Three different answers go out: back in fifteen minutes, nothing looks wrong on our side, we are checking. Engineering fixes it in forty minutes. In those forty minutes support resolves nothing else, two customers ask about it in public, and by evening two dealers have held their orders because they could not tell their own customers anything. The outage lasted forty minutes; the damage lasted two days.
What follows builds incident communication end to end: what a status page solves and what it does not, when the first post goes out, what belongs in an outage notice, update cadence, how severity levels are defined, where the page should physically live, who writes and who approves, who gets notified, where transparency stops, and what happens once the incident closes.
What does a status page solve, and what does it not?
A status page does not shorten the outage. Its job is to stop the same question being asked a hundred times while the outage runs. With one address everyone can look at, support redirects instead of answering, sales knows what to tell a customer, and managers stop interrupting engineers. Most tickets opened during an incident are not requests for a fix but versions of "is it just me?", and they arrive in the minutes when the team is most stretched. Answering them before they are opened buys back the thing you need most during an outage: attention.
What it does not solve is equally clear. A customer with a contractual commitment, a partner whose integration has stopped, a project going live that morning — none of them are served by a page. The page scales broadcast; it does not replace one-to-one contact. Handling the message flood on one screen is covered in the shared inbox.
When does the first post have to go out?
Most teams hold the first post until the cause is known, which ties communication to engineering's pace while the customer is already living the problem. The working threshold: if anyone outside your company can see the outage, the post goes up. The first post contains no cause — only what is not working, who is affected, and when you will write again. There is a number for this, independent of how long the fix takes: time to first post. Resolution time belongs to engineering, time to first post belongs to communication, and teams that track both run incidents with noticeably less chaos.
The usual objection is that posting early alerts customers who had not noticed; it rarely holds. The first person to notice something is broken is the user staring at the screen, and when you say nothing they assume the problem is theirs, switch browsers, reset a password, then open a ticket. Telling customers before they ask is the subject of proactive customer support.
What belongs in an outage notice?
A notice is a status report, not an apology letter. Keep it short and flat; customers scan it, they do not read it. The seven components below give the same skeleton to the first post and to every update after it.
- Impact line: What is not working, in the words the user sees: cannot log in, orders not saving, reports not opening.
- Scope: Who is affected — everyone, one region, one module. Leaving scope out is the most expensive omission on this list.
- Still working: The components that are unaffected. This line alone stops unaffected customers from opening tickets.
- Workaround: Any alternative path — the mobile app still works, the upload can wait until later.
- Current state: Investigating, identified, fix in progress, monitoring. Four words that mean the same thing to everyone.
- What not to do: Do not resubmit the order, do not retry the payment. Without this line you spend the hour after recovery clearing duplicates.
- Next update: With a time on it. You will post then even if nothing has changed.
The one thing that does not belong is a guess at the cause. The first suspect is usually wrong, and retracting a published cause costs far more than never naming one. The same discipline applies one-to-one in delivering bad news to customers.
How is update cadence set?
Silence is the most expensive option during an incident. The gap between a post at 10:20 and the next at 11:45 is where the entire ticket pile gets created. Cadence is set by commitment, not progress: every post names the time of the next one, and you keep it.
Workable intervals scale with severity. Half-hourly during a full outage, hourly during a partial one, every few hours for degraded performance is sustainable for most teams. Keeping the interval you promised matters more than making it short; a missed update time tells the customer the incident has escaped your control.
What do you write when there is no news?
There is no such thing as an empty update, only a badly written one. Even with no progress, three things can be said: what you have ruled out, who is working on what right now, and when the next post lands. "Work continues" says none of them; "we have ruled out the database and are now working with the payment provider" carries information without progress.
How do you define severity levels?
Levels built on adjectives fail. "Critical" and "high" mean different things to different people. A working definition ties each level to two concrete outcomes: who gets woken up, and which channel is used.
| Level | What it means | Notification and who steps in |
|---|---|---|
| Full outage | The core service works for nobody | Status page, email, calls to key accounts; on-call team and a manager |
| Partial outage | One module or one region is down | Status page and email to affected accounts; on-call team |
| Degraded performance | Working but slow, intermittent errors | Component note on the status page; internal channel |
| Planned maintenance | A window announced in advance | Advance email and a calendar entry on the page; maintenance owner |
| Single-account issue | Affects one customer only | Never on the status page; handled as a support ticket |
The last row is the most argued over, and it belongs where it is. A single-account problem posted publicly turns the page into noise, and a noisy page goes unread. How a breached threshold routes itself to the right person automatically is covered in escalation automation.
The level most often misjudged is degraded performance. When a system stops entirely, users understand and wait; when it slows down they retry, switch browsers, then write in. Slowness generates more tickets per affected user than a full outage, so teams that keep degradation off the page hide the incident type that produces the most contact.
Where should the status page live?
This is the most common mistake and the most quietly punished: hosting the status page on the system whose status it reports. When the application goes down, a page on the same servers goes with it; when your DNS is affected, a subdomain page is unreachable too. Independence is not a technical detail — it is the whole reason the page exists.
Keep the page on separate hosting and, where possible, a separate domain, and print its address on invoices, in contracts and in support signatures in advance — nobody should be searching for it mid-incident. The data-side equivalent of the same logic is in backup and disaster recovery.
Who writes, and who approves?
Two roles have to be separated during an incident: the person fixing it and the person describing it. When one person does both, either the update is late or the fix is slow. In a small team that means two names on the rota, not two departments.
Approval is what usually delays the first post. Routing every sentence through management or legal is reasonable in peacetime and disastrous mid-incident. Approve the template, not the post: write three on a calm day, get them signed off, fill in the blanks when it matters. Language belongs in the template too: use the name the customer knows — order confirmations are delayed, not the queue service is down — and avoid the passive voice, because "an interruption has been experienced" reads like an event nobody owns. For incidents that grow into brand-level events, see crisis communication.
Everyone, or only the affected?
A status page is a pull channel: whoever wonders comes and looks. A notification is a push channel, and every push has a cost. An outage email to an unaffected customer reminds them you have outages and lowers the odds your next real alert gets read. How alerts go blunt is covered in notification fatigue.
The arrangement that works is layered. An in-app banner shows only to people currently trying to use the thing, which makes it your most accurate channel. Email goes to subscribers. Accounts with a contractual commitment get a direct message with a name on it, and that list must exist beforehand. What those commitments cover is the subject of defining service levels in a customer contract.
The real price of an outage is not the minutes it lasts; it is that for those minutes your customer cannot tell their own customer anything.
Where does transparency stop?
The standard advice is to publish everything, and it has a limit. Teams that post every three-minute blip make the page unreadable, and the subscriber list dies quietly. Define the threshold in advance: how many users, for how many minutes, before it goes public. Log what falls below the line internally.
The opposite mistake is more dangerous. A company that keeps real outages off the page to preserve a green history ends up showing "all systems operational" to a customer living through one, and the page loses its value in that second. The test is simple: if customers noticed, it goes up. A status page is not a report card. Security incidents follow separate rules — publish only what is confirmed, and where personal data may be involved, take advice from your own legal counsel on notification duties.
What happens after the incident closes?
A resolved notice is not the end; a post-incident report follows. It stays short: what happened, who was affected and for how long, why, what you are changing, and what you are deliberately not changing. That last item is unusual and the most reassuring, because nobody believes a company that claims to have changed everything after every outage.
Knowing who reads it sets the tone. It is written for the IT lead and the procurement reviewer on the other side, not for your engineers: it goes into a vendor file and comes back out at the next renewal. The date and owner of each remediation matter more than the technical detail, and it should say what you changed rather than that it will never happen again. Whether customers can help themselves during an incident depends on self-service support and knowledge bases.
Four numbers are worth tracking. Time to first post is the only real performance measure on the communication side. Tickets opened per incident tells you whether the page is working; after a good post that number falls. Subscriber churn warns that the page has turned into noise, and the days between the incident and the published report measure whether you keep your word.
Where to start
The minimum system takes an afternoon: a page hosted somewhere else, three approved templates, a five-line severity table, a ready list of committed accounts, two names on the rota. Rehearse it before a real outage — measure how many minutes the first post takes and make that a target. Planning outages, data loss and key-person risk together is the subject of business continuity planning.
The hard part of incident communication is not writing the text; it is knowing who is affected while the incident runs. Rocketly keeps customer records, contact history, tasks and reminders and team notes in one place, so the affected-account list takes minutes; open a free account and set up your own announcement flow.