How Do You Design a Status Page People Trust?
What makes a status page actually useful?
Answering one question before anything else: is it you or is it me. Nielsen Norman Group's first usability heuristic is visibility of system status, which it defines as keeping users informed about what is going on through appropriate feedback within a reasonable amount of time. A status page is that heuristic applied to your whole product.
Most status pages fail at the top of the page. They show a grid of component names, a legend, and a 90 day uptime chart, and the one sentence a worried customer needs has to be assembled from all of it.
So this is a design walkthrough rather than a tooling comparison: what the page should say, how the indicators should look, and what an update needs to contain to be worth publishing.
Why does a status page belong in your design system?
Because it is read at the worst possible moment. Every other page on your site is browsed by someone calm. This one is opened by someone whose work has just stopped, often on a phone, often while they are also writing a message to their own team.
That audience changes every design decision. Hierarchy matters more than completeness, plain language matters more than precision, and anything requiring interpretation is a failure. The page has one job and it is comprehension under stress.
It is also a trust artefact. A status page that is honest during an incident does more for credibility than any amount of marketing copy about reliability, and one that stays green while customers are clearly broken does lasting damage. Our notes on designing a trust centre cover the neighbouring page.
What should the page show above everything else?
One sentence of plain state, at the top, in a size you cannot miss. Either everything is working, or something specific is not. That sentence should be readable without scrolling, without a legend, and without counting coloured squares.
Underneath it, the current incident if there is one, with its most recent update first. Reverse chronological order within an incident is right, because the newest information is the only information most readers need, and the history is for people who arrive late.
Everything else is secondary: component breakdowns, uptime history, scheduled maintenance, subscription options. Those earn their place lower down. The common mistake is leading with the component grid, which is a data structure rather than an answer.
How many status levels should you have?
As few as you can act on. Every level you add is a judgement call someone has to make at three in the morning, and levels that are hard to distinguish get used inconsistently, which destroys their meaning faster than having too few.
In practice most teams need working, degraded, and down, plus a separate state for planned maintenance. Degraded is the one that earns its keep, because partial failure is the common case and forcing it into either working or down produces a page that lies in both directions.
Define each level in writing before you need it, and define it by customer experience rather than by infrastructure. Degraded should mean something a customer would notice, not something a dashboard noticed. Our notes on building an incident response plan cover writing those definitions down.
Why can status not be colour alone?
Because a meaningful number of readers will not see the difference. WCAG success criterion 1.4.1, Use of Color, is a Level A requirement and states that colour is not used as the only visual means of conveying information, indicating an action, prompting a response, or distinguishing a visual element.
Level A is the baseline, not the stretch goal. A row of green and amber dots with no other distinguishing feature fails it, and on a status page that failure is particularly cruel, because the whole point of the page is to convey one piece of information quickly.
The fix is cheap. Pair every colour with a shape, an icon, or a word. The WCAG understanding document gives the same pattern for forms, describing required fields marked with both red text and an accompanying icon so people who cannot perceive the colour difference still get the message.
What contrast do status indicators need?
At least 3 to 1 against what sits next to them. WCAG success criterion 1.4.11, Non-text Contrast, is Level AA and requires a contrast ratio of at least 3 to 1 against adjacent colours for the visual information required to identify user interface components and states.
It also covers graphics that carry meaning. The criterion applies to parts of graphics required to understand the content, which is exactly what a status dot or an outage icon is. A pale green circle on a white card is decorative, not informative.
This catches a lot of otherwise careful designs, because status colours tend to be chosen for a brand palette rather than for legibility against the surface they sit on. Check them against the actual card background, in both light and dark mode. Our notes on contrast measurement cover how to check properly.
What should an incident update say?
What is broken, who it affects, what you are doing, and when you will next speak. Four things, in plain language. Nielsen Norman's ninth heuristic applies directly here: error messages should be expressed in plain language with no error codes, precisely indicate the problem, and constructively suggest a solution.
The next update time is the part teams leave out and customers want most. Without it, every reader has to decide for themselves how often to refresh, and your support inbox fills with people asking for an estimate you could have published once.
Avoid the vocabulary of the inside. Nobody outside your team knows what a queue consumer is, and during an incident is the worst time to teach them. Describe the symptom the customer is experiencing, then the cause in ordinary words if it helps. The same rules apply to error messages inside the product.
Should you publish a postmortem?
For anything that cost customers real time, yes. Google's SRE book defines a postmortem as a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root causes, and the follow-up actions to prevent the incident from recurring. That is a good outline for a public version too.
Google's framing is blameless by design. The SRE book calls blameless postmortems a tenet of SRE culture and describes assuming everyone involved had good intentions and did the right thing with the information they had. A public postmortem that reads as blame does not build trust, it signals a bad internal culture.
Decide the trigger in advance. The SRE book lists criteria including user-visible downtime or degradation beyond a threshold, data loss of any kind, on-call intervention such as a rollback, and resolution time above a threshold, and advises teams to define postmortem criteria before an incident occurs. Deciding during the incident guarantees inconsistency.
Where should the status page live?
Somewhere that survives your own outage. If the page is served by the same infrastructure as the product, it will be down exactly when it is needed, which is the oldest mistake in this category and still a common one.
A separate subdomain on separate hosting is the standard answer, and it should be reachable with no login. A status page behind authentication is useless to the person who cannot log in, which is frequently the symptom being reported.
Link to it from the places people look during a failure: the footer, the help centre, and the error state in the product itself. Nobody searches for a status page before they need one, so discovery has to happen at the moment of trouble. Our notes on uptime monitoring cover the detection side that feeds it.
What do we build for clients?
A simple page, honestly maintained, over an elaborate one nobody updates. The most common failure we see is not a badly designed status page, it is a beautiful one that has said all systems operational for two years because updating it was never anyone's job.
So we design around the update workflow first. Who can post, how fast, from where, and with what template. If posting an update takes more than a couple of minutes under pressure, it will not happen, and an unmaintained status page is worse than none because it actively misleads.
We also keep the component list short. A page listing thirty internal services pushes the interpretation work onto the customer. Group by what customers actually use, name those groups in their language, and keep the infrastructure detail in your own monitoring.
What makes a status page worth building at all?
Being believed. The value is entirely in whether customers trust what it says, and that trust is built by being accurate during bad weeks rather than by looking good during calm ones. One honest degraded label earns more than a year of green.
Our bet is that status pages keep getting more important as more B2B products become dependencies of other products. When your customer's customer is affected, the page stops being a courtesy and becomes part of how your buyers assess risk during procurement.
If you are designing one, or you have one that nobody updates, we are happy to look at it with you and say what we would change. Find us at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.