A status page is the customer-facing record of service availability. Done well, it builds trust during incidents — customers see the company is on it, communicating clearly, and recovering quickly. Done poorly, it's a tool for cover-up that erodes trust faster than the incident itself.
This page covers what makes status pages actually work.
Each major service or feature has a status:
The granularity matters. "API is down" is more useful than "everything is down" or "subsystem 47b is down."
Current issues with timeline of updates. Most recent on top.
History of resolved issues. Useful for customers evaluating reliability.
Planned downtime announced in advance.
Customers can subscribe to email/SMS/RSS updates.
The status reflects reality. If the API is down, the page says so within minutes.
The temptation: keep it green to avoid bad metrics. Customers notice; trust erodes.
Updates posted within minutes of incident start. Customers shouldn't have to call support to find out something's wrong — they should already know from the status page.
"We are investigating elevated error rates on the orders API. ~30% of requests affected. ETA unknown; updating in 15 minutes."
Better than: "We are investigating an issue."
"We don't yet know the cause" is OK. "We are confident the issue is resolved" with a cause stated is fine. Hedging when it isn't appropriate erodes trust.
"We are committed to providing world-class service while we investigate this temporary connectivity blip."
Customer translation: "Something is broken; they're using marketing words."
"We are experiencing intermittent issues with some services."
Useless. Which services? What issues? What's affected?
The page goes red 30 minutes after customers start calling support. Clearly wasn't actively monitored.
Customers know the service is broken; status page is green. Confidence destroyed.
Resolved incidents disappear from the page. No record; no learning visible.
The dominant choice. Established; full-featured.
Modern alternatives. Comparable features; sometimes better pricing.
Open-source. For privacy-sensitive companies.
Built on top of monitoring data. Reasonable for specific needs.
For most companies, hosted services are right. The page itself is rarely a competitive differentiator.
During an incident:
"We are investigating reports of [specific symptom]. Affected: [scope]. We will update in 15 minutes."
Even if you don't have details. Acknowledge; commit to next update.
Every 15-30 minutes during active incidents. Even if nothing has changed: "still investigating; ETA unknown."
When the immediate impact is contained: "We have applied a mitigation; users should see normal behavior. We continue to investigate root cause."
"The incident is resolved. We will publish a postmortem within [N business days]."
The full report: timeline, root cause, what we're doing differently. Customer-facing version may be shorter than internal version.
Match status to actual impact:
Don't over-grade (every issue is "major") or under-grade (everything is "degraded" until resolved).
Some companies have:
The internal page can have more detail, more components, more granular status.
Status pages are about trust as much as information. Customers extend trust to companies that:
A company that hides outages is presumed less trustworthy than one that shares them and explains. Even if the latter has more outages.
This is counterintuitive but important. Cover-up costs more than the original incident.