Keeping the service available.
WHY IT MATTERS MORE THAN IN OTHER SOFTWARE
Customers pay continuously, and an outage is visibly a failure to deliver what they paid for.
WHAT TO MEASURE
Availability, from outside your own infrastructure Error rate Response time
WHY FROM OUTSIDE
Internal monitoring reports that your servers are healthy while customers cannot reach them.
WHAT TO MONITOR
The critical user journeys, continuously.
WHAT SYNTHETIC CHECKS SHOULD DO
Sign in and perform the core action, repeatedly.
WHAT ALERTING SHOULD WAKE SOMEONE FOR
Customers unable to use the product.
WHAT IT SHOULD NOT
Anything that can wait until morning.
WHY THAT DISTINCTION
Alert fatigue means real incidents are missed.
WHAT TO PREPARE BEFORE YOU NEED IT
A way to communicate during an outage, hosted separately A rollback procedure, tested Someone responsible
WHY HOSTED SEPARATELY
A status page on the same infrastructure is down during the outage.
WHAT TO COMMUNICATE DURING AN INCIDENT
That you are aware What is affected When you will next update
WHAT NOT TO DO
Go silent Speculate about causes Understate the impact
WHAT TO PUBLISH AFTERWARDS
What happened, and what prevents recurrence.
WHY
Handled well, an incident increases trust.