Knowing the state of everything.
WHAT TO MONITOR
Availability, from outside
Capacity: processor, memory, storage, network
Saturation and queueing Error rates Latency Certificate and domain expiry Backup completion Replication lag Hardware health
WHY FROM OUTSIDE
A system cannot reliably report its own unavailability.
WHAT TO ALERT ON
Conditions requiring action now.
WHAT NOT TO ALERT ON
Anything nobody will act on immediately.
WHY
Alerts people ignore train them to ignore alerts.
WHAT TO REVIEW
Every alert that fired, and whether it was actionable.
WHAT TO DELETE
Alerts that were not.
WHAT EVERY ALERT NEEDS
A documented response.
WHY
An alert with no documented response is answered by improvisation.
WHAT TO MONITOR BEYOND TECHNICAL HEALTH
Business outcomes: orders, sign-ups, mail delivered.
WHY
Every technical check can pass while nothing useful is happening.
WHAT TO PLACE OUTSIDE YOUR OWN INFRASTRUCTURE
The monitoring itself, and the alerting path.
WHY
Otherwise an outage takes your monitoring with it.