Knowing what is happening.
WHAT TO MONITOR
Availability, from outside Error rate Response time distribution Resource utilisation Queue depth and job failures Scheduled work completing Business outcomes
WHY FROM OUTSIDE
A service cannot reliably report its own unavailability.
WHY BUSINESS OUTCOMES
Every technical check can pass while a defect has stopped all orders.
WHAT PRODUCES NO ERRORS
A silently failing integration A scheduled job that stopped running A queue nobody consumes
WHAT TO ALERT ON
Unavailability Error rate rising Latency degrading Queue depth growing A scheduled job not completing A sharp fall in expected activity
WHERE TO SEND ALERTS
Somewhere independent of the system.
WHAT TO AVOID
Alerts nobody acts on Alerts so frequent they are ignored
WHAT TO REVIEW
Every alert that fired, and whether it was actionable.
WHAT TO RECORD
Metrics over time, so trends and regressions are visible.
WHAT TO TEST
That alerts actually reach someone who can act.