Deciding what matters.
WHAT TO MONITOR FIRST
Whether the service works, from outside.
WHY FROM OUTSIDE
A system cannot reliably report its own unavailability.
WHAT ELSE TO MONITOR
Error rates Response times
Resource use: processor, memory, disk
Queue depth, where queues exist Certificate and domain expiry Whether scheduled tasks ran Whether backups completed
WHAT THOSE LAST THREE HAVE IN COMMON
They fail silently, and nobody notices for weeks.
WHAT TO MONITOR THAT IS BUSINESS-SPECIFIC
Whether orders are being placed Whether sign-ups are occurring Whether payments are succeeding
WHY THOSE MATTER MOST
A technically healthy system that has stopped taking orders is an outage.
WHAT NOT TO MONITOR
Everything, producing noise nobody reads.
WHAT TO ALERT ON
Things requiring action now.
WHAT TO RECORD WITHOUT ALERTING
Everything else, for investigation.
WHAT TO REVIEW
Whether your monitoring would have caught your last incident.