Knowing before your customers do.
WHAT TO MONITOR ON ANY SERVER
Whether it responds Disk space Load and memory Whether key services are running Certificate expiry Backup success
WHY EXTERNAL MONITORING MATTERS MOST
A machine monitoring itself cannot report that it is unreachable.
WHAT TO CHECK FROM OUTSIDE
That sites respond, with the expected content.
WHY CONTENT RATHER THAN A RESPONSE CODE
A server returning a friendly error page returns a successful code.
WHAT TO ALERT ON
Conditions requiring action.
WHAT NOT TO ALERT ON
Anything that resolves itself, or can wait.
WHY
Alerts that are routinely ignored train people to ignore alerts.
WHAT THRESHOLDS TO SET
Ones giving time to act before failure.
WHAT THAT MEANS FOR DISK
Warn well before full, not at ninety-nine per cent.
WHY
A disk filling at a steady rate gives hours of warning if you ask for it.
WHAT TO ROUTE ALERTS TO
Somewhere actually seen.
WHAT TO TEST
That alerts arrive, by deliberately triggering one.
WHY
Untested alerting fails silently, which is worse than none.
WHAT TO REVIEW
Alerts that fired and were dismissed, monthly.
WHAT TO ADJUST
Thresholds that produce noise.