Watching many machines.
WHAT TO MONITOR PER MACHINE
Availability, from outside Load, memory, storage and storage operations
Service status: web, mail, database, panel
Certificate expiry Backup completion Security updates pending
WHAT TO MONITOR PER ACCOUNT
Resource use against limits Outbound mail volume Unusual process activity
WHY PER ACCOUNT MATTERS
Server-level figures tell you something is wrong, not who is causing it.
WHAT TO MONITOR FOR CUSTOMERS
That their sites actually load and return expected content.
WHY
A server can be healthy while a site is broken.
WHAT TO ALERT ON
Service unavailability Resource exhaustion approaching Backup failure Blocklist listing Certificate expiry
WHAT TO AVOID
Alerts nobody acts on Alerts firing for every machine when one shared dependency fails
WHAT TO DO ABOUT THAT SECOND POINT
Dependency-aware alerting, so the cause is reported rather than every symptom.
WHAT TO PLACE OUTSIDE YOUR OWN INFRASTRUCTURE
Monitoring and alerting.
WHAT TO REVIEW
Which alerts fired, and whether each was useful.