Watching the infrastructure.
WHAT TO MONITOR
Link status and utilisation Error and discard rates Latency between points
Device health: temperature, power supplies, fans
Configuration changes
WHY ERROR RATES MATTER
Rising errors precede failure, and cause degraded performance before anything appears down.
WHAT THAT MEANS
A link passing traffic with rising errors is failing quietly.
WHAT TO ALERT ON
Loss of a link Loss of redundancy Utilisation approaching capacity Rising error rates Unexpected configuration change
WHAT TO BASELINE
Normal traffic patterns, so departures are visible.
WHAT TO RECORD
Utilisation over time, for capacity planning.
WHAT ELSE TO MONITOR
Whether devices are reachable by their management path Whether backups of configurations are current
THAT LAST ONE
Device configurations should be backed up, so a failed device can be replaced quickly.
WHAT TO TEST
That a configuration backup can actually restore a device.
WHAT TO DOCUMENT
The topology, accurately, and update it with every change.