Watching the fabric.
WHAT TO MONITOR
Link status and utilisation Error, discard and retransmission counters
Device health: temperature, power supplies, fans
Routing session status Configuration changes
WHY ERROR COUNTERS MATTER MOST
They rise before a link fails, and they degrade performance while everything appears up.
WHAT TO ALERT ON
Loss of a link Loss of redundancy Rising errors Utilisation approaching capacity Unexpected configuration change
WHY LOSS OF REDUNDANCY SPECIFICALLY
Running on a single path after a silent failure is the state in which the next failure becomes an outage.
WHAT TO BASELINE
Normal traffic patterns, so departures are visible.
WHAT DEPARTURES MAY INDICATE
An attack A misconfiguration A compromised system sending outbound
WHAT TO COLLECT
Flow data, showing what is talking to what.
WHY
It answers questions nothing else can during an incident.
WHAT TO RETAIN
Enough history to investigate.
WHAT TO BACK UP
Device configurations, automatically.
WHAT TO TEST
That a configuration backup can actually restore a device.