Removing single points of failure.
WHAT THE PROBLEM IS
A load balancer distributing traffic is itself a single point of failure.
WHAT SOLUTIONS EXIST
A pair with a floating address, moved on failure Several balancers behind anycast Several addresses published, with clients trying alternatives A managed service handling it
WHAT A FLOATING ADDRESS REQUIRES
A protocol allowing one node to claim it, and reliable detection of failure.
WHAT SPLIT-BRAIN IS
Both nodes believing the other has failed, and both claiming the address.
WHAT CAUSES IT
Loss of communication between the pair, without actual failure.
WHAT MITIGATES IT
A third party to arbitrate, and fencing of the failed node.
WHAT TO TEST
Failover, deliberately, under load.
WHY
Failover that has never been exercised frequently does not work.
WHAT ELSE TO MAKE REDUNDANT
Upstream connectivity Power paths The configuration source
WHAT TO CHECK
That redundancy is genuinely independent rather than sharing a component.
WHAT TO MONITOR
Which node is active, and whether failover occurred unnoticed.