Surviving a server failure.
WHAT THE OPTIONS ARE
A replica promoted manually Automatic failover with a manager A cluster where several nodes accept writes Managed services handling it for you
WHAT MANUAL PROMOTION PROVIDES
Simplicity, and human judgement.
WHAT IT COSTS
Downtime while someone acts.
WHAT AUTOMATIC FAILOVER PROVIDES
Faster recovery.
WHAT IT RISKS
Promoting during a network problem rather than a real failure.
WHAT THAT PRODUCES
Two servers believing they are primary, accepting conflicting writes.
WHY THAT IS THE WORST OUTCOME
Reconciling divergent data is far harder than downtime.
WHAT PREVENTS IT
A mechanism ensuring only one can be primary.
WHAT THE APPLICATION NEEDS
A way to reach whichever server is current.
WHAT PROVIDES THAT
A proxy, a virtual address, or service discovery.
WHY THAT IS FREQUENTLY THE HARD PART
Failing over the database achieves nothing if applications still point at the old one.
WHAT TO ESTABLISH BEFORE BUILDING ANY OF IT
How much downtime is actually acceptable.
WHY
The honest answer is often more than the complexity costs.
WHAT TO PREFER WHEN STARTING
A well-tested replica and a documented promotion procedure.
WHAT TO PRACTISE
The failover, deliberately.