Assuming things break.
THE PRINCIPLE
Everything fails eventually. Design so failure is contained.
WHAT TO ASSUME
Networks are unreliable External services fail Disks fill Processes are killed Servers restart
WHAT TO BUILD IN
Timeouts on every external call Retries for temporary failures, with increasing delay Graceful behaviour when a dependency is unavailable
WHAT GRACEFUL MEANS
The rest of the application continues working.
WHAT THAT LOOKS LIKE
A recommendation service failing, and the page still loading without recommendations.
WHAT NOT TO DO
Let one non-essential dependency take down everything.
WHAT TO AVOID
Retrying indefinitely Retrying failures that will not resolve Retrying in a way that worsens an overload
WHAT TO IMPLEMENT FOR REPEATED FAILURE
Stopping calls to a failing service temporarily, rather than continuing to wait.
WHAT TO MAKE SAFE
Operations that may be repeated, since retries happen.
WHAT TO TEST
Behaviour when a dependency is unavailable.
HOW
Disable it deliberately, in a test environment.