Designing for Failure Print

  • devopsinfrastructure, devops, guide, howto, solution, zillionkinghost, hosting, support
  • 0

Assuming things break.

THE PRINCIPLE

Everything fails eventually. Design so failure is contained.

WHAT TO ASSUME

Networks are unreliable External services fail Disks fill Processes are killed Servers restart

WHAT TO BUILD IN

Timeouts on every external call Retries for temporary failures, with increasing delay Graceful behaviour when a dependency is unavailable

WHAT GRACEFUL MEANS

The rest of the application continues working.

WHAT THAT LOOKS LIKE

A recommendation service failing, and the page still loading without recommendations.

WHAT NOT TO DO

Let one non-essential dependency take down everything.

WHAT TO AVOID

Retrying indefinitely Retrying failures that will not resolve Retrying in a way that worsens an overload

WHAT TO IMPLEMENT FOR REPEATED FAILURE

Stopping calls to a failing service temporarily, rather than continuing to wait.

WHAT TO MAKE SAFE

Operations that may be repeated, since retries happen.

WHAT TO TEST

Behaviour when a dependency is unavailable.

HOW

Disable it deliberately, in a test environment.


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot