When the other system is down.
WHAT TO DECIDE IN ADVANCE
Whether your system continues without it.
WHAT THE OPTIONS ARE
Fail the operation Queue it for later Proceed with a degraded result Fall back to another provider
WHAT TO CHOOSE
Depends on whether the operation can wait.
WHAT QUEUING REQUIRES
Somewhere durable to hold the work Processing when the provider returns Visibility of the backlog
WHAT DEGRADED OPERATION LOOKS LIKE
Showing cached data with a notice Accepting an order and confirming later
WHAT TO TELL USERS
That something is delayed, plainly.
WHY
Silence produces repeated attempts and support contacts.
WHAT TO AVOID
Retrying continuously against a down service.
WHY
It delays their recovery and exhausts yours.
WHAT A CIRCUIT BREAKER PREVENTS
Exactly that.
WHAT TO MONITOR
Provider availability, independently of their status page.
WHY INDEPENDENTLY
Status pages lag, and sometimes understate.
WHAT TO PREPARE
A documented procedure for each critical provider.
WHAT IT SHOULD CONTAIN
How to detect it What to switch off What to tell customers How to process the backlog afterwards
WHAT TO DO AFTERWARDS
Reconcile, since partial processing is likely.