Assuming things break.
WHAT WILL FAIL
Sources unavailable Credentials expiring Schemas changing Volumes exceeding capacity Downstream systems rejecting data The orchestrator itself
WHAT TO DECIDE PER TASK
Whether a failure should stop everything, or only its branch.
WHAT RETRIES SUIT
Transient failures: timeouts, temporary unavailability.
WHAT THEY DO NOT FIX
Bad data, or a genuine error.
WHY THAT MATTERS
Retrying a task that fails deterministically wastes time and delays the alert.
WHAT TO CONFIGURE
A limited number of retries, with increasing delay.
WHAT TO ALERT ON
Final failure, not each retry.
WHAT A TIMEOUT PROTECTS AGAINST
A task hanging indefinitely and blocking everything behind it.
WHAT TO SET
A timeout on every task, above normal duration.
WHAT PARTIAL FAILURE REQUIRES
Either completing or leaving no trace, never half-written output.
WHAT TO USE
Staging and atomic promotion.
WHAT TO DOCUMENT PER PIPELINE
What to do when it fails, specifically.
WHY
Whoever is called at night did not build it.