Knowledgebase

Designing Pipelines for Failure Print

  • dataengineering, data, domainrenewal, errors, troubleshooting, staging, guide, howto
  • 0

Assuming things break.

WHAT WILL FAIL

Sources unavailable Credentials expiring Schemas changing Volumes exceeding capacity Downstream systems rejecting data The orchestrator itself

WHAT TO DECIDE PER TASK

Whether a failure should stop everything, or only its branch.

WHAT RETRIES SUIT

Transient failures: timeouts, temporary unavailability.

WHAT THEY DO NOT FIX

Bad data, or a genuine error.

WHY THAT MATTERS

Retrying a task that fails deterministically wastes time and delays the alert.

WHAT TO CONFIGURE

A limited number of retries, with increasing delay.

WHAT TO ALERT ON

Final failure, not each retry.

WHAT A TIMEOUT PROTECTS AGAINST

A task hanging indefinitely and blocking everything behind it.

WHAT TO SET

A timeout on every task, above normal duration.

WHAT PARTIAL FAILURE REQUIRES

Either completing or leaving no trace, never half-written output.

WHAT TO USE

Staging and atomic promotion.

WHAT TO DOCUMENT PER PIPELINE

What to do when it fails, specifically.

WHY

Whoever is called at night did not build it.


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot