Knowledgebase

Handling Production Incidents Print

  • backenddevelopment, backend, downtime, guide, howto, solution, zillionkinghost, hosting
  • 0

When something breaks.

WHAT TO DO FIRST

Establish scope: what is affected, and how many.

WHAT TO DO SECOND

Communicate, early.

WHY

Silence during an outage is worse than the outage.

WHAT TO ASK

What changed?

WHY

Most incidents follow a change, and reverting is frequently faster than diagnosing.

WHAT TO DO

Revert first, investigate afterwards.

WHAT TO RECORD DURING

Times, observations and actions.

WHY

Memory is unreliable afterwards, and the timeline is what the review depends on.

WHAT TO AVOID

Several people changing things simultaneously Changing several things at once Persisting alone while service is down

WHAT TO DELEGATE

Communication, so whoever is diagnosing can diagnose.

WHAT TO DO AFTERWARDS

A review, without blame, establishing what happened and what allowed it.

WHAT TO PRODUCE

Actions preventing recurrence, with owners.

WHAT TO AVOID

Reviews producing no changes Blame, which ensures the next incident is concealed


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot