Controlling what is altered.
WHY IT MATTERS
Most infrastructure incidents follow a change.
WHAT EVERY CHANGE NEEDS
A description An assessment of risk and impact A plan A way to reverse it Approval proportionate to risk Notification of those affected A record of what was done
WHAT TO SCHEDULE IN WINDOWS
Anything risky.
WHAT TO FREEZE
Changes during periods of critical business activity.
WHAT TO REQUIRE FOR WORK ON REDUNDANT SYSTEMS
Awareness that redundancy is reduced during the work.
WHY
While one path is maintained, the other is a single point of failure.
WHAT TO NEVER DO
Work on both paths simultaneously Assume redundancy means the work carries no risk
WHAT TO RECORD ALWAYS
What changed, when, and by whom.
WHY
It answers the first question in every incident.
WHAT TO AUTOMATE
Configuration, as code, reviewed and versioned.
WHAT THAT PROVIDES
Repeatability, review, and a record of every change.
WHAT TO TEST
Changes, somewhere other than production, where practical.