Controlled deployment.
WHAT THE STRATEGIES ARE
Shadow deployment: running alongside, predictions not used
Canary: a small proportion of traffic
Gradual rollout Comparison testing against the current model
WHAT SHADOW DEPLOYMENT PROVIDES
Real-world behaviour with no risk.
WHAT IT REVEALS
Latency under real load Inputs unlike the evaluation set Failures in the serving path
WHY IT IS UNDERUSED
It requires infrastructure, and feels like extra work.
WHAT IT PREVENTS
Discovering all of that in production.
WHAT COMPARISON TESTING MEASURES
Whether the new model produces better outcomes, not merely better offline metrics.
WHY THAT DISTINCTION MATTERS
Offline improvement frequently does not translate.
WHAT TO MEASURE
The business outcome, over enough time.
WHAT TO PREPARE BEFORE RELEASE
A way to revert immediately Defined criteria for reverting Monitoring of the relevant metrics
WHAT TO NEVER DO
Deploy without the ability to revert.
WHAT TO KEEP AVAILABLE
The previous model, deployable.
WHAT TO ANNOUNCE
That predictions may change, and why.