Choosing what to deploy.
WHAT TO COMPARE ON
Performance on the metric that matters Performance on the segments that matter Latency and resource cost Interpretability Maintenance burden
WHY SEGMENTS SPECIFICALLY
Aggregate performance conceals failure on important subgroups.
WHAT TO ALWAYS CHECK
Performance by group: customer type, region, volume band.
WHAT TO WEIGH AGAINST ACCURACY
Inference cost, at expected volume Training cost and frequency Complexity the team must maintain
WHY THAT LAST POINT IS UNDERRATED
A marginally better model nobody can maintain is worse than a simpler one.
WHAT A SMALL IMPROVEMENT IS WORTH
Frequently nothing, once deployment cost is counted.
WHAT TO ESTABLISH
The improvement required to justify the change.
WHAT TO TEST BEFORE COMMITTING
Behaviour on unusual inputs Behaviour when features are missing Stability across retraining
WHY STABILITY MATTERS
A model whose predictions swing between retrainings undermines trust.
WHAT TO DOCUMENT
Why this model was chosen over the alternatives.
WHAT TO RETAIN
The comparison, so the decision can be revisited.