The summary.
THE DEFINITION OF THE TARGET IS THE PROJECT
Churn, default, failure and relevance each have several defensible definitions, and each produces a different model. Settle it before anything else.
TARGETING THE MOST LIKELY IS USUALLY THE WRONG STRATEGY
Some would have acted anyway. Uplift modelling predicts the change the intervention causes, and it requires a held-back control group — always keep one.
SELECTION BIAS SHAPES WHAT YOU CAN LEARN
In credit, you only see outcomes for accepted applicants. In fraud, blocked activity never produces labels. In maintenance, successful maintenance removes the evidence.
Let a small random sample through where risk permits, so labels keep arriving.
ADVERSARIAL MODELS DECAY FASTER THAN OTHERS
Because the behaviour changes deliberately. Keep rules alongside models so a new pattern can be blocked immediately while retraining runs.
TEST DOCUMENT AND LANGUAGE SYSTEMS ON LOCAL MATERIAL
Clean scans and benchmark languages give misleading accuracy. Local documents, local varieties and code-switching all need explicit evaluation.
NAIVE TRAINING ON CLICKS TEACHES A MODEL TO REPRODUCE THE EXISTING RANKING
Position bias must be corrected or randomised away.