Knowing whether they work.
WHAT OFFLINE EVALUATION DOES
Measures performance against historical data.
WHAT IT CANNOT CAPTURE
How users respond to recommendations they never saw.
WHY THAT MATTERS
Historical data reflects what the previous system showed, not what users would have liked.
WHAT THAT BIAS IS CALLED
Feedback loop bias, and it is pervasive.
WHAT ONLINE EVALUATION DOES
Tests with real users, comparing variants.
WHAT TO MEASURE
Immediate engagement
Downstream outcomes: purchases, retention, satisfaction
Diversity of what is shown Fairness across user groups
WHY DOWNSTREAM OUTCOMES MATTER MOST
Clicks are easy to increase and frequently mean nothing.
WHAT TO BE CAREFUL WITH
Running experiments too briefly Measuring only what is easy Ignoring effects on groups that are small in the data
WHAT NOVELTY EFFECT IS
Improved engagement simply because something changed.
WHAT TO DO ABOUT IT
Run long enough for it to subside.
WHAT TO MONITOR CONTINUOUSLY
Whether performance degrades as behaviour changes.
WHY
Models trained on past behaviour become stale.
WHAT TO PLAN
Retraining, and detection of drift.