Testing whether a change helps.
WHY OFFLINE IMPROVEMENT IS INSUFFICIENT
A better metric does not guarantee a better outcome.
WHAT CAUSES THE GAP
The metric being a proxy Users responding to the change itself Effects on parts of the system nobody modelled
WHAT AN ONLINE EXPERIMENT DOES
Compares outcomes between users receiving the new model and those receiving the current one.
WHAT TO RANDOMISE
Assignment, at the right unit.
WHAT THE RIGHT UNIT IS
Usually the user, not the request.
WHY
A user seeing both versions produces an inconsistent experience and contaminated results.
WHAT TO DECIDE BEFORE STARTING
The primary metric The minimum effect worth detecting How long to run What guardrail metrics must not worsen
WHY GUARDRAILS
A change improving one metric while damaging another is not an improvement.
WHAT TO RUN LONG ENOUGH FOR
Novelty effects to subside, and weekly cycles to complete.
WHAT TO AVOID
Stopping when the result looks good Testing many variants without accounting for it Concluding from a difference within the noise
WHAT TO DO WITH A NEGATIVE RESULT
Accept it, and record why the offline signal misled.