Knowing whether it still works.
WHAT TO MONITOR
Operational health: availability, latency, errors
Input distributions Prediction distributions Performance, where outcomes become known Fallback and referral rates Cost
WHY INPUT DISTRIBUTIONS MATTER
Changes there precede performance degradation, and are visible immediately.
WHAT PREDICTION DISTRIBUTION SHIFTS INDICATE
Either the population changed, or something upstream broke.
WHAT THE HARD PROBLEM IS
Outcomes arriving late, or never.
WHAT THAT MEANS
Actual performance may not be measurable for weeks.
WHAT TO DO ABOUT IT
Monitor proxies: distributions, and business outcomes.
WHAT FEEDBACK DELAY REQUIRES
Accepting that degradation is detected late, and monitoring what you can see now.
WHAT TO COMPARE AGAINST
The training distribution, recorded at training time.
WHY RECORDED THEN
You cannot compare against something you did not save.
WHAT TO ALERT ON
Significant distribution shift Performance below a threshold Error and fallback rates rising Prediction volume far from expected
WHAT TO REVIEW REGULARLY
Sample predictions, by eye.
WHY
Metrics miss failures that are obvious on inspection.