Knowledgebase

Monitoring Models in Production Print

  • machinelearningengineering, machine, performance, errors, troubleshooting, affiliate, guide, howto
  • 0

Knowing whether it still works.

WHAT TO MONITOR

Operational health: availability, latency, errors

Input distributions Prediction distributions Performance, where outcomes become known Fallback and referral rates Cost

WHY INPUT DISTRIBUTIONS MATTER

Changes there precede performance degradation, and are visible immediately.

WHAT PREDICTION DISTRIBUTION SHIFTS INDICATE

Either the population changed, or something upstream broke.

WHAT THE HARD PROBLEM IS

Outcomes arriving late, or never.

WHAT THAT MEANS

Actual performance may not be measurable for weeks.

WHAT TO DO ABOUT IT

Monitor proxies: distributions, and business outcomes.

WHAT FEEDBACK DELAY REQUIRES

Accepting that degradation is detected late, and monitoring what you can see now.

WHAT TO COMPARE AGAINST

The training distribution, recorded at training time.

WHY RECORDED THEN

You cannot compare against something you did not save.

WHAT TO ALERT ON

Significant distribution shift Performance below a threshold Error and fallback rates rising Prediction volume far from expected

WHAT TO REVIEW REGULARLY

Sample predictions, by eye.

WHY

Metrics miss failures that are obvious on inspection.


Was this answer helpful?
Back

Are you happy with your experience? Leave us a review on Trustpilot.


Trustpilot