Knowing whether it works.
WHAT ACCURACY MEASURES
The proportion of correct predictions.
WHY IT MISLEADS
With imbalanced classes, predicting the majority always gives high accuracy and no value.
WHAT TO USE INSTEAD
Precision: of what was predicted positive, how much was correct
Recall: of what was actually positive, how much was found
Their combination, where both matter
WHAT TO CHOOSE BY
Which error is more costly.
WHAT THAT MEANS PRACTICALLY
Missing a fraudulent transaction and blocking a legitimate one have different costs, and the threshold should reflect that.
WHAT A CONFUSION MATRIX SHOWS
Exactly which errors are being made.
WHY THAT MATTERS
Aggregate metrics conceal that errors concentrate in one class.
WHAT TO EVALUATE
Performance across segments, not only overall.
WHY
A model can perform well overall and badly for a group.
WHAT TO ESTABLISH
A threshold for deployment, decided before training.
WHY BEFORE
Otherwise the threshold moves to whatever the model achieved.
WHAT TO COMPARE AGAINST
The baseline, and the current process.