Measuring categorical predictions.
WHAT THE CONFUSION MATRIX SHOWS
Correct positives, correct negatives, false positives and false negatives.
WHY IT IS THE STARTING POINT
Every classification metric derives from it.
WHAT ACCURACY MEASURES
The proportion correct overall.
WHEN IT MISLEADS
Whenever classes are imbalanced.
WHAT PRECISION MEASURES
Of those predicted positive, how many were.
WHAT RECALL MEASURES
Of those actually positive, how many were found.
WHAT THE TRADE-OFF IS
Raising one usually lowers the other.
WHEN PRECISION MATTERS MORE
When acting on a false positive is costly: blocking a legitimate customer, alerting a person.
WHEN RECALL MATTERS MORE
When missing a positive is costly: disease screening, fraud, safety.
WHAT A COMBINED MEASURE PROVIDES
A single figure balancing them, weighted as you choose.
WHAT THE RECEIVER CURVE SHOWS
The trade-off across all thresholds.
WHY IT MISLEADS ON IMBALANCED DATA
It looks good even when precision is poor.
WHAT TO USE INSTEAD
The precision-recall curve.
WHAT TO DECIDE EXPLICITLY
The threshold, from the relative costs.