Where it performs and where it fails.
WHAT IT DOES WELL
Detecting specific findings it was trained on, at scale, without fatigue Flagging studies for priority review Measuring consistently Screening large volumes
WHERE IT HAS SHOWN GENUINE VALUE
Screening programmes, where volume is high and the target finding is well defined.
WHAT IT DOES BADLY
Findings it was not trained on Images unlike its training data Unusual presentations Anything requiring clinical context
THE FAILURE MODE
Confident misclassification. The system does not know when it is outside familiar territory.
A normal report on an abnormal study is the dangerous outcome.
WHAT THIS MEANS PRACTICALLY
A radiologist reviews. Always.
An AI flag prompts attention. An AI clearance does not remove the need to look.
THE AUTOMATION BIAS RISK
Readers shown an AI assessment are influenced by it, including when it is wrong.
Some protocols require independent reading before revealing the AI output.
WHAT TO ASK OF ANY SYSTEM
What was it trained on, what does it miss, and how was that measured.