Making computers interpret images.
WHAT IT DOES
- Classification: what is in this image
- Detection: where are the objects
- Segmentation: which pixels belong to which object
- Recognition: identifying specific individuals or items
- Reading: extracting text
WHERE IT IS USED
Medical imaging support Quality inspection in manufacturing Document scanning and data extraction Security and access control Agriculture, assessing crops Retail, tracking stock Autonomous vehicles
WHAT IT DOES WELL
Consistent, high-volume visual tasks where examples are plentiful.
Frequently better than humans at narrow tasks, and tireless.
WHAT IT DOES BADLY
Unusual situations unlike its training data Poor quality images Tasks requiring context beyond the image
THE FAILURE MODE
Confident misclassification. The system does not know when it is outside familiar territory.
WHERE CARE IS ESSENTIAL
Medical, safety and security applications, where an error has serious consequences.
Human review of consequential decisions is not optional in those contexts.
FOR PRACTICAL BUSINESS USE
Document data extraction is the most commonly useful application.