Handling more than text.
THE DEFINITION
A model working with several kinds of input or output: text, images, audio, video.
WHAT IT ENABLES
Describing an image Answering questions about a photograph or chart Extracting text and data from documents Generating images from descriptions Transcribing speech
PRACTICAL USES
Reading a receipt and extracting figures Explaining a diagram Diagnosing a problem from a screenshot Converting a handwritten note to text
WHAT IT DOES WELL
Describing what is visibly present, extracting text, answering questions about clear images.
WHAT IT DOES BADLY
Precise measurement or counting Poor quality images Anything requiring specialist visual judgement
THE CAUTION
Do not rely on it for readings where an error matters.
Verify extracted data against the source.
RELATED TERMS
Computer vision, optical character recognition.