Extracting structure from documents.
WHAT THE STAGES ARE
Classification: what kind of document
Text recognition, for images and scans Field extraction Validation Routing for review where confidence is low
WHAT MAKES IT DIFFICULT
Layout variation between issuers Poor scan and photograph quality Handwriting Fields expressed differently
WHAT TO TEST WITH
Documents photographed by actual users, on their actual devices.
WHY THAT MATTERS SO MUCH
Clean scans give misleading accuracy figures.
WHAT LOCAL DOCUMENTS ADD
Formats and issuers absent from general training data.
WHAT THAT REQUIRES
Evaluation on local documents, and fine-tuning where necessary.
WHAT VALIDATION SHOULD CHECK
Format and checksum rules Internal consistency between fields Plausibility of values
WHY VALIDATION MATTERS MOST
It converts an uncertain extraction into a confident one or a flagged one.
WHAT TO ROUTE TO A PERSON
Anything below a confidence threshold, or failing validation.
WHAT TO MEASURE
Proportion processed without review Error rate among those accepted automatically
WHY THAT SECOND FIGURE
It is the one that determines whether automation is safe.