Producing the answers to learn from.
WHERE LABELS COME FROM
Recorded outcomes, which are best Human annotation Weak supervision from rules or heuristics Existing systems' decisions
WHAT RECORDED OUTCOMES PROVIDE
Ground truth, at no annotation cost.
WHAT TO PREFER ALWAYS
Those, where they exist.
WHAT HUMAN ANNOTATION REQUIRES
Clear guidelines Examples of difficult cases Consistency checking between annotators A route to escalate ambiguity
WHAT INTER-ANNOTATOR AGREEMENT MEASURES
How consistently different people apply the guidelines.
WHY IT MATTERS
Low agreement means the task is ill-defined, and a model cannot exceed the consistency of its labels.
WHAT TO DO ABOUT LOW AGREEMENT
Improve the guidelines, before annotating more.
WHAT TO NEVER DO
Label a large set before checking agreement on a small one.
WHAT ACTIVE LEARNING IS
Selecting the most informative examples to label next.
WHAT IT SAVES
Annotation effort, substantially.
WHAT TO LABEL FIRST
A small set, to establish the guidelines work.
WHAT TO RE-CHECK
Labels for examples the model finds most confusing.
WHY
They are disproportionately likely to be wrong.