Detecting prohibited content.
WHAT THE TASKS ARE
Classifying content against policy categories Prioritising for human review Detecting coordinated behaviour Matching known prohibited material
WHAT MAKES IT DIFFICULT
Context determines meaning Policies are nuanced and change Language, slang and code evolve deliberately Errors harm in both directions
WHAT BOTH DIRECTIONS MEANS
Missing harmful content, and removing legitimate content.
WHAT TO SET THRESHOLDS FROM
The relative harm of each, per category.
WHY PER CATEGORY
The correct balance for spam differs entirely from that for safety-critical categories.
WHAT TO AUTOMATE FULLY
Only high-confidence matches against known material.
WHAT TO ROUTE FOR REVIEW
Everything else above a threshold.
WHAT REVIEWERS NEED
Context, policy guidance, and manageable volume.
WHAT TO PROVIDE USERS
Notice of action taken, and an appeal route.
WHAT TO MONITOR
Appeal overturn rates, by category.
WHAT HIGH RATES INDICATE
A threshold set wrongly, or a policy applied inconsistently.
WHAT TO BE AWARE OF
Performance differing by language, which disadvantages some users systematically.
WHAT TO TEST
Local languages and varieties, specifically.