The summary.
SAMPLE ABOVE TWICE THE HIGHEST FREQUENCY, AND FILTER BEFORE SAMPLING
Aliasing cannot be corrected afterwards.
Convolution underlies both classical filtering and convolutional neural networks — it is the same operation.
CLASSICAL IMAGE METHODS REMAIN BEST FOR GEOMETRIC AND MEASUREMENT TASKS
They need no training data. Learned approaches win on classification and detection.
SPEECH SYSTEMS PERFORM WORSE ON ACCENTS UNLIKE THEIR TRAINING DATA
Test with recordings from your actual users in their actual conditions — published accuracy rarely reflects local speech.
Self-supervised learning is what makes systems for under-resourced languages feasible.
VOICE CLONING NEEDS VERY LITTLE AUDIO NOW
Which means voice alone must not authorise anything consequential, and impersonation fraud is already occurring.
A CONVERSATIONAL SYSTEM THAT CONVERSES BUT COMPLETES NOTHING IS WORSE THAN A FORM
Design the tasks first, measure task completion rather than conversation length, and hand over to a person promptly on repeated failure.
Read transcripts — they reveal what users actually want, which differs from what was designed for.