Producing speech.
WHAT THE TASK IS
Generating natural-sounding speech from text.
WHAT EARLIER APPROACHES DID
Concatenated recorded fragments, or generated speech parametrically.
WHAT THEY SOUNDED LIKE
Recognisably synthetic.
WHAT NEURAL APPROACHES CHANGED
Speech close to indistinguishable from recording.
WHAT THE COMPONENTS TYPICALLY ARE
A model producing an intermediate acoustic representation A vocoder producing the waveform
WHAT VOICE CLONING DOES
Reproduces a specific voice from a sample.
HOW LITTLE AUDIO IT NEEDS NOW
Very little, in some systems.
WHY THAT IS A SERIOUS CONCERN
Voice has been used as an authentication factor and as evidence of identity.
WHAT THAT MEANS PRACTICALLY
Voice alone should not authorise anything consequential.
WHAT ELSE IT ENABLES
Impersonation in fraud, which is already occurring.
WHAT TO ADVISE
Verification through a separate channel for anything involving money.
WHAT LEGITIMATE USES EXIST
Accessibility Audio content production Interfaces where reading is impractical Preserving the voice of people losing speech
WHAT TO ESTABLISH BEFORE CLONING ANY VOICE
Consent, and what obligations apply.