Architectures for ordered data.
WHAT RECURRENT NETWORKS DID
Processed sequences step by step, carrying a hidden state.
WHAT LIMITED THEM
Sequential processing, preventing parallelism Difficulty carrying information across long distances
WHAT ATTENTION DOES
Lets every position attend directly to every other, weighted by relevance.
WHY THAT SOLVED BOTH PROBLEMS
Distance no longer matters, and the computation parallelises.
WHAT SELF-ATTENTION MEANS
Positions within one sequence attending to each other.
WHAT THE TRANSFORMER ARCHITECTURE IS
Layers of self-attention and feed-forward computation, with normalisation and skip connections.
WHY IT DOMINATES
It scales, parallelises, and transfers across domains.
WHAT POSITIONAL INFORMATION IS NEEDED FOR
Attention has no inherent notion of order, so position must be supplied.
WHAT THE COST IS
Attention scaling with the square of sequence length.
WHAT THAT LIMITS
Context length, and why extending it is a research focus.
WHAT VARIANTS ADDRESS IT
Approximations reducing that cost.
WHERE TRANSFORMERS ARE NOW USED
Language, vision, audio, time series, and protein structure.
WHAT TO USE FOR MOST SEQUENCE TASKS
A pre-trained transformer, adapted.