Architectures for images and grids.
WHAT A CONVOLUTION DOES
Applies a small learned filter across the input, producing a feature map.
WHY THAT SUITS IMAGES
Patterns are local, and the same pattern may appear anywhere.
WHAT PARAMETER SHARING PROVIDES
Far fewer parameters than connecting everything.
WHAT POOLING DOES
Reduces spatial size, providing tolerance to small shifts.
WHAT DEPTH PRODUCES
Early layers detecting edges and textures, later layers detecting parts and objects.
WHAT THE STANDARD PATTERN IS
Convolution, normalisation, activation, repeated, with periodic downsampling.
WHAT ARCHITECTURES TO USE
Established ones, pre-trained, rather than designed from scratch.
WHY
They are heavily tuned and available with pre-trained weights.
WHAT AUGMENTATION PROVIDES
Artificially expanded training data: flips, crops, rotations, colour changes.
WHY IT MATTERS SO MUCH
It is the most effective regulariser available for image tasks.
WHAT TO BE CAREFUL WITH
Augmentations that change the label Augmentations unlike real variation
WHAT TO ALWAYS DO
Start from a pre-trained model.
WHAT THAT ACHIEVES
Useful results from hundreds of images rather than millions.