Running training practically.
WHAT DETERMINES COST
Hardware type and hours Data transfer Storage of datasets and checkpoints Failed and repeated runs
WHAT REDUCES IT
Starting from a pre-trained model Training on a sample first Early stopping Mixed precision Interruptible instances, where restarts are handled
WHAT INTERRUPTIBLE CAPACITY REQUIRES
Checkpointing, and resumption.
WHAT CHECKPOINTING SHOULD SAVE
Model parameters Optimiser state The position in the data The epoch and step
WHY OPTIMISER STATE SPECIFICALLY
Resuming without it degrades training.
WHAT TO DO BEFORE ANY EXPENSIVE RUN
Verify the pipeline end to end on a small sample.
WHY
Discovering a bug after hours of training is an avoidable expense.
WHAT DISTRIBUTED TRAINING PROVIDES
Using several devices or machines.
WHAT LIMITS IT
Communication between devices.
WHAT TO ESTABLISH FIRST
Whether one device suffices.
WHAT MATTERS LOCALLY
Foreign-currency billing, and limited access to capable hardware.
WHAT THAT ARGUES FOR
Smaller models, transfer learning, and careful budgeting.
WHAT TO SET
Spend limits, before starting.