Getting the most from hardware.
WHAT TO DO FIRST
Measure, with a profiler, on representative data.
WHAT TO LOOK FOR
Where time is spent Cache miss rates Branch misprediction rates Instructions retired per cycle
WHAT LOW INSTRUCTIONS PER CYCLE INDICATES
Stalling, usually on memory.
WHAT TO ADDRESS FIRST
Algorithmic complexity, which dominates everything else.
WHY
No amount of micro-optimisation fixes a poor algorithm.
WHAT TO ADDRESS NEXT
Memory access patterns.
WHAT TO ADDRESS LAST
Instruction-level details.
WHAT COMMONLY HELPS
Contiguous data layout Structure of arrays rather than array of structures, for bulk processing Avoiding unnecessary allocation Reducing indirection
WHAT BRANCHLESS CODE PROVIDES
Avoiding mispredictions, where branches are unpredictable.
WHAT IT COSTS
Readability, and it is frequently slower where branches predict well.
WHAT TO NEVER DO
Optimise without measuring before and after.
WHAT TO PRESERVE
Correctness, verified by tests, at every step.
WHAT TO DOCUMENT
Why unusual code exists, so nobody simplifies it back.