Drawing conclusions from biological data.
WHY STATISTICS DOMINATES
Biological data is noisy, variable between individuals, and expensive to collect.
WHAT MULTIPLE TESTING IS
Performing many tests simultaneously, which produces apparent findings by chance.
WHY IT MATTERS SO MUCH HERE
Testing tens of thousands of genes guarantees false positives at conventional thresholds.
WHAT TO APPLY
Correction for multiple testing, always.
WHAT THE COMMON APPROACHES ARE
Controlling the proportion of false discoveries among findings More conservative corrections controlling any false positive
WHAT BATCH EFFECTS ARE
Systematic differences arising from when or where samples were processed.
WHY THEY ARE DANGEROUS
They can perfectly mimic the biological effect you are looking for.
WHAT PREVENTS THEM
Randomising sample processing across batches, at collection time.
WHAT TO DO WHEN THEY EXIST
Model them explicitly, and report that you did.
WHAT POWER MEANS
The ability to detect an effect of a given size.
WHY IT MATTERS
Underpowered studies produce findings that do not replicate.
WHAT TO ESTABLISH BEFORE COLLECTING
How many samples the question requires.
WHAT TO REPORT
Everything tested, not only what was significant.