Verifying correctness.
WHY IT DIFFERS FROM SOFTWARE TESTING
The code may be correct while the data is wrong, and both must be checked.
WHAT TO TEST ABOUT THE CODE
Logic, against known inputs and expected outputs.
WHAT TO TEST ABOUT THE DATA
Uniqueness of keys Absence of nulls where required Referential integrity between models Accepted values in categorical columns Ranges of numerical columns Row counts within expected bounds
WHAT TO ASSERT ON EVERY MODEL
Key uniqueness and non-nullity.
WHAT RECONCILIATION TESTS DO
Compare a total against an independently known figure.
WHY THEY ARE THE MOST VALUABLE
They catch errors every structural test misses.
WHAT TO COMPARE AGAINST
The source system's own reported totals, where available.
WHAT TO DO WHEN THEY DISAGREE
Investigate before publishing, not after someone notices.
WHAT UNIT TESTS LOOK LIKE HERE
Small fixed inputs, run through the transformation, compared to expected output.
WHAT THEY CATCH
Logic errors, particularly in edge cases.
WHAT TO TEST EXPLICITLY
Nulls Empty inputs Duplicate inputs Boundary dates
WHAT TO RUN TESTS ON
Every change, before deployment.
WHAT TO DO WHEN A TEST FAILS IN PRODUCTION
Stop the pipeline, rather than publishing.