Precomputing for speed and cost.
WHAT THEY ARE
Tables holding pre-aggregated results at a coarser grain.
WHY THEY EXIST
Aggregating large volumes repeatedly is slow and expensive.
WHAT THEY PROVIDE
Fast dashboards Predictable cost Lower load on the warehouse
WHAT THEY COST
Additional pipelines to maintain Potential divergence from the detail Less flexibility
WHAT TO AGGREGATE
Combinations queried frequently.
HOW TO KNOW WHICH
Query logs, which record what people actually ask.
WHAT GRAIN TO CHOOSE
Coarse enough to be small, fine enough to serve the questions.
WHAT TO BE CAREFUL WITH
Measures that cannot be re-aggregated, such as distinct counts and ratios.
WHY DISTINCT COUNTS SPECIFICALLY
Summing distinct counts across groups double-counts entities appearing in several.
WHAT TO DO INSTEAD
Store the underlying sets, or recompute at the needed grain.
WHAT TO TEST
That the aggregate matches the detail, exactly.
WHEN TO TEST IT
On every build.
WHY
Divergence between a summary and the detail destroys trust in both.
WHAT TO REBUILD PERIODICALLY
Aggregates built incrementally.