Sizing the infrastructure.
WHAT TO MEASURE FIRST
Actual resource use per workload, under realistic load.
WHY FIRST
Everything else depends on it, and guesses are consistently wrong.
WHAT TO PLAN FOR
Peak, not average Room for a node to fail Room for rolling updates
WHY ROLLING UPDATES NEED HEADROOM
Extra instances exist temporarily during the change.
HOW MUCH HEADROOM
Enough that losing the largest node does not exhaust capacity.
WHAT NODE SIZE TO CHOOSE
Large enough for your biggest workload, with room Small enough that losing one is tolerable
WHAT TOO-LARGE NODES CAUSE
Coarse scheduling, and a big loss when one fails.
WHAT TOO-SMALL NODES CAUSE
Workloads that cannot be placed Overhead multiplied across many nodes
WHAT TO ACCOUNT FOR
System components on every node, which consume a meaningful share.
WHAT PEOPLE FORGET
That usable capacity is well below the advertised total.
WHAT TO MONITOR
Requested against available Actual against requested
WHAT A LARGE GAP BETWEEN THOSE MEANS
Requests are set too high, and you are paying for unused capacity.
WHAT TO REVIEW
Those figures, quarterly.
WHAT TO PLAN BEFORE GROWTH
When another node is needed.