The machines underneath.
WHAT A NODE RUNS
The workloads, plus agents managing them.
WHAT TO MONITOR
Readiness
Resource pressure: memory, disk, processes
Disk filling with images and logs
WHY DISK PRESSURE MATTERS
The node evicts pods, and they move elsewhere, spreading the problem.
WHAT TO CONFIGURE
Image cleanup thresholds.
WHAT DRAINING A NODE DOES
Moves workloads off it, respecting disruption budgets.
WHEN TO DRAIN
Before maintenance or replacement.
WHAT CORDONING DOES
Stops new workloads being placed, without moving existing ones.
WHAT TO PREFER FOR UPGRADES
Replacing nodes rather than upgrading in place.
WHY
It is repeatable, and it verifies workloads reschedule.
WHAT THAT REQUIRES
Workloads that tolerate being moved.
WHAT DOES NOT TOLERATE IT WELL
Anything with local storage.
WHAT TO CHECK BEFORE DRAINING
What is running that cannot move.
WHAT NODE POOLS PROVIDE
Groups of nodes with different characteristics.
WHERE THAT HELPS
Separating system components from applications Different sizes for different workloads
WHAT TO AVOID
A single enormous node.
WHY
Its failure takes everything, and scheduling becomes coarse.
WHAT TO PREFER
Several smaller nodes.