Running containers at scale.
WHAT ORCHESTRATION PROVIDES
Scheduling containers across machines Restarting failed ones Scaling by count Service discovery Rolling updates and rollback Configuration and secret distribution
WHAT THE DOMINANT PLATFORM IS
A widely adopted system that has become the standard.
WHAT IT COSTS
Substantial complexity Operational expertise Ongoing maintenance of the platform itself
WHEN IT IS WARRANTED
Many services Several teams deploying independently Genuine need for automated scaling and self-healing
WHEN IT IS NOT
A handful of services One team Predictable load
WHAT MOST ORGANISATIONS AT THAT SCALE NEED
Simpler deployment onto a few machines.
WHAT TO CONSIDER INSTEAD
Managed platforms, which remove the burden of operating the control plane.
WHAT TO ALWAYS CONFIGURE
Resource requests and limits Health checks that exercise dependencies Rolling update strategy
WHAT HAPPENS WITHOUT LIMITS
One workload consumes a node and affects everything on it.
WHAT TO MONITOR
Node capacity, pending workloads, and restart counts.
WHAT RESTART LOOPS INDICATE
A failing workload that orchestration is masking.