The summary.
OVERCOMMIT PROCESSORS MODERATELY AND MEMORY VERY CAUTIOUSLY
Exhausting memory causes swapping, and performance collapses for everything on the host.
Monitor contention, not merely allocation. And storage performance is frequently the real constraint.
SNAPSHOTS ARE NOT BACKUPS
They depend on the original storage and they grow. Long-lived ones cause storage exhaustion and severe degradation — take one before a change and remove it after.
CONTAINERS SHARE THE HOST KERNEL
Which makes them insufficient isolation between untrusted customers. Use virtual machines, or containers inside per-customer machines.
CONSOLE AND OUT-OF-BAND ACCESS ARE WHAT MAKE RECOVERY POSSIBLE
Without them, a customer's firewall mistake becomes a support ticket every time, and an unbootable machine needs someone physically present.
That management interface must never be reachable from the internet.
ORCHESTRATION IS WARRANTED BY SCALE, NOT BY FASHION
At a handful of services with one team, simpler deployment is better. Always set resource limits, or one workload consumes a node.
WHAT MATTERS MORE THAN THE PLATFORM CHOSEN
Backups, monitoring, and recovery that has actually been tested.