Liveness, readiness and startup.
WHAT A LIVENESS PROBE DETERMINES
Whether the container should be restarted.
WHAT A READINESS PROBE DETERMINES
Whether it should receive traffic.
WHAT A STARTUP PROBE DETERMINES
Whether it has finished starting, suspending the others until it has.
WHY THE STARTUP PROBE EXISTS
Slow-starting applications otherwise fail liveness during startup and restart forever.
WHAT A LIVENESS PROBE SHOULD CHECK
That the process is functioning.
WHAT IT SHOULD NOT CHECK
Dependencies.
WHY
A database outage otherwise restarts every application pod, making recovery harder.
WHAT A READINESS PROBE MAY CHECK
Dependencies, since withholding traffic is the correct response.
WHAT THE PARAMETERS CONTROL
Delay before starting How often How long to wait How many failures count
WHAT THE COMMONEST MISTAKE IS
A liveness probe that is too aggressive.
WHAT THAT PRODUCES
Restarts under load, when the application is merely slow.
WHAT THAT LOOKS LIKE
Pods restarting during traffic peaks.
WHAT TO DO
Loosen the thresholds, and check the application's actual behaviour.
WHAT TO IMPLEMENT IN THE APPLICATION
A lightweight endpoint for liveness, and a fuller one for readiness.
WHAT TO KEEP THEM
Cheap, since they run constantly.