Observing many short-lived things.
WHAT DIFFERS FROM MONITORING SERVERS
Instances appear and disappear Names change Aggregate behaviour matters more than any instance
WHAT TO MONITOR
Resource use per container Restarts Pods not ready Pods unable to be placed Request rates, errors and latency
WHY RESTARTS MATTER MOST
They indicate crashes that would otherwise be invisible.
WHAT TO ALERT ON
Repeated restarts Pods pending for a long time Memory terminations Nodes under pressure
WHAT A PENDING POD USUALLY MEANS
No node has room, or storage cannot be attached.
WHAT MEMORY TERMINATIONS INDICATE
Limits too low, or a leak.
WHAT TO COLLECT
Metrics with labels identifying the workload, not the instance.
WHY
Instances are ephemeral; the workload persists.
WHAT TO CENTRALISE
Logs, because containers disappear with theirs.
WHAT TO ADD TO EVERY LOG LINE
The workload, namespace and a correlation identifier.
WHAT TRACING PROVIDES
The path of one request across services.
WHY THAT MATTERS HERE
A single request may touch several containers, and logs alone do not connect them.
WHAT TO ESTABLISH
A baseline of normal.