Understanding a distributed system.
WHAT THE THREE KINDS ARE
- Metrics: numbers over time
- Logs: records of events
- Traces: the path of one request
WHY ALL THREE
Each answers a question the others cannot.
WHAT METRICS ANSWER
Whether something is wrong, and how widely.
WHAT LOGS ANSWER
What happened in one place.
WHAT TRACES ANSWER
Where the time went across services.
WHAT TO INSTRUMENT
Request rate, error rate and duration, per service.
WHY THOSE THREE
They describe almost every problem users experience.
WHAT TO ADD
Saturation: how full the resources are.
WHAT LABELS TO USE
Ones identifying the workload, not the instance.
WHY
Instances are ephemeral, and per-instance labels multiply endlessly.
WHAT THAT PROBLEM IS CALLED
High cardinality.
WHAT IT CAUSES
Monitoring systems consuming enormous resources.
WHAT TO NEVER USE AS A LABEL
Anything unbounded: user identifiers, request identifiers, paths with values in them.
WHAT TO PROPAGATE FOR TRACING
A correlation identifier, through every call.
WHAT TO SAMPLE
Traces, since keeping all of them is expensive.
WHAT TO KEEP FULLY
Traces of requests that errored.