Numbers over time.
WHAT A METRIC IS
A value recorded repeatedly, showing change over time.
WHAT TO RECORD
Requests per period Error rate Response time Resource use Queue depth
Business quantities: orders, sign-ups, payments
WHY TIME SERIES MATTER
A single reading means little. A trend means a great deal.
WHAT TO LOOK AT FOR RESPONSE TIMES
Not the average.
WHY
An average conceals the slow requests, which are the ones people notice.
WHAT TO USE INSTEAD
Percentiles, showing what the slowest portion of requests experience.
WHAT TO ESTABLISH
What normal looks like for each metric.
WHY
Without it, you cannot tell whether a value is a problem.
WHAT TO ALERT ON
Departure from normal, sustained.
WHAT TO CORRELATE
Metrics with deployments.
WHY
It answers whether a release caused a change.
WHAT TO RETAIN
Enough history to see seasonal patterns.
WHAT TO REVIEW
Trends monthly, not only during incidents.