Understanding running systems.
WHAT THE OPERATIONS SUITE PROVIDES
Metrics for every service Logs collected centrally Alerting on conditions Distributed tracing Error reporting Profiling
WHAT TO COLLECT
Request rates, errors and latency Resource utilisation Application logs, structured Audit logs, recording who did what
WHY STRUCTURED LOGS
They can be searched and aggregated by field, rather than by text matching.
WHAT TO ALERT ON
Availability failures Error rate rising Latency exceeding a threshold Quota approaching Unexpected cost increases Administrative changes
WHAT TO AVOID
Alerting on everything, which produces noise nobody reads.
WHAT TO ESTABLISH FIRST
What normal looks like.
WHAT TRACING PROVIDES
Where time went, across services.
WHAT TO PROPAGATE
A correlation identifier through every call.
WHAT TO BE CAREFUL WITH
Log volume, which is billed and grows quickly.
WHAT TO CONFIGURE
Retention, and exclusion of high-volume low-value entries.
WHAT TO NEVER LOG
Credentials, tokens or unnecessary personal data.