Understanding what is running.
WHAT MONITORING PROVIDES
Metrics for every resource Logs, collected centrally Alerts on conditions Application performance insight
WHAT TO COLLECT
Resource metrics Application logs and traces Activity logs, recording who did what Security signals
WHAT TO ALERT ON
Availability failures Error rates rising Resource limits approaching Unexpected cost increases Administrative changes
WHY THAT LAST ONE
Unexpected configuration changes warrant immediate attention.
WHAT TO AVOID
Alerting on everything, which produces noise nobody reads.
WHAT TO ESTABLISH
What normal looks like, before setting thresholds.
WHAT APPLICATION INSIGHT PROVIDES
Request rates, response times, failures, and dependency performance.
WHAT TO INSTRUMENT
Your own applications, with correlation identifiers.
WHY
So a request can be followed across services.
WHAT TO RETAIN
Logs long enough to investigate incidents.
WHAT TO BE CAREFUL WITH
Log volume, which is billed and can become substantial.
WHAT TO REVIEW
What you are collecting, and whether it is used.