Knowing what is happening.
WHAT TO LOG PER REQUEST
The endpoint and method The caller The status code The duration A correlation identifier
WHY THE CORRELATION IDENTIFIER
It links a request across services and to the caller's report.
WHAT TO NEVER LOG
Credentials and tokens Payment details Personal data beyond what is necessary
WHY
Logs are widely readable and long retained.
WHAT TO MONITOR
Request rate Error rate, split by class Latency at percentiles Rate limit rejections
WHY PERCENTILES RATHER THAN AVERAGES
Averages hide the slow requests callers notice.
WHAT TO ALERT ON
Error rate rising Latency rising A caller suddenly dominating traffic Authentication failures increasing
WHAT THE LAST ONE USUALLY MEANS
A broken integration, or an attack.
WHAT TO TRACK PER CALLER
Volume, errors and latency.
WHY PER CALLER
It identifies whose integration is broken, and who is causing load.
WHAT TO PROVIDE CALLERS
Visibility of their own usage and errors.
WHY
It removes a large category of support requests.
WHAT TO RETAIN
Enough history to investigate a report from last week.