Knowing they still work.
WHAT TO MONITOR
Success and failure rates, per integration Latency Queue depth and age Reconciliation discrepancies Last successful run, for scheduled work
WHY LAST SUCCESSFUL RUN MATTERS MOST
An integration that silently stopped is the commonest failure, and the hardest to notice.
WHAT TO ALERT ON
No successful run within an expected period Error rate above normal A queue growing rather than draining Any reconciliation discrepancy
WHAT NOT TO ALERT ON
Individual failures that retry successfully.
WHY
They are normal, and alerting on them trains people to ignore alerts.
WHAT TO BUILD
A dashboard showing every integration's state at a glance.
WHAT IT SHOULD SHOW
Healthy, degraded or failed When it last succeeded How much is pending
WHY THAT VIEW MATTERS
It answers the first question during any incident.
WHAT TO RECORD PER OPERATION
What was attempted, when, and the outcome.
WHAT TO PROVIDE
A way to find a specific record's integration history.
WHY
It is what support needs when a customer says something did not arrive.
WHAT TO TEST PERIODICALLY
That the alerting works, by inducing a failure.
WHAT TO REVIEW
Integrations nobody has looked at in months.