Knowing what is happening.
WHAT TO MONITOR PER DEVICE
Last contact time Battery or power status Signal quality Firmware version Configuration version Error counts Data volume
WHY LAST CONTACT MATTERS MOST
It is the primary indicator that something is wrong.
WHAT TO MONITOR ACROSS THE FLEET
Proportion online Version distribution Error rates Aggregate data consumption
WHAT VERSION DISTRIBUTION REVEALS
Devices that never updated, which are a standing risk.
WHAT TO ALERT ON
Devices silent beyond their expected interval Battery levels approaching depletion Error rates rising Data consumption far above expectation
WHY DATA CONSUMPTION SPECIFICALLY
Unexpected volume indicates a fault or a compromise, and it costs money.
WHAT TO AVOID
Alerting on every individual device in a large fleet.
WHY
It is unmanageable, and real problems are lost.
WHAT TO ALERT ON INSTEAD
Proportions crossing thresholds Sites or groups failing together
WHY GROUPS MATTER
Correlated failure indicates a common cause: connectivity, power, or a bad update.
WHAT TO PROVIDE OPERATORS
A view of the fleet's health at a glance, and drill-down when needed.