Growing with the fleet.
WHAT GROWS
Message volume, linearly with devices Storage, continuously Query load, with users Alert evaluation, with rules and devices
WHAT BREAKS FIRST
Storage cost, usually Ingestion, at bursts Queries over accumulated history
WHAT TO PLAN EARLY
Retention and downsampling.
WHY EARLY
Retrofitting them across accumulated data is painful.
WHAT TO PARTITION BY
Time, primarily Customer or device, where isolation matters
WHAT CACHING SUITS
Current state, queried constantly by dashboards.
WHY
It is the same query repeated, and it need not reach storage.
WHAT TO PRECOMPUTE
Aggregates displayed frequently.
WHAT TO AVOID
Dashboards querying raw history on every refresh.
WHY
It is the commonest cause of unexpected cost and slowness.
WHAT TO MEASURE PER CUSTOMER
Messages, storage and query load.
WHY
It reveals which customers cost disproportionately.
WHAT TO TEST
The reconnection burst, at projected fleet size.
WHAT TO PLAN FOR
A fleet several times larger than current.
WHY
Growth in device count is frequently sudden, following a single large customer.