Working within platform limits.
WHAT QUOTAS ARE
Limits on how much of a service you may consume.
WHAT TYPES EXIST
Rate limits, per minute or per second Allocation limits, on total resources Daily limits, on some APIs
WHY THEY EXIST
To protect the service and to limit runaway consumption.
WHAT HAPPENS WHEN EXCEEDED
Requests are rejected with a specific error.
WHAT TO DO ABOUT THAT ERROR
Retry with increasing delay, and randomisation.
WHY RANDOMISATION
So many clients do not retry simultaneously and sustain the problem.
WHAT TO NEVER DO
Retry immediately, repeatedly.
WHAT TO REQUEST
A quota increase, where the limit is genuinely too low.
WHAT TO PROVIDE WHEN REQUESTING
Justification and expected volume.
WHAT TO DESIGN FOR
Batching, where the API supports it Caching, so the same request is not repeated Backing off gracefully rather than failing
WHAT TO MONITOR
Consumption against limits, with alerting before the limit is reached.
WHY BEFORE
Discovering a limit by hitting it means an outage.
WHAT TO SET DELIBERATELY
Lower quotas than the maximum, as a cost control.