Diagnosing slowness.
WHAT TO ESTABLISH FIRST
Which operation is slow, specifically Whether it is always slow or sometimes Whether it worsens under load
WHAT TO MEASURE
Response time distribution, not the average.
WHY
Averages conceal the experience of the worst-affected requests.
WHAT TO LOOK AT
The slow end of the distribution.
WHAT TO PROFILE
Where time is actually spent within a request.
WHAT THE USUAL CAUSES ARE
Database queries, by a wide margin Queries issued in a loop Missing indexes External calls without timeouts Work that should be in the background Blocking operations in an event-driven runtime
WHAT TO CHECK FIRST
Query count and query time per request.
WHY
It is the answer more often than anything else.
WHAT TO AVOID
Optimising without measuring Assuming the cause
WHAT TO ESTABLISH
A baseline, so improvement can be demonstrated.
WHAT TO MONITOR CONTINUOUSLY
Response times, so regressions are noticed rather than reported.