What happened?
An org-wide cache clear for a high-volume org caused a cache stampede / thundering herd into Postgres. A large amount of traffic that would normally be served from cache missed at the same time, causing many concurrent DB hydrations. PgBouncer client connections spiked from a normal ~700-850 range to ~7,600, and pool wait time peaked at ~113s. This caused API requests to hit application timeouts, returning fail-open 202s for entitlement paths and 5xxs for DB-backed paths.
Total non-403 requests: 2,695,878
Fail-open 202 responses: 148,527 5.51%
5xx responses: 145,823 5.41%
The cache clear invalidated too much hot data at once for a high-volume org. The subsequent cold-cache traffic produced a thundering herd of DB reads and cache rehydrations. The database connection pool saturated, which led to connection waits and query timeouts.
Resolution:
Traffic was partially blocked/rate-limited and the cache/DB pressure drained. By 18:04 UTC, PgBouncer max wait had returned to 0s, and recent checks showed no ongoing fail-open 202s or 5xxs on the affected check/track paths.
Follow-Ups:
- Spinning up a dedicated PG bouncer to separate the CPU load from the primary DB
- Set better alerts on DB CPU to more proactively catch issues
- Add concurrency gating to DB queries
As always please let us know if you have any questions.