The issue has now been fully resolved.
--- POST MORTEM ---
Autumn had three periods of partial outage overnight on July 23rd (UTC): 12:00–12:27AM, 4:10–4:20AM, and 9:00–9:26AM. During these windows, requests to our API failed with 503s - around 3% of requests in the first window, 12% in the second (peaking at 30% for a few minutes), and 5% in the third. In total, about 655k requests failed. Our database was overloaded by our own background jobs.
No events were lost, and we are replaying all customer / entity creation events.
What happened?
Our API is backed by a Postgres database, and two background systems write to it heavily: one that syncs usage balances, and one that resets entitlements at the end of billing periods. Neither had a limit on how much database work it could create at once.
Three different triggers hit that same weakness:
12:00AM (27 minutes). Several large customers send us traffic spikes at the top of every hour. At midnight, the spike landed while a legacy reset job was already consuming most of the database's capacity, and the balance sync work generated tens of thousands of tiny concurrent transactions. The database hit 100% CPU and the connection pool was exhausted.
4:10AM (10 minutes). A long running script began executing. Each call unintentionally queued background reset work - roughly 244,000 jobs in 10 minutes.
9:00AM (26 minutes). The same hourly traffic spike arrived, and this time the first transient failures triggered a retry loop from one client - 200,000 retries in 24 minutes. The retries kept the database pinned long after the original spike had passed.
What we're doing
The root cause across all three incidents is the same: background database work had no global limit, so any burst could take down the database for everyone.
- We've shipped a rebuilt entitlement reset job: it processes small bounded batches, can never overlap itself, and has a kill switch we can flip instantly.
- We've shipped the ability to move heavy customers' usage tracking to async processing, which takes it off the critical database path.
- We're adding hard concurrency limits and back-pressure to all background database work, so a burst queues instead of overwhelming the database.
- We're adding admission control so a retry storm from one client can't extend an outage for everyone else.
We are happy to help and address any questions or concerns you may have
- Ayush
Resolved
The issue has now been fully resolved.
--- POST MORTEM ---
Autumn had three periods of partial outage overnight on July 23rd (UTC): 12:00–12:27AM, 4:10–4:20AM, and 9:00–9:26AM. During these windows, requests to our API failed with 503s - around 3% of requests in the first window, 12% in the second (peaking at 30% for a few minutes), and 5% in the third. In total, about 655k requests failed. Our database was overloaded by our own background jobs.
No events were lost, and we are replaying all customer / entity creation events.
What happened?
Our API is backed by a Postgres database, and two background systems write to it heavily: one that syncs usage balances, and one that resets entitlements at the end of billing periods. Neither had a limit on how much database work it could create at once.
Three different triggers hit that same weakness:
12:00AM (27 minutes). Several large customers send us traffic spikes at the top of every hour. At midnight, the spike landed while a legacy reset job was already consuming most of the database's capacity, and the balance sync work generated tens of thousands of tiny concurrent transactions. The database hit 100% CPU and the connection pool was exhausted.
4:10AM (10 minutes). A long running script began executing. Each call unintentionally queued background reset work - roughly 244,000 jobs in 10 minutes.
9:00AM (26 minutes). The same hourly traffic spike arrived, and this time the first transient failures triggered a retry loop from one client - 200,000 retries in 24 minutes. The retries kept the database pinned long after the original spike had passed.
What we're doing
The root cause across all three incidents is the same: background database work had no global limit, so any burst could take down the database for everyone.
- We've shipped a rebuilt entitlement reset job: it processes small bounded batches, can never overlap itself, and has a kill switch we can flip instantly.
- We've shipped the ability to move heavy customers' usage tracking to async processing, which takes it off the critical database path.
- We're adding hard concurrency limits and back-pressure to all background database work, so a burst queues instead of overwhelming the database.
- We're adding admission control so a retry storm from one client can't extend an outage for everyone else.
We are happy to help and address any questions or concerns you may have
- Ayush
Identified
The issue has been identified and we are working on resolving it.