Autumn

Join our DiscordSubscribe to updates
Powered by
Privacy policy

·

Terms of service
Write-up
Redis Provider Issue
Partial outage
View the incident
What happened?

Our Redis provider runs a weekly maintenance window at Sunday 08:00 UTC. During this window they applied an upgrade to our instance that failed (post-mortem from them will be shared soon). The instance immediately hit 100% CPU at 08:00 BST, and the replica in our second AZ was affected too.

With the cache effectively down, every request became a cache miss and fell through to Postgres, which saturated within minutes. DB-backed customer/entity reads and get_or_create returned errors. To keep Postgres alive, we shed load by blocking and rate-limiting the heaviest paths until the cache recovered.

Resolution

The provider identified a bad rollback and reverted to their stable version, restoring the instance. We lifted our load-shedding as the cache warmed and CPU normalized, and buffered track events were replayed from the queue.

Follow-ups
  • Keep the cache warm so a cold cache can't trigger a Postgres thundering herd (warm-only cutover to a replacement instance).

  • Make Redis disposable we'll be coming up with a better plan to make sure our app is more resilient to Redis failures.

  • Better alerting to catch instance degradation the moment it starts.

As always, let us know if you have any questions.