Our Redis provider runs a weekly maintenance window at Sunday 08:00 UTC. During this window they applied an upgrade to our instance that failed (post-mortem from them will be shared soon). The instance immediately hit 100% CPU at 08:00 BST, and the replica in our second AZ was affected too.
With the cache effectively down, every request became a cache miss and fell through to Postgres, which saturated within minutes. DB-backed customer/entity reads and get_or_create returned errors. To keep Postgres alive, we shed load by blocking and rate-limiting the heaviest paths until the cache recovered.
The provider identified a bad rollback and reverted to their stable version, restoring the instance. We lifted our load-shedding as the cache warmed and CPU normalized, and buffered track events were replayed from the queue.
Keep the cache warm so a cold cache can't trigger a Postgres thundering herd (warm-only cutover to a replacement instance).
Make Redis disposable we'll be coming up with a better plan to make sure our app is more resilient to Redis failures.
Better alerting to catch instance degradation the moment it starts.
As always, let us know if you have any questions.