The issue has now been fully resolved. Thank you for your patience and we will be publishing a post mortem on this soon.
--- POST MORTEM ---
Autumn had an outage between 8:15PM and 8:50PM UTC. This affected ~30% of requests to our API. One of our downstream providers, Redis Cloud, went down. We had a single point of failure with them.
What happened?
We were investigating a minor Redis issue on our end. While we were investigating, we decided to upgrade our cluster.
An issue on their end caused our entire setup to fail. We were unable to restore it or deploy new instances.
What we’re doing
We take full responsibility and recognise that we are at the point where a single point of failure is unacceptable:
We are making failovers and redundancy our p0.
We’re defining a formal process for touching any critical infra.
We have just hired a infra/dev-ops expert to scale our services reliably
We’ll be prioritising SDK safety measures (eg, fail-open by default)
What you can do
An outage by us does not need to take your app down. This incident disproportionately affected customers that blocked usage on an Autumn error. This should never be the case.
If Autumn errors, allow requests to go through. The worst case is that users temporarily get additional usage. We can work with you to re-sync it after.
For extra safety, use our single webhook to replicate customer state into your own DB. If Autumn is down, you can fallback to that.
You can simulate errors from the Autumn API by removing your API key from your .env
Any questions you have on the incident, or resolution, we're always here.
Resolved
The issue has now been fully resolved. Thank you for your patience and we will be publishing a post mortem on this soon.
--- POST MORTEM ---
Autumn had an outage between 8:15PM and 8:50PM UTC. This affected ~30% of requests to our API. One of our downstream providers, Redis Cloud, went down. We had a single point of failure with them.
What happened?
We were investigating a minor Redis issue on our end. While we were investigating, we decided to upgrade our cluster.
An issue on their end caused our entire setup to fail. We were unable to restore it or deploy new instances.
What we’re doing
We take full responsibility and recognise that we are at the point where a single point of failure is unacceptable:
We are making failovers and redundancy our p0.
We’re defining a formal process for touching any critical infra.
We have just hired a infra/dev-ops expert to scale our services reliably
We’ll be prioritising SDK safety measures (eg, fail-open by default)
What you can do
An outage by us does not need to take your app down. This incident disproportionately affected customers that blocked usage on an Autumn error. This should never be the case.
If Autumn errors, allow requests to go through. The worst case is that users temporarily get additional usage. We can work with you to re-sync it after.
For extra safety, use our single webhook to replicate customer state into your own DB. If Autumn is down, you can fallback to that.
You can simulate errors from the Autumn API by removing your API key from your .env
Any questions you have on the incident, or resolution, we're always here.
Investigating
The provider has been notified of the issue and are actively investigating (one cluster restored). We will provide updates as soon as possible.