Skip to main content
awsdown

AWS outage post-mortems

Every major AWS incident, analyzed: a timestamped timeline, the root cause in plain language, who felt it, and whether it earned SLA credits.

ELBmajor

The 2012 ELB Outage: The Deleted State That Ruined Christmas Eve

A developer running maintenance accidentally deleted Elastic Load Balancing state data in us-east-1. Without that state, load balancers could not be scaled or modified correctly, so a growing share of them degraded over Christmas Eve, most visibly taking Netflix offline for millions of viewers.

7 min read

DynamoDBcritical

The 2015 DynamoDB Outage: A Metadata Service Under Its Own Load

A brief network disruption made DynamoDB storage nodes re-request their partition assignments from a metadata service at the same moment that larger tables had made those requests slower and heavier. The metadata service could not keep up, storage nodes took themselves out of service, and DynamoDB errors in us-east-1 cascaded into EC2, SQS, and other services.

8 min read

EBScritical

The Great AWS Outage of 2011: When EBS Re-Mirrored Itself to Death

A network change accidentally routed high-volume EBS traffic onto a low-capacity network in one us-east-1 Availability Zone. Volumes lost their mirrors and tried to re-mirror all at once, exhausting capacity and creating a re-mirroring storm that stuck EBS and EC2 for days.

9 min read

Kinesiscritical

The 2020 Kinesis Outage: One Service That Broke Half of AWS

A routine capacity addition pushed the Kinesis front-end fleet past an operating-system thread limit in us-east-1. Because Kinesis quietly powers CloudWatch, Cognito, and much of AWS itself, the failure rippled into dozens of services for most of a business day.

8 min read

Lambdamajor

AWS Lambda Outage in us-east-1 (June 13, 2023): What Happened

A degradation in the subsystem that manages Lambda’s execution capacity caused elevated invocation error rates in us-east-1 for about three hours. Because Lambda sits inside so many AWS features - including STS and parts of the console - the blast radius reached far beyond “serverless” workloads.

6 min read

S3critical

The 2017 S3 Outage: A Typo That Broke the Web - Lessons That Still Apply

A playbook command executed with a mistyped parameter removed far more S3 index capacity than intended in us-east-1. The subsystems needed a full restart - something S3 hadn’t done in years at that scale - and for ~4 hours, a huge share of the web (including AWS’s own status dashboard) broke.

7 min read

Recent incident log

Smaller incidents from the live feed (last 90 days). Major events graduate into full post-mortems above.

  • [RESOLVED] Increased Error Rates

    August 19, 2026 · 3h 32m

    Amazon EC2 · eu-west-2 · minor

  • [RESOLVED] Increased 5xx Errors

    July 16, 2026 · 3h 38m

    Amazon CloudFront · global · major

  • [RESOLVED] Increased Launch Template API Error Rates

    July 6, 2026 · 2h 8m

    Amazon EC2 · us-east-1 · minor

  • [RESOLVED] Increased Error Rates and Latencies

    June 30, 2026 · 51m

    Amazon EC2 · eu-north-1 · minor