A developer running maintenance accidentally deleted Elastic Load Balancing state data in us-east-1. Without that state, load balancers could not be scaled or modified correctly, so a growing share of them degraded over Christmas Eve, most visibly taking Netflix offline for millions of viewers.
A brief network disruption made DynamoDB storage nodes re-request their partition assignments from a metadata service at the same moment that larger tables had made those requests slower and heavier. The metadata service could not keep up, storage nodes took themselves out of service, and DynamoDB errors in us-east-1 cascaded into EC2, SQS, and other services.
A network change accidentally routed high-volume EBS traffic onto a low-capacity network in one us-east-1 Availability Zone. Volumes lost their mirrors and tried to re-mirror all at once, exhausting capacity and creating a re-mirroring storm that stuck EBS and EC2 for days.
A routine capacity addition pushed the Kinesis front-end fleet past an operating-system thread limit in us-east-1. Because Kinesis quietly powers CloudWatch, Cognito, and much of AWS itself, the failure rippled into dozens of services for most of a business day.
A latent race condition in DynamoDB’s DNS management automation left the regional endpoint with an empty DNS record. Because half of AWS depends on DynamoDB in us-east-1, the failure cascaded into EC2 launches, Lambda, and NLB health checks for ~15 hours.
A degradation in the subsystem that manages Lambda’s execution capacity caused elevated invocation error rates in us-east-1 for about three hours. Because Lambda sits inside so many AWS features - including STS and parts of the console - the blast radius reached far beyond “serverless” workloads.
An automated scaling activity triggered unexpected behavior in the network connecting AWS’s internal control plane to the main network. The resulting congestion blinded AWS’s own monitoring, broke the consoles and EC2 APIs, and degraded services from SQS to Ring doorbells for about seven hours.
A playbook command executed with a mistyped parameter removed far more S3 index capacity than intended in us-east-1. The subsystems needed a full restart - something S3 hadn’t done in years at that scale - and for ~4 hours, a huge share of the web (including AWS’s own status dashboard) broke.
7 min read
Recent incident log
Smaller incidents from the live feed (last 90 days). Major events graduate into full post-mortems above.