AWS incident history
2 active incidents and a full searchable archive of resolved AWS outages.
Sourced live from the official AWS Health Dashboard. How we verify.
Showing 1-10 of 25 incidents
[RESOLVED] Increased Error Rates
Between August 18 5:00 PM and August 19 11:00 AM PDT, we experienced elevated errors launching EC2 instances in a newly launched Availability Zone (euw2-az4) in the EU-WEST-2 Region. After the new Availability Zone launch, we began experiencing errors when using a default VPC. We discovered the root cause of the issue on August 19 at 9:00 AM and began deploying a change to resolve the issue at 9:30 AM. While the change was underway, we began to see incremental improvements in new instance launches, with full recovery at 11:00 AM. Existing running instances and resources were not affected. Some regional services, such as Lambda functions or Aurora databases, were not available at the launch of the new Availability Zone and service availability will be added over time. Customers attempting to create resources before the services become available will see a message reporting that it is not supported in the Availability Zone. The issue has been resolved and the service is operating normally.
[RESOLVED] Increased 5xx Errors
Between 12:45 AM and 4:18 AM PDT, we experienced increased 5xx errors for CloudFront customers utilizing VPC Origins connectivity. Our engineers were automatically engaged and immediately began investigating the root cause. By 2:57 AM PDT, we identified the root cause of the issue as an internal constraint on the fleet that manages connections to private VPC origins. When this constraint was reached, the system responsible for distributing routing configuration to our network processors failed to load the updated configuration data correctly, affecting routing of VPC Origin connections. At 3:52 AM PDT, we took multiple mitigation actions that led to to full recovery at 4:18 AM PDT. Now that the issue has been mitigated, customers who temporarily changed their origin type can safely revert these changes. Customers utilizing other origin types were not affected by this issue. The issue has been resolved and the service is operating normally.
[RESOLVED] Increased Launch Template API Error Rates
Between 2:56 AM and 6:54 AM PDT, we experienced increased error rates when calling EC2 Launch Template APIs in the US-EAST-1 Region. During this time, affected customers may have experienced errors when creating, modifying, or describing Launch Templates. Other AWS services that rely on Launch Templates were also impacted. Amazon EC2 instances and Amazon EKS workloads already running on provisioned nodes continued to operate normally. Cluster modification operations, and Managed Node Group creation were also impacted. For EKS Auto Mode, impact was limited to operations requiring new capacity or changes, including node provisioning and pod scheduling. Our engineers were automatically engaged and immediately began investigating the root cause. We identified the root cause as a congestion issue within an EC2 internal subsystem responsible for processing EC2 launch template workflows. At 3:26 AM PDT, we took mitigation actions by introducing throttling for the affected APIs and we saw some recovery which was communicated directly with a subset of customers via the 'Your Account view' of the AWS Health Dashboard. We took multiple additional mitigation paths, incrementally lifting these throttle limits, and by 6:54 AM PDT, the issue was fully mitigated. We recommend that customers retry any failed requests. The issue has been resolved and all AWS services are now operating normally.
[RESOLVED] Increased Error Rates and Latencies
Between 1:42 PM and 2:25 PM PDT we experienced increased error rates and latencies for EC2 APIs in the EU-NORTH-1 Region. This issue also affected new instance launches. Other AWS Services that launch new instances or call EC2 APIs as part of their workflows were also affected by this issue. Existing EC2 instances were unaffected by this issue. During this time, customers would have received an Internal Server Error in the Management Console and APIs. Engineers were automatically engaged and began investigating the root cause. We identified the root cause as a planned configuration change. This change was reverted and we began observing recovery at 2:19 PM. By 2:25 PM, the issue was fully mitigated. We do not expect this issue to reoccur. Since the issue was mitigated at 2:25 PM, we have been processing a backlog for ELB workflows and expect this backlog to complete within the next 30 minutes. We recommend customers retry requests that failed during this time. The issue has been resolved and all services are operating normally.
[RESOLVED] Increased API Error Rates
Between 4:00 PM and 4:46 PM, we experienced increased error rates for the Route53 APIs. This issue did not impact resolution of existing DNS records. Engineers were automatically engaged and immediately began investigating the issue. During this time, customers may have received 500s for Route53 APIs and the Route53 Management Console. We have identified the root cause and have mitigated this issue. Other AWS Services that call the Route53 APIs in their workflows may also have been impacted during this time. We recommend retrying any failed operations or stuck workflows. We do not expect this issue to reoccur. The issue has been resolved and the service is operating normally.
[RESOLVED] Increased Error Rate and Latency
Starting May 7 4:20 PM PDT, we experienced increased impaired EC2 instances and degraded EBS volumes in a single facility (data center) within a single Availability Zone (use1-az4) in the US-EAST-1 Region. The issue was caused by a thermal event resulting in a loss of power. As part of our recovery effort, we shifted traffic away from the impacted Availability Zone for most services at May 7 5:06 PM. AWS services, like Elastic Load Balancing, Elastic Kubernetes Service, ElastiCache, Redshift, OpenSearch, Managed Streaming for Apache Kafka among others, that depend on the affected EC2 instances and EBS volumes in this Availability Zone, also experienced elevated error rates and latencies for some workflows and/or configurations. Our main effort during the event mitigation strategy was to bring back our cooling systems capacity. By May 8 1:50 PM, we were able to stabilize cooling system capacity to pre-event levels, which helped us to restore the majority of the impaired EC2 instances and EBS volumes. A small number of instances and EBS volumes remain impaired and we continue to work to recover all affected remaining resources. We will communicate with customers who are still impacted via the Your Account view of the AWS Health Dashboard. Customers that require further assistance with this event may contact AWS Support through the AWS Management Console or the AWS Support Center.
[RESOLVED] Increased Connectivity Issues
Between 3:58 AM and 4:40 AM PDT, we experienced increased error rates and increased launch failures for EC2 instances in a single Availability Zone (euw3-az2) in the EU-WEST-3 Region. During this time, customers attempting to launch new EC2 instances in the affected Availability Zone would have experienced launch failures. Additionally, a subset of existing EC2 instances and EBS volumes in this Availability Zone were impacted and became unreachable. We have identified the root cause to be a loss of power to infrastructure within the affected Availability Zone. Engineers were engaged at 4:02 AM and immediately began working to restore power and assess the scope of impact. By 4:20 AM, power was successfully restored to the affected infrastructure. We then focused our efforts on recovering impacted EC2 instances and EBS volumes. By 4:40 AM, all impacted EC2 instances and EBS volumes had been fully recovered and were operating normally. No additional action is required for EC2 instances and EBS volumes that were impacted during the power loss event, as these have been fully recovered. While EC2 and EBS have recovered, some AWS services may take additional time to fully recover as they process backlogs and complete their own recovery procedures. The issue has been resolved and the service is operating normally.
[RESOLVED] Increased Error Rates
Between 11:27 AM and 12:20 PM PST we experienced substantial error rates for S3 PUT/GET requests in EU-CENTRAL-2 Region. Engineers were engaged immediately based on automated alarming. We identified the root cause as an issue with a subsystem responsible for assembling objects bytes in storage. At 12:04 PM PST, we implemented mitigations and began observing early signs of recovery for S3. Error rates continued to improve, and other AWS Services continued to recover until 12:50 PM PST when we observed full recovery. We continue to work toward backfilling Cloudwatch logs, and expect that to continue over the next couple hours. We recommend customers retry any failed requests. The issue has been resolved and all services are operating normally.
Increased Error Rates
We are providing an update on the ongoing service disruption. The Middle East (Bahrain) Region (ME-SOUTH-1) has suffered damage due to the conflict in the Middle East and is currently unavailable. Customers should recover their resources in other Regions from remote backups. Relevant billing operations are currently suspended while we restore normal operations in this AWS Region. This process is expected to take several months.
Increased Error Rates
We are providing an update on the ongoing service disruption. The Middle East (UAE) Region (ME-CENTRAL-1) has suffered damage as a result of the conflict in the Middle East and is currently unable to reliably support customer applications. While some workloads continue to function normally, we strongly recommend customers migrate all accessible resources to other Regions and restore inaccessible resources from remote backups as soon as possible. Relevant billing operations are currently suspended while we restore normal operations in this AWS Region. This process is expected to take several months.
Page 1 of 3