The failure of two power components at a Virginia data center affected some EC2 operations on December 9th, Amazon Web Services says.
Amazon Web Services has attributed a 44-minute outage in part of its Northern Virginia data center last week to the failure of power supply in one "availability zone" in the data center, which was soon followed by a second failure of a component in the redundant system.
Users of the Amazon EC2 cloud with workloads in Amazon's Northern Virginia data center experienced problems early in the morning of December 9, with some operations in a part of the data center interrupted during a five-hour period.
Amazon started notifying customers of a problem at 4:08 a.m. Eastern. By 9:41 a.m., it's Amazon Service Health Dashboard reported that "we have completed recovery of most instances affected by this event."
The postings first mentioned a connectivity issue, then acknowledged a power issue. In following up on the postings, InformationWeek asked Amazon whether the power issue was inside the data center or an issue with an external supplier.
Amazon spokesmen responded that "a single component of the redundant power distribution system failed in this zone. Prior to completing the repair of this unit, a second component, used to assure redundant power paths, failed as well, resulting in a portion of the servers in that availability zone losing power."
Forty-three minutes after the first notice, Amazon Web Services posted messages, such as: "The underlying power issue has been addressed, instances have begun to recover" at 4:51 a.m. Eastern and "most affected instances and are operating normally" at 5:11 a.m.
The next day it added an explanation: "A single component of the redundant power distribution system failed in this zone. Prior to completing the repair of this unit, a second component, used to assure redundant power paths, failed as well, resulting in a portion of the servers in that availability zone losing power."
The actual time of the outage, according to a monitoring service that gathers information by pinging traffic over the Internet and off the accounts it maintains inside Amazon facilities, indicated 3:34 a.m. to 4:19 a.m. Eastern.
How Enterprises Are Attacking the IT Security EnterpriseTo learn more about what organizations are doing to tackle attacks and threats we surveyed a group of 300 IT and infosec professionals to find out what their biggest IT security challenges are and what they're doing to defend against today's threats. Download the report to see what they're saying.
IT Strategies to Conquer the CloudChances are your organization is adopting cloud computing in one way or another -- or in multiple ways. Understanding the skills you need and how cloud affects IT operations and networking will help you adapt.