The topic of major outages in modern data centers raises questions that were not addressed in the first article — we decided to expand on it.

According to statistics from the Uptime Institute, the majority of incidents in data centers are related to power supply failures — they account for 39% of occurrences. Following that is the human factor — contributing another 24% of outages. The third most significant cause (15%) was failures in the cooling system, while natural disasters took fourth place (12%). The total share of other issues is only 10%. While not questioning the data of a reputable organization, let's highlight something common in different outages and try to understand whether they could have been avoided. Spoiler: in most cases, they could.
The Science of Contacts
In simple terms, there are only two problems with power supply: either there's no contact where it should be, or there is contact where it shouldn't be. One can argue for a long time about the reliability of modern uninterruptible power supply systems, but they don’t always save the day. Take, for instance, the notorious case involving British Airways' data center, owned by the parent company, International Airlines Group. There are two such facilities located near Heathrow Airport — Boadicea House and Comet House. In the former, an accidental power shutdown occurred on May 27, 2017, which led to an overload and failure of the UPS system. As a result, some IT equipment was physically damaged, and it took three days to rectify the last outage.
The airline had to cancel or postpone over a thousand flights, with around 75,000 passengers unable to depart on time — compensation payments amounted to $128 million, excluding the costs associated with restoring the data centers' functionality. The reasons behind the blackout remain unclear. If we believe the results of the internal investigation voiced by International Airlines Group's CEO, Willie Walsh, it occurred due to an engineer error. Nevertheless, the uninterruptible power supply system should have withstood such a shutdown — that’s what it was installed for. The data center was managed by specialists from the outsourcing company CBRE Managed Services, which led British Airways to attempt to claim damages through the London courts.

Power outages occur in similar scenarios: first, a shutdown happens due to the energy supplier's fault, sometimes due to bad weather or internal issues (including staff errors), and then the uninterruptible power supply system fails to handle the load or a brief interruption in the sine wave causes failures of many services, which take a lot of time and money to restore. Can such accidents be avoided? Absolutely. If the system is designed correctly, however, even the creators of large data centers are not immune to mistakes.
Human Factor
When the immediate cause of an incident is the incorrect actions of data center personnel, the problems most often (but not always) affect the software part of the IT infrastructure. Such accidents happen even in large corporations. In February 2017, due to a typo made by a member of the technical operations group at one of the data centers, some servers of Amazon Web Services were shut down. The mistake occurred during the debugging of the billing process for customers of Amazon Simple Storage Service (S3). The employee attempted to delete a certain number of virtual servers used by the billing system, but mistakenly affected a larger cluster.

As a result of the engineer's error, servers running important software modules of Amazon's cloud storage were deleted. Primarily, the indexing subsystem, which contains information about the metadata and location of all S3 objects in the US region US-EAST-1, was affected. The subsystem used for data storage and managing the available storage space was also impacted. After the virtual machines were deleted, these two subsystems required a full restart, and Amazon engineers faced a surprise — for an extended period, the public cloud storage could not handle customer requests.
The impact was significant, as many major resources utilize Amazon S3. The outages affected Trello, Coursera, IFTTT, and, most unfortunately, services of Amazon's major partners from the S&P 500 list. It's hard to quantify the damage in such cases, but it was estimated to be in the hundreds of millions of dollars. As you can see, a single incorrect command is enough to disrupt the service of the largest cloud platform. This is not an isolated incident; on May 16, 2019, during maintenance, the Yandex.Cloud service virtual machines of users in the ru-central1-c zone that had ever been in SUSPENDED status. Client data has already suffered here, some of which was irretrievably lost. Of course, people are imperfect, but modern information security systems have long been able to monitor the actions of privileged users before executing the commands they input. If such solutions were implemented in Yandex or Amazon, similar incidents could be avoided.

Frozen cooling
In January 2017, a major incident occurred at the Dmitrov data center of the company Megafon. At that time, the temperature in the Moscow region dropped to -35 °C, which led to the cooling system of the facility failing. The operator's press service did not elaborate on the causes of the incident — Russian companies are extremely reluctant to discuss accidents at their facilities; in terms of publicity, we lag significantly behind the West. There was a rumor circulating on social media about the coolant freezing in the pipes laid along the street and an ethylene glycol leak. If this is to be believed, the maintenance service could not promptly obtain 30 tons of refrigerant due to long holidays and had to resort to makeshift solutions, organizing improvised free cooling while violating operational rules. The severe cold exacerbated the problem — in January, Russia unexpectedly experienced winter, although no one was expecting it. As a result, staff had to power down part of the server racks, which caused some of the operator's services to be unavailable for two days.

One could say this is related to a weather anomaly, but such severe frosts are not unusual for the capital region. Winter temperatures in the Moscow region can drop to much lower levels, which is why data centers are built to operate reliably at −42°C. Most often, cooling systems fail in the cold due to insufficient glycol concentration and excess water in the coolant solution. There can also be problems with pipe installation or design calculations and testing related primarily to the desire to cut costs. As a result, a serious accident occurs that could have been easily avoided.
Natural Disasters
Most often, storms and/or hurricanes disrupt the engineering infrastructure of the data center, leading to service outages and/or physical equipment damage. Weather-related incidents happen quite frequently. In 2012, Hurricane Sandy swept along the U.S. West Coast with heavy rainfall. The Peer 1 data center located in a skyscraper in Lower Manhattan , after saltwater flooded the basements. The facility's emergency generators were placed on the 18th floor, and their fuel supply was limited — regulations introduced in New York after the 9/11 attacks prohibit storing large amounts of fuel on upper floors.
The fuel pump also failed, so staff had to manually transport diesel for the generators for several days. The team's heroism saved the data center from a major disaster, but was it really necessary? We live on a planet with a nitrogen-oxygen atmosphere and abundant water. Storms and hurricanes are commonplace here (especially in coastal areas). Designers should probably consider the associated risks and plan a proper uninterruptible power supply system. Or at least choose a more suitable location for the data center than a skyscraper on an island.
Everything else
In this category, the Uptime Institute highlights a variety of incidents, making it difficult to identify a typical one. Copper cable thefts, vehicles crashing into data centers, power line supports, transformer substations, fires that damage fiber optics due to excavators, rodents (rats, rabbits, and even wombats, which are marsupials), as well as enthusiasts practicing shooting at wires — the menu is extensive. Power outages can even be caused by illegal marijuana plantation. In most cases, the culprits of the incident are specific individuals, meaning we are once again dealing with the human factor, where the problem has a name and a surname. Even if at first glance the accident is related to a technical failure or natural disasters, it can be avoided with proper design and operation of the facility. The exceptions are cases of critical damage to data center infrastructure or destruction of buildings and structures due to natural disasters. These are indeed force majeure circumstances, while all other issues arise from the connection between the computer and the chair — arguably, the most unreliable part of any complex system.
Source: habr.com
