Dallas-Fort Worth Data Center Update

Dallas-Fort Worth Data Center Update

Message from Rackspace CEO Lanham Napier, July 9, 2009

Rackspace Community,

Some of our customers have been directly affected by recent outages in a portion of our Dallas-Fort Worth Data Center. Others of you may have heard about it and are following it closely. An interruption like this is not up to our Fanatical Support standards and we are working hard to prevent such incidents from occurring in the future.

On behalf of Rackspace, I sincerely apologize for these disruptions. We know these failures negatively impacted the lives and businesses of our customers. After the disruptions occurred we did our best to recover quickly and explain what happened in a transparent fashion. Since we have seen erroneous and incomplete information on the web and in the media, we wanted to share with you with the most up-to-date and accurate information.

What happened

First, some context. Our DFW Data Center has three phases, or sections, and the outages were caused by a malfunction in our power infrastructure in Phase 1. We have redundancy in place, and this redundancy generally works as intended, but these outages show that we clearly have room for improvement.

While we take any outage seriously, it is important to know that this is not a pervasive issue across Rackspace. We operate nine facilities worldwide and the problems in DFW are not affecting our other data centers. Unfortunately, these localized incidents in DFW have had a disproportionate impact on some customers.

Here’s a quick recap of the outages and near-term resolution activities:

  1. We had a power interruption on June 29, 2009 in Phase 1 of our DFW Data Center, and we moved some of our customers to generator power. The generators then experienced a failure, which caused those customers to lose power to their servers for approximately 40 minutes. We have since performed maintenance and upgrades to those generators, with the help of experts from companies like Cummins, GE and Eaton, and the generators are now stable.
  2. We experienced another power interruption on July 7, 2009. Again, we moved customers to generator power. During this outage we also suffered a loss of network connectivity due to the power disruption. The part of the power infrastructure that failed (a “bus duct”) prevented proper operation of our UPS for that section, so some customers lost power to their servers for about 20 minutes before we could get them onto generator power. We have since replaced the failed bus duct, and that section of the data center is back to normal and running on utility power.

If you would like more detailed information on the June 29 interruption, please refer to the June 29 Incident Review. The Incident Report for July 7th is forthcoming. I also wanted to speak to all of you in some way other than anonymous copy on a screen, so this morning, I recorded a video which follows this letter.

What we’re doing about it

The above resolution steps address the near-term issues. Now we are digging into the actions we need to take to prevent these types of outages in the future. Let me be clear: data centers will experience power interruptions, parts will break, and servers will go down. No data center is completely risk-free. But we can manage and mitigate the risk to acceptable levels, better than we have today, and we can make sure our recovery is as quick as physically possible. I have no doubt that we will get better and stronger from this situation.

Our main actions include the following steps:

  1. Put our best people on it, and bring in the experts. I am personally going to locate myself in our DFW data center until I am satisfied that our repairs and maintenance are complete. We have assembled our best talent from the US and the UK to focus on the issues there. And we have brought in top talent from our vendors, as well as knowledgeable outside consultants, to assist us.
  2. Assess the status of the infrastructure. We are combing through the power systems in DFW and assessing every link in the chain. Based on the advice of our experts, we will update every piece that needs updating to ensure the performance we require.
    • Phase I has four zones within it. At this point we have completed work on the major power systems for each zone by remediating known deficiencies at the generator and UPS levels.
    • We will continue our work through the smaller components of each zone including switches, breakers and ducting. At this point we have completed all of the work on the smaller components for one of our zones and preventative maintenance on the other three zones is underway.
    • We will complete all of this work as soon as possible with minimal disruption to customers.
  3. Improve standard operating procedures. We are going to increase the frequency of our testing, monitoring and measurement programs within DFW. Our maintenance schedules will change. And the level of detail we review internally and share externally will increase.
  4. Invest. We will continue to invest in our infrastructure.