Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Disasters are unpredictable, but Microsoft datacenters and operations personnel prepare for disasters to provide continuity of operations if unexpected events occur. Resilient architecture and up-to-date tested continuity plans mitigate potential damage and promote swift recovery of datacenter operations. Crisis management plans provide clarity on roles, responsibilities, and mitigation activities before, during, and after a crisis. The roles and contacts defined in these plans facilitate effective escalation up the chain of command during crisis situations.
Business resilience
Under Microsoft Cloud Operations and Innovation (CO+I) Business Continuity Program, datacenters are required to test the continued operation and response to crisis events. Each Microsoft managed datacenter has its own business continuity plan, created by using the key subject matter expertise from CO+I Resilience Center of Excellence and Datacenter Operations to ensure that site-specific context is factored into emergency preparedness. These plans describe roles, responsibilities, personnel safety procedures, notification criteria, escalations steps, and checklists for different disaster scenarios.
Microsoft's CO+I organization's Resilience function is governed by the Enterprise Business Continuity Management program and follows the Enterprise Policies and Standards. The Business Continuity Council, departmental leadership, and ultimately Microsoft's Senior Leadership Team review the program's performance on a periodic basis.
Crisis management and pandemic response
The Crisis Management Program is an integral part of Microsoft's response to major events given its global presence. Microsoft's Datacenter Crisis Management Plan is based on industry-best practices and includes the critical components required to allow for a tactical approach in responding to major events. In addition, CO+I Resilience Center of Excellence developed and continues to maintain a Pandemic and Infectious Disease Plan that is used to respond to infectious diseases that may have an operational impact. As part of our pandemic response, the resilience support team provides critical and timely local disease intelligence to Redmond-based Microsoft leadership to facilitate a comprehensive mitigation strategy.
Microsoft has established an organization-wide Enterprise Resilience and Crisis Management (ERCM) framework that serves as a guideline for developing Business Continuity Program across the company. The program includes Business Continuity Policy, Implementation Guidelines, Business Impact Analysis (BIA), Risk Assessment, Dependency Analysis, and procedures for monitoring and improving the program. Enterprise Resilience Office manages the governance and performance reporting across Microsoft. The CO+I Resilience program is coordinated through the CO+I Resilience Center of Excellence to ensure that the program adheres to a coherent long-term vision and mission, and is consistent with enterprise program standards, methods, policies, and metrics. CO+I Resilience Center of Excellence established a series of Standards designed to provide additional governance to the CO+I organization.
The CO+I Technology Resilience Plans (TRPs) are intended for various Engineering Groups within CO+I for the recovery from high-severity incidents or disasters to help ensure that our critical technology remains available.
The Business Resilience Plan (BRP) and TRP includes scope and applicable dependencies for the services, restoration procedures, and communications with the Incident Management team. Dedicated plan owners review and approve the BRP and TRP at least annually and make them available to all applicable users. The plans are tested per the defined testing schedule as part of applicable standards.
Resiliency Program
Microsoft defined the BRP to serve as a guide to respond, recover, and resume operations during a serious adverse event. The BRP covers the key personnel, resources, services, and actions required to continue critical business processes and operations. The development of the BRP is based on recommended guidelines of Microsoft's Enterprise Resilience Office.
In scope for this plan are Microsoft's critical business processes, defined as needed within 24 hours or less. These processes are determined during a BIA, in which Microsoft estimated potential operational and financial impacts if they couldn't perform a process and determined the Recovery Time Objective (RTO) and Recovery Point Objective (RPO). Following the BIA, a Non-Technical Dependency Analysis is performed to determine the specific people, applications, vital records, and user requirements necessary to perform the process.
Microsoft periodically tests the BRP to assess its effectiveness, usability, and to identify areas where risks can be eliminated or mitigated. When applicable, third parties are involved in the test if there are dependencies associated with them. The results of testing are documented, validated, and approved by appropriate personnel. This information is used to create and prioritize work items.
Datacenter resilience program
As part of the datacenter resilience program, the CO+I Resilience Center of Excellence team develops the methods, policies, and metrics that address the information security requirements needed for the organization's business continuity. The team develops TRPs for the continued operations of critical processes and required resources if disruptions occur.