Group Purchasing
Group Purchasing

OT Disaster Recovery Quick Start Guide

OT Disaster Recovery Quick Start Guide (PDF, 0.69MB)Published: 26 Sep, 2025

The OT Disaster Recovery Quick Start Guide, published by SANS ICS in version 1.0, provides a framework for building a Disaster Recovery Plan (DRP) for Industrial Control System (ICS) and Operational Technology (OT) environments. Authored by Mike Hoffman and Saltanat Mashirova, the tri-fold guide covers the plans that make up Business Continuity Management, the three phases of OT recovery, and the criteria that determine when an incident qualifies as a disaster.

Key concepts:

  • An OT Disaster Recovery Plan (DRP) is a subset of the broader Business Continuity Management (BCM) framework, which also includes the Business Continuity Plan, Continuity of Operations Plan, Crisis Communications Plan, and Occupant Emergency Plan
  • Recovery Time Objective (RTO) defines the maximum amount of time a system resource can remain unavailable before causing unacceptable impact to other systems and the Maximum Tolerable Downtime (MTD)
  • Recovery Point Objective (RPO) represents the point in time, prior to a disruption, to which mission or business data can be recovered; unlike RTO, RPO is not considered part of the MTD
  • Maximum Tolerable Downtime (MTD) is the key parameter driving recovery strategy, since it defines the total time a business will accept for a production process outage
  • OT recovery follows three phases: Activation and Notification, Recovery, and Reconstitution
  • Disaster severity is classified on a four-level scale from SD-3 (anomaly, no process downtime) to SD-0 (major multi-site incident with downtime exceeding the MTD)
  • Only SD-0 scenarios are formally considered disasters that invoke the disaster recovery process or Emergency Response Team
  • Dependency analysis in OT recovery covers six relationship types: direct dependency, indirect dependency, co-dependency, interdependency, co-dependency off a shared system, and redundant dependency
  • Large-scale OT disaster recovery involves more than restoring from backups; it requires accounting for automation functions, recovery priorities, dependencies, and data reconstitution of automation systems

The guide's central point is that OT disaster recovery is fundamentally different from IT disaster recovery because of the cyber-physical nature of industrial systems. Recovery isn't complete once data is restored; systems must also be validated for correct physical operation and reconstituted with the most current operational parameters before a plant can safely resume production. Contributors: Mike Hoffman and Saltanat Mashirova, SANS ICS.

FAQ

RTO (Recovery Time Objective) defines the maximum acceptable downtime for a system before it impacts other systems and the MTD, while RPO (Recovery Point Objective) defines the point in time to which data must be recoverable after an outage; RPO is not considered part of the MTD.

The three phases are Activation and Notification (initial response after a disruption is detected), Recovery (restoring systems once they are safe and cyber-ready), and Reconstitution (final validation that recovered systems have the correct operational parameters before restart).

An incident is classified as a disaster when it reaches Shutdown 0 (SD-0) severity: a major or widespread incident affecting multiple sites and essential functions, with downtime equal to or exceeding the Maximum Tolerable Downtime (MTD).

No. An OT DRP is a subset of the broader Business Continuity Management framework, which also includes the Business Continuity Plan, Continuity of Operations Plan, Crisis Communications Plan, and Occupant Emergency Plan.

Because industrial control systems are cyber-physical, recovery also requires validating that automation functions operate correctly, resolving system dependencies, and reconciling any concurrent data processing before a plant can safely resume operations.

Authors