Cloud SCADA Security • Incident Response

OT Incident Response Plan for Cloud SCADA Operations

Merobix Engineering • • 11 min read

Most industrial operators have an emergency response plan for a leak, a fire, or a lightning strike - and no equivalent for the day an operator credential starts making setpoint changes nobody authorized. An OT incident response plan is that missing document: who takes command, which playbook applies, what gets isolated and what absolutely does not, how evidence is preserved, and when the regulator gets called. This guide walks through building one for a cloud SCADA operation - the roles, the scenario playbooks, the vendor's share of the work, the tabletop exercises that make it real, and the evidence that lets you prove afterward what actually happened.

Back to Blog

Part of the SCADA Security Hub - 60+ guides on SCADA security, compliance & certifications.

4Phases in the NIST-Style Incident Lifecycle
72Hours for Covered Pipeline Operators to Report to CISA
2Tabletop Exercises Per Year, Minimum

Why OT Incident Response Is Different from IT

The standard IT incident playbook - isolate the machine, image it, wipe it, restore from backup - is built on an assumption OT cannot make: that turning things off is safe. In an industrial operation, the compromised system may be supervising a compressor station, a chlorine feed, or a separator train. Disconnecting it is itself a process event with safety consequences, and the cure can genuinely be worse than the disease. Oldsmar-style scares and the Colonial Pipeline shutdown both illustrate the pattern: the operational decision - keep running, degrade, or stop - dominates the technical one.

That inversion drives everything else in this guide. OT incident response puts safety first, availability second, and confidentiality third - the reverse of IT's ordering. It requires people IT plans never include: console operators, process engineers, operations leadership with authority to slow or stop production. Containment options must be engineered and pre-approved in calm conditions - through management of change, not improvisation at 2 AM. And "fall back to manual" must be a rehearsed procedure with named people who still know how, not a phrase in a binder. For the deeper background on why IT playbooks break in plants, see OT vs IT security differences.

The Incident Lifecycle, Applied to OT

The classic NIST SP 800-61 lifecycle - preparation; detection and analysis; containment, eradication and recovery; post-incident activity - holds up well in OT as a skeleton, and its successor, the CSF 2.0-aligned SP 800-61 Revision 3, carries the same ideas into NIST's current risk-management guidance. What changes is the content of each phase:

Phase Classic IT Emphasis OT / Cloud SCADA Emphasis
PreparationTooling, contacts, backupsAll of that, plus pre-approved containment options, manual-operation procedures, spare hardware, printed runbooks that survive a network outage
Detection & analysisEDR and SIEM alerts on endpointsPlatform security alerts, auth anomalies, unexplained setpoint or command activity, telemetry that stops making physical sense
ContainmentIsolate host, reset credentialsRevoke sessions and disable accounts first (non-disruptive), then graduated isolation - always weighed against process impact and safety
Eradication & recoveryReimage, restore, patchVerified-clean restores, gateway re-verification, controlled reconnection site by site, heightened monitoring
Post-incidentLessons-learned meetingLessons learned plus regulatory reporting closure, evidence retention, and control changes fed into MOC

One cloud-specific note on containment: in a cloud SCADA platform, your fastest containment tools are identity-layer actions - revoking sessions, disabling accounts, rotating API keys - which take effect immediately and touch nothing physical. Network isolation, the classic IT reflex, is the second rung on the ladder, used when identity actions are not enough.

Roles: Who Does What When the Pager Goes Off

An IR plan without named humans is a wish. The core roles, which in a small operator may be four people wearing six hats:

Write an on-call rotation, publish an escalation path with phone numbers (printed - the incident may take email with it), and pre-authorize the big calls in writing: who may disable any account instantly, who may isolate a site, who may declare fallback to manual. Decision rights argued about during an incident are decision rights you do not have.

Scenario Playbooks Worth Writing

Generic plans produce generic paralysis. Write short, specific playbooks - one to two pages each - for the incidents you actually expect:

Each playbook ends the same way: evidence steps, reporting triggers, and the named role that owns each action.

Detection: You Cannot Respond to What You Cannot See

Every playbook above begins with a signal, which makes detection capability a prerequisite of response, not a separate topic. In practice the signals that matter in cloud SCADA are: repeated authentication failures and lockout events, privilege probes, cross-tenant denial attempts, logins from unexpected origins, commands and setpoint writes outside normal patterns, and telemetry freshness anomalies. Merobix generates these detections at runtime and pushes them through a durable notification outbox - with retries and dead-letter handling rather than fire-and-forget email - to on-call staff by email, SMS, or webhook, and into your SIEM so OT events land in the console your security team already watches. Whatever platform you run, verify two things: that these event classes exist at all, and that their delivery is assured - an alert that silently fails to send is a playbook that never starts. Our threat detection and SIEM guide treats this layer in full.

The Vendor's Share: IR in the Shared Responsibility Model

Cloud SCADA splits incident response the same way it splits everything else. The vendor owns platform-level incidents: infrastructure compromise, platform vulnerabilities, cross-tenant events. For those, your contract should commit them to timely customer notification with defined windows, preserved evidence and audit data relevant to your tenant, status updates during the incident, and a post-incident summary. Ask to see their internal IR readiness the same way you would ask about pen testing and validation - an honest vendor will describe their detection stack and their practice cadence, including what is still maturing.

The operator owns incidents in their own scope: their users and credentials, their sites, their field devices, their regulatory reporting. Crucially, the vendor's controls become your response tools here - logout-everywhere and session revocation, account disablement, API key rotation, audit exports, alert escalation. During evaluation, rehearse this: in a demo, have the vendor walk your compromised-credential playbook end to end and show each control functioning. Five minutes of that is worth fifty pages of assurances.

Evidence: The Audit Trail Is Your Star Witness

After containment, every question becomes historical: What did the attacker touch? Which commands executed? When did it start? Your answers are only as good as your records - and only as defensible as those records are tamper-proof. Preserve, from the first hour: the platform audit trail (logins, failures, configuration changes, setpoint writes and acknowledgements, with named accounts and timestamps), security alerts and detections, boundary network logs, gateway and device logs, session and account details for affected users, and a handwritten-if-necessary timeline of every response action your team takes.

Two properties separate usable evidence from arguable evidence. Immutability: Merobix keeps chained, tamper-evident audit records for sensitive events, with chain verification - so you can demonstrate the log was not edited after the fact, which matters to regulators, insurers, and courts alike. Attribution: command records with full lifecycle states (pending, applied, verified, failed) and telemetry stamped with device, site, and sample identity let you reconstruct not just who clicked, but what the process actually received. Export early, snapshot storage before rebuilding anything, and loop in counsel and your insurer before destroying anything they may later need. The full evidence story is in our audit trail guide.

Reporting Obligations: Know Your Clock Before You Need It

Regulatory reporting windows are short and unforgiving of improvisation. Covered pipeline operators under the TSA security directives (SD Pipeline-2021-01 series) must report cybersecurity incidents to CISA - no later than 72 hours after identifying one under the current SD Pipeline-2021-01F revision, which lengthened the original 24-hour window. Water utilities have AWIA emergency response plan obligations and state-level rules that vary. Electric sector entities face NERC CIP reporting requirements. Publicly traded companies have SEC disclosure timelines for material incidents, and CIRCIA rulemaking is extending federal reporting duties across critical infrastructure - exact obligations vary by sector and change, so confirm yours with counsel now, not mid-incident. Your plan should pre-stage everything the report will need: entity identifiers, contact channels, a report template, and the internal decision path for "is this reportable?" Industry-specific detail lives in our compliance series, starting with oil and gas SCADA compliance.

Tabletops: The Difference Between a Plan and a Document

A tabletop exercise is a structured walkthrough of a scenario with the real team making real decisions against a ticking clock - no systems touched, everything learned. Run at least two per year, rotating scenarios from your playbook list, and make them uncomfortable: start one at shift change, make the incident commander unreachable for the first thirty minutes, land a regulator deadline mid-exercise. Include operations and engineering every time; an IR tabletop with only IT people in the room rehearses the wrong incident. End each exercise with a findings list - every hesitation, missing phone number, and ambiguous authority - assign owners, and open the next exercise by checking the last list. Pair tabletops with technical drills from your DR program (restore tests, failover exercises) so the paper plan and the technical capability converge.

Key takeaway: OT incident response is a safety discipline wearing a cybersecurity jacket: safety first, availability second, confidentiality third. Build it as named roles with pre-authorized decisions, short scenario playbooks, assured detection and alert delivery, immutable evidence, pre-staged regulatory reporting, and twice-yearly tabletops that include the control room. And make your platform carry its share - session revocation, audit exports, SIEM delivery, and durable alerting are response tools you should see working, in a demo, before you ever need them. The platform side of that list is documented on the Merobix security architecture page.

Frequently Asked Questions

How is OT incident response different from IT incident response?

The priorities invert. IT incident response optimizes for confidentiality - isolate the machine, wipe it, restore from backup. In OT, availability and safety come first: you cannot simply power off a controller that is running a compressor station or a water treatment train, and containment actions themselves can create process hazards. OT response therefore requires process engineers and operations leadership in the incident team, containment options pre-approved through management of change, fallback-to-manual procedures, and playbooks written around process impact rather than data loss. The lifecycle phases are the same; the decision criteria inside them are not.

What should an OT incident response plan include?

Six things: severity definitions tied to process and safety impact, not just data; a named incident team with an on-call rotation covering operations, engineering, security, and leadership, plus decision authority for actions like isolating a site or falling back to manual; scenario playbooks for the incidents you actually expect - ransomware, compromised credentials, unauthorized control activity, telemetry anomalies, and vendor-side incidents; communication and regulatory reporting procedures with time limits, such as TSA-directed reporting to CISA for covered pipeline operators; evidence-preservation steps; and recovery criteria defining how you verify the system is clean before reconnecting. Keep printed copies - the plan must survive the incident it describes.

Who is responsible for incident response in cloud SCADA - vendor or operator?

Both, along the shared-responsibility line. The vendor detects and responds to platform-level incidents - infrastructure compromise, cross-tenant issues, platform vulnerabilities - and owes you timely notification, preserved evidence, and status updates, with commitments written into the contract. The operator owns response for their own users, credentials, sites, and field equipment: disabling compromised accounts, deciding process fallbacks, and meeting their own regulatory reporting duties. Your plan should include vendor-related scenarios and name the vendor contact path, and vendor capabilities like session revocation, audit exports, and SIEM delivery become your response tools during an operator-side incident.

How often should you run OT incident response tabletop exercises?

At least twice a year, with different scenarios each time - for example a ransomware outbreak reaching the OT boundary in spring, and a compromised operator credential making unauthorized setpoint changes in fall. Include operations and process engineering, not just IT security, and inject realistic complications: a decision-maker unreachable, alarms arriving at 2 AM, a regulator deadline mid-incident. Track findings as action items with owners and revisit them at the next exercise. Regulated sectors may face specific expectations - water utilities covered by AWIA must maintain emergency response plans, and pipeline operators under TSA directives are expected to exercise their cybersecurity incident response plans.

What evidence should be preserved during a SCADA security incident?

Preserve the audit trail (logins, failed attempts, configuration changes, setpoint writes and acknowledgements), security alerts and detection records, network logs at the OT boundary, gateway and device logs, affected user account details and session records, and timestamps for every response action you take. Tamper-evident audit records are what make this evidence defensible - Merobix maintains immutable, chained audit records with verification, so investigators can prove the log was not altered after the fact. Export and snapshot early, keep a written incident timeline from the first minute, and involve legal counsel and insurers before wiping or rebuilding anything they may need.

Sources & Further Reading

Response Tools You Can Rehearse

Session revocation, account lockdown, immutable audit trails, and assured alert delivery to your SIEM - walk your compromised-credential playbook against the live platform before you ever need to.

Request a Demo → See Our Security Architecture
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →