SCADAPENTEST

GD-04 · Permit

Rules of engagement for an OT security test

Updated 10 min read

A plant already knows how to let outsiders do dangerous work: it issues a permit. It names the authority, defines the boundary, requires isolation, sets the conditions for stopping and demands a return-to-service check. Security testing on operational technology should borrow that structure wholesale, because the failure modes are the same and the vocabulary is already familiar to everyone who has to sign.

Why an IT scope document is not enough

A typical IT rules-of-engagement document names IP ranges, dates, a contact for out-of-hours, and a statement that the client accepts the risk of disruption. Every part of that is inadequate for a plant. IP ranges do not describe a fieldbus. A date range is not a window. And no security contact can accept, on behalf of an operator, the risk of an unplanned process trip: the person who can accept that risk is the one responsible for the process and the safety case.

The published record is unambiguous about what happens when this is done casually. NIST records a natural gas utility whose consultancy “carelessly ventured into a part of the network that was directly connected to the SCADA system”, after which “the penetration test locked up the SCADA system, and the utility was not able to send gas through its pipelines for four hours”. Nobody had intended to test SCADA at all. The scope was corporate IT.

The eight controls that must exist before the first packet

1. A named process authority

One person, named in the document, with the standing to stop everything instantly and no involvement in running the test. On a shift-operated plant this means a named role per shift, and the name of whoever is actually on duty during the window. The authority is not the IT manager and not the project sponsor. It is the person who would call a shutdown for any other reason.

2. An abort phrase on a live channel

A spoken stop word, agreed in advance, on an open voice channel that stays open for the duration of any intrusive activity. Email and ticketing systems are not stop controls. Define what happens after the phrase is used: testing halts, tooling is disconnected, the tester states the last three actions taken, and nothing resumes without the authority saying so. Practise it once at the start of the window so that everyone has heard it work.

3. An exclusion list in writing

Devices, addresses, protocols and functions that are out of scope, agreed before the engagement rather than argued during it. On a process plant the safety instrumented system is always on this list. So is anything with a single point of failure the operator cannot restore within the shift, anything mid-way through a validated batch, and any device the vendor has flagged as fragile.

4. Backups verified restorable

Controller projects, engineering workstation images, HMI projects and device configuration. Not “a backup exists”, but a restore performed once, to prove the backup is real and to measure how long it takes. That measurement changes decisions: a four-hour restore makes a device an outage-only target no matter how interesting it looks.

5. A rollback plan owned by the vendor

Who restores what, in what order, and how quickly with the vendor on a normal support response rather than a heroic one. This is also the moment to check the support contract. NIST notes that in some cases “third-party security solutions are not allowed due to OT vendor licensing and service agreements, and service support can be lost if third-party applications are installed without vendor acknowledgement or approval”. Written notice to the vendor costs an email and protects the contract that keeps the line running.

6. Logging aligned to the process historian

Every test action time-stamped against a synchronised clock so that the log can be laid alongside the historian trend. When a pump behaves oddly two hours into the window, the question “was that us?” has to be answerable in minutes. Without aligned logging it becomes a post-incident investigation, and the testing stops either way.

7. A defined window and the right audience

Start and stop times published to the shift that will be on duty, not only to the project team that commissioned the work. An operator who sees an unexplained alarm and has not been told there is a test running will respond as if it is real, which is correct behaviour and an avoidable cost.

8. Testers who understand the process

People who can read a piping and instrumentation diagram and a ladder diagram, and who know what a fail-safe action looks like. NIST removed the independent-tester requirement from its high baseline for exactly this reason: “Specific expertise is necessary to conduct effective penetration testing on OT systems, and it may not be feasible to identify independent personnel with the appropriate skillset or knowledge to perform penetration testing on an OT environment.” Independence is worth less than competence when the failure mode is physical.

Abort criteria, written as conditions rather than judgement

The abort phrase is the mechanism. The criteria are what triggers it, and they should be objective enough that a tester with no process knowledge would call it correctly.

  • Any device stops responding to the operator, whether or not it is in scope.
  • Any alarm the operator did not expect, including nuisance alarms, until it is explained.
  • Any change in a process value that was not requested by the control system.
  • Any loss of communication to a remote site or an outstation.
  • Any deviation from the agreed device list, including a device discovered mid-test.
  • Any doubt at all from the process authority, with no requirement to justify it at the time.

Windows, outages and the difference between them

Operators use “window” loosely, and the distinction matters more than any technical decision in the plan.

What each level of tolerance permits
ToleranceWhat it permitsWhat it does not
NonePassive capture, documentation and ruleset review, offline project and logic analysis, engineering workstation and historian review, full testing of the edgeAnything that writes or probes below the demilitarised zone
A booked short windowSupervised, device-by-device active checks on the supervisory band, starting with safe active scanning and stopping at the first unexpected responseField-device interrogation and exploitation: recovery outlasts the window
A planned shutdownActive testing of controllers against a named list, with backups taken first and the vendor reachableAnything involving the safety instrumented system, still and always
Method risk classification follows NIST SP 800-82r3, control CA-8 and appendices C.3.4 and E.2.3.

A short window is not a small shutdown. If restoring a hung controller takes ninety minutes and the window is thirty, the controller is not in scope for that window regardless of how carefully anyone plans. The arithmetic, not the ambition, decides.

Where the process cannot absorb anything, the compensating control NIST names is “a replicated, virtualized, or simulated system”. Treat the replica as part of the rules of engagement rather than as a fallback: state which findings will be produced there, on which firmware version, and how they will be evidenced as applicable to the plant.

Environment-specific clauses

Four kinds of estate need four extra paragraphs, and they are the paragraphs that get missed.

  • Continuous process plant. Exclude the safety instrumented system in writing and treat it separately through the functional-safety lifecycle. Name the batch or campaign state during which no testing occurs at all.
  • Distributed utility network. Document, per remote site, what the outstation does when it loses communications and for how long it tolerates the loss. An unattended site that fails to a safe state on a lost link will do so during a test as readily as during a fault.
  • Building management. Identify and exclude the life-safety branches: smoke control, stairwell pressurisation, fire dampers and access egress are not comfort systems, and they are frequently on the same controller family as the air handling.
  • Discrete manufacturing line. Agree a maximum acceptable line stop in minutes, who is authorised to call it, and what it costs. This turns an argument about risk appetite into a number that the plant manager can approve.

Two access paths also deserve their own clauses wherever they exist. A vendor support tunnel needs written notice and a decision about whether it is tested with the vendor present. And radio work is receive-only on site: transmitting into a live telemetry band is a different activity from surveying it, and belongs in a screened environment.

The deliverable most OT reports leave out

An OT report that lists only findings is incomplete, and in front of a board it is misleading. The plant needs to see, in the same document, an explicit statement of what could not be tested safely and what it would take to test it.

“We found nothing in the control network” and “we could not look in the control network” describe completely different situations, and a report that silently merges them sells assurance that was never produced. The untested residue is not an admission of failure. It is the input to the next outage plan, the justification for a replica, and the honest boundary of the current risk picture.

What the residue section should contain

  • Each method that was excluded, and the specific constraint that excluded it.
  • What that method would have been expected to find, stated as a class of issue rather than a guess.
  • What it would take to run it: a window of a given length, a replica of given fidelity, or a vendor engagement.
  • The interim compensating measure, if any, and who owns it.
  • A recommended point in the maintenance calendar at which it should be revisited.

A one-page checklist

  • Named process authority per shift, with contact details, and the statement that they may stop the test without justification.
  • Abort phrase, the open voice channel it is used on, and the defined post-abort procedure.
  • Scope boundary stated as an enforcement point, with the party who verifies it in the logs.
  • Exclusion list: devices, addresses, protocols, functions, and the safety instrumented system explicitly.
  • Backups taken and one restore proven, with the measured restore time recorded in the document.
  • Rollback plan and vendor notice, including who holds the support contract and what it forbids.
  • Time synchronisation and logging aligned to the historian.
  • Window start and stop, published to the shift on duty.
  • Environment-specific clauses for the estate type, and clauses for vendor tunnels and radio work.
  • Agreed report structure, including the untested-residue section, before the testing begins.

For which methods should appear on the permitted side of the boundary, see what can safely be tested on a live OT network. For where the obligation to assess comes from, see NIS2 and OT security testing, and for the framework the remediation will be measured against, which part of IEC 62443 applies to you. The approach selector generates the control list for a specific environment.

Sources

  1. NIST SP 800-82r3, Guide to Operational Technology (OT) Security National Institute of Standards and Technology · 2023 Control overlay CA-8 and the rationale for removing CA-8(1); section 5.3.1 on fail-to-a-known-state design; section 2.5 on vendor support constraints; appendix C.3.4 for the pipeline case.
  2. Defending OT Operations Against Ongoing Pro-Russia Hacktivist Activity CISA with FBI, NSA, EPA, DOE, USDA, FDA, MS-ISAC, CCCS and NCSC-UK · 2024 Recommends practising and maintaining the ability to operate systems manually, and backing up HMI engineering logic, configuration and firmware to enable fast recovery.
  3. DarkSide Ransomware: Best Practices for Preventing Business Disruption from Ransomware Attacks (AA21-131A) CISA and FBI · 2021 Source for the recommendation to identify IT and OT interdependencies and to regularly test manual controls so critical functions can keep running if OT is taken offline.

FAQ-950 · Questions

Related questions

Who should sign the rules of engagement?
At minimum the process authority, the person accountable for the security programme, and the testing provider. On a regulated site the safety function should see the exclusion list, and where a vendor holds the support contract or the remote-access path they should acknowledge the notice in writing. A document signed only by IT and the provider has not obtained permission from anyone who can grant it.
What if we discover a device that is not on the list mid-test?
Stop and add it, or stop and exclude it. Discovering an undocumented device is a finding in its own right and usually a valuable one, but testing it on the strength of having just found it is exactly the sequence that produced the four-hour pipeline outage in the NIST case file. The rule to write is that anything not on the list is out of scope until the authority puts it on the list.
Can the testing provider accept the risk on our behalf?
No, and be suspicious of one that offers to. A provider can accept professional liability for negligence, but nobody outside the operator can accept the consequences of a process trip: the production loss, the environmental release and the safety case are the operator’s. What a good provider does instead is make the risk explicit enough that the operator can decide, and refuse work where the constraint and the requested method are incompatible.
How long should the document be?
Short enough that the shift supervisor reads it. Two pages of controls, a device list as an annex and a one-page abort card taped next to the operator station will be used; a forty-page contract schedule will not. The permit-to-work culture on most plants already knows this, which is another reason to borrow the format.