GD-04 · Permit
Rules of engagement for an OT security test
A plant already knows how to let outsiders do dangerous work: it issues a permit. It names the authority, defines the boundary, requires isolation, sets the conditions for stopping and demands a return-to-service check. Security testing on operational technology should borrow that structure wholesale, because the failure modes are the same and the vocabulary is already familiar to everyone who has to sign.
Why an IT scope document is not enough
A typical IT rules-of-engagement document names IP ranges, dates, a contact for out-of-hours, and a statement that the client accepts the risk of disruption. Every part of that is inadequate for a plant. IP ranges do not describe a fieldbus. A date range is not a window. And no security contact can accept, on behalf of an operator, the risk of an unplanned process trip: the person who can accept that risk is the one responsible for the process and the safety case.
The published record is unambiguous about what happens when this is done casually. NIST records a natural gas utility whose consultancy “carelessly ventured into a part of the network that was directly connected to the SCADA system”, after which “the penetration test locked up the SCADA system, and the utility was not able to send gas through its pipelines for four hours”. Nobody had intended to test SCADA at all. The scope was corporate IT.
The eight controls that must exist before the first packet
1. A named process authority
One person, named in the document, with the standing to stop everything instantly and no involvement in running the test. On a shift-operated plant this means a named role per shift, and the name of whoever is actually on duty during the window. The authority is not the IT manager and not the project sponsor. It is the person who would call a shutdown for any other reason.
2. An abort phrase on a live channel
A spoken stop word, agreed in advance, on an open voice channel that stays open for the duration of any intrusive activity. Email and ticketing systems are not stop controls. Define what happens after the phrase is used: testing halts, tooling is disconnected, the tester states the last three actions taken, and nothing resumes without the authority saying so. Practise it once at the start of the window so that everyone has heard it work.
3. An exclusion list in writing
Devices, addresses, protocols and functions that are out of scope, agreed before the engagement rather than argued during it. On a process plant the safety instrumented system is always on this list. So is anything with a single point of failure the operator cannot restore within the shift, anything mid-way through a validated batch, and any device the vendor has flagged as fragile.
4. Backups verified restorable
Controller projects, engineering workstation images, HMI projects and device configuration. Not “a backup exists”, but a restore performed once, to prove the backup is real and to measure how long it takes. That measurement changes decisions: a four-hour restore makes a device an outage-only target no matter how interesting it looks.
5. A rollback plan owned by the vendor
Who restores what, in what order, and how quickly with the vendor on a normal support response rather than a heroic one. This is also the moment to check the support contract. NIST notes that in some cases “third-party security solutions are not allowed due to OT vendor licensing and service agreements, and service support can be lost if third-party applications are installed without vendor acknowledgement or approval”. Written notice to the vendor costs an email and protects the contract that keeps the line running.
6. Logging aligned to the process historian
Every test action time-stamped against a synchronised clock so that the log can be laid alongside the historian trend. When a pump behaves oddly two hours into the window, the question “was that us?” has to be answerable in minutes. Without aligned logging it becomes a post-incident investigation, and the testing stops either way.
7. A defined window and the right audience
Start and stop times published to the shift that will be on duty, not only to the project team that commissioned the work. An operator who sees an unexplained alarm and has not been told there is a test running will respond as if it is real, which is correct behaviour and an avoidable cost.
8. Testers who understand the process
People who can read a piping and instrumentation diagram and a ladder diagram, and who know what a fail-safe action looks like. NIST removed the independent-tester requirement from its high baseline for exactly this reason: “Specific expertise is necessary to conduct effective penetration testing on OT systems, and it may not be feasible to identify independent personnel with the appropriate skillset or knowledge to perform penetration testing on an OT environment.” Independence is worth less than competence when the failure mode is physical.
Abort criteria, written as conditions rather than judgement
The abort phrase is the mechanism. The criteria are what triggers it, and they should be objective enough that a tester with no process knowledge would call it correctly.
- Any device stops responding to the operator, whether or not it is in scope.
- Any alarm the operator did not expect, including nuisance alarms, until it is explained.
- Any change in a process value that was not requested by the control system.
- Any loss of communication to a remote site or an outstation.
- Any deviation from the agreed device list, including a device discovered mid-test.
- Any doubt at all from the process authority, with no requirement to justify it at the time.
Windows, outages and the difference between them
Operators use “window” loosely, and the distinction matters more than any technical decision in the plan.
| Tolerance | What it permits | What it does not |
|---|---|---|
| None | Passive capture, documentation and ruleset review, offline project and logic analysis, engineering workstation and historian review, full testing of the edge | Anything that writes or probes below the demilitarised zone |
| A booked short window | Supervised, device-by-device active checks on the supervisory band, starting with safe active scanning and stopping at the first unexpected response | Field-device interrogation and exploitation: recovery outlasts the window |
| A planned shutdown | Active testing of controllers against a named list, with backups taken first and the vendor reachable | Anything involving the safety instrumented system, still and always |
A short window is not a small shutdown. If restoring a hung controller takes ninety minutes and the window is thirty, the controller is not in scope for that window regardless of how carefully anyone plans. The arithmetic, not the ambition, decides.
Where the process cannot absorb anything, the compensating control NIST names is “a replicated, virtualized, or simulated system”. Treat the replica as part of the rules of engagement rather than as a fallback: state which findings will be produced there, on which firmware version, and how they will be evidenced as applicable to the plant.
Environment-specific clauses
Four kinds of estate need four extra paragraphs, and they are the paragraphs that get missed.
- Continuous process plant. Exclude the safety instrumented system in writing and treat it separately through the functional-safety lifecycle. Name the batch or campaign state during which no testing occurs at all.
- Distributed utility network. Document, per remote site, what the outstation does when it loses communications and for how long it tolerates the loss. An unattended site that fails to a safe state on a lost link will do so during a test as readily as during a fault.
- Building management. Identify and exclude the life-safety branches: smoke control, stairwell pressurisation, fire dampers and access egress are not comfort systems, and they are frequently on the same controller family as the air handling.
- Discrete manufacturing line. Agree a maximum acceptable line stop in minutes, who is authorised to call it, and what it costs. This turns an argument about risk appetite into a number that the plant manager can approve.
Two access paths also deserve their own clauses wherever they exist. A vendor support tunnel needs written notice and a decision about whether it is tested with the vendor present. And radio work is receive-only on site: transmitting into a live telemetry band is a different activity from surveying it, and belongs in a screened environment.
The deliverable most OT reports leave out
An OT report that lists only findings is incomplete, and in front of a board it is misleading. The plant needs to see, in the same document, an explicit statement of what could not be tested safely and what it would take to test it.
“We found nothing in the control network” and “we could not look in the control network” describe completely different situations, and a report that silently merges them sells assurance that was never produced. The untested residue is not an admission of failure. It is the input to the next outage plan, the justification for a replica, and the honest boundary of the current risk picture.
What the residue section should contain
- Each method that was excluded, and the specific constraint that excluded it.
- What that method would have been expected to find, stated as a class of issue rather than a guess.
- What it would take to run it: a window of a given length, a replica of given fidelity, or a vendor engagement.
- The interim compensating measure, if any, and who owns it.
- A recommended point in the maintenance calendar at which it should be revisited.
A one-page checklist
- Named process authority per shift, with contact details, and the statement that they may stop the test without justification.
- Abort phrase, the open voice channel it is used on, and the defined post-abort procedure.
- Scope boundary stated as an enforcement point, with the party who verifies it in the logs.
- Exclusion list: devices, addresses, protocols, functions, and the safety instrumented system explicitly.
- Backups taken and one restore proven, with the measured restore time recorded in the document.
- Rollback plan and vendor notice, including who holds the support contract and what it forbids.
- Time synchronisation and logging aligned to the historian.
- Window start and stop, published to the shift on duty.
- Environment-specific clauses for the estate type, and clauses for vendor tunnels and radio work.
- Agreed report structure, including the untested-residue section, before the testing begins.
For which methods should appear on the permitted side of the boundary, see what can safely be tested on a live OT network. For where the obligation to assess comes from, see NIS2 and OT security testing, and for the framework the remediation will be measured against, which part of IEC 62443 applies to you. The approach selector generates the control list for a specific environment.
Sources
- NIST SP 800-82r3, Guide to Operational Technology (OT) Security Control overlay CA-8 and the rationale for removing CA-8(1); section 5.3.1 on fail-to-a-known-state design; section 2.5 on vendor support constraints; appendix C.3.4 for the pipeline case.
- Defending OT Operations Against Ongoing Pro-Russia Hacktivist Activity Recommends practising and maintaining the ability to operate systems manually, and backing up HMI engineering logic, configuration and firmware to enable fast recovery.
- DarkSide Ransomware: Best Practices for Preventing Business Disruption from Ransomware Attacks (AA21-131A) Source for the recommendation to identify IT and OT interdependencies and to regularly test manual controls so critical functions can keep running if OT is taken offline.
FAQ-950 · Questions
Related questions
Who should sign the rules of engagement?
What if we discover a device that is not on the list mid-test?
Can the testing provider accept the risk on our behalf?
How long should the document be?
DOC-960 · Keep reading
More guides
- What can safely be tested on a live OT network The methods that run while the process runs, the ones that never should, and the honest limits of each. Open guide
- NIS2 and OT security testing: what Article 21 actually requires Which industrial sectors are in scope, what the directive really says about assessment, and why nobody has to run an annual pentest of the control network. Open guide
- IEC 62443: which part applies to you The series has four groups and three audiences. Knowing which one you are removes most of the reading. Open guide