RISK: Auditing Industrial Control Systems for Too-Late-to-Recover Vulnerabilities

arXiv:2609.38528 · cs.CR · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "RISK: Auditing Industrial Control Systems for Too-Late-to-Recover Vulnerabilities".

Elias: The security of industrial control systems (ICS) requires attention to recovery after detection, as existing efforts often focus only on detection.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, to wrap up the discussion on "RISK: Auditing Industrial Control Systems for Too-Late-to-Recover Vulnerabilities," the paper effectively introduced a systematic way to model and test how attacks can specifically target and deplete the recovery margin in industrial control systems.

Elias: I think what stands out about this work is its methodical approach—developing RISK to holistically analyze PLC logic, detection, and recovery procedures to generate concrete TLTR attack scenarios.

Priya: And from a data perspective, the validation across various testbeds and a real-world plant shows that these scenarios are not just theoretical problems but are statistically prevalent in operational environments, with seventy-six percent of confirmed cases falling under this too-late-to-recover category.

Nadia: Exactly, and the main implication for the world is that we need to move beyond just detecting an attack and start rigorously auditing whether the system can actually recover safely once it's detected; this paper provides a concrete tool for that kind of resilience testing.

Elias: The authors demonstrate that existing ICS vetting tools often miss these specific scenarios because they aren't sensitive to the temporal dynamics of recovery failure, which points toward a necessary evolution in security tooling.

Priya: It’s about shifting the focus from simply finding a breach to ensuring that when a breach happens, the system retains enough operational slack to safely return to its intended state.

Nadia: That’s the core message: understanding how an attack drains recovery margins is essential for building truly resilient industrial control systems, and RISK is a framework designed specifically for that kind of deep audit.

Conclusion: Nadia: So, we've seen how RISK systematically models how attacks can drain an ICS's recovery margin before detection, but what does that title actually mean in practice?

Elias: I think the title is spot on because it focuses squarely on that 'too-late-to-recover' problem, which sounds like a very specific type of failure we see in operational systems.

Priya: From my side, I'm focused on what this means for the actual data; it suggests that standard detection methods might be insufficient if they don't also model the recovery timeline accurately.

Nadia: Exactly; it’s not just about *if* you get an alert, but whether you can actually fix things after the alert without causing more damage, which is a really tangible risk.

Elias: And regarding the authors, I'm looking at their methodology to see if they've made any assumptions in their modeling that could be broken by a clever cryptographer.

Priya: I'm curious about what kind of data they used—did it capture enough detail on those recovery procedures to make these predictions really robust?

Nadia: That’s the million-dollar question; if this framework is accurate, it means we can start testing systems for resilience against attacks that aim to destroy their ability to recover safely.

Elias: I see the implication for security tools being forced to evolve beyond just finding immediate threats toward predicting long-term operational failures based on recovery constraints.

Priya: So, the real impact is shifting the focus from attack surface reduction to system survivability under duress, which is a huge change for industrial safety standards.

Nadia: Right, and if we can't find these vulnerabilities cheaply or easily, it might mean that securing critical infrastructure becomes much more complex than we currently imagine.

Syed Ghazanfar Abbas, Gang Wang, Dongyan Xu

Purdue University · University of Illinois Urbana-Champaign

cs.CR

Submitted: 2026-09-29

Updated: 2026-09-29

Comments: 17 pages

Code: https://github.com/Fortiphyd/GRFICSv2

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: The security of industrial control systems (ICS) requires attention to recovery after detection, as existing efforts often focus only on detection.

Key concepts

Too-Late-to-Recover (TLTR) Vulnerability
This vulnerability exists when an attack manipulates a process so effectively that the available safety margin is completely exhausted before detection. Once detected, the system's recovery procedures cannot successfully return the plant to its safe state because there is no remaining buffer to handle the situation.
RISK Framework
RISK is an automated system that holistically models an ICS by combining PLC programs, detection rules, and recovery plans. It uses a large language model agent to suggest process manipulations and validates these attacks using a virtual PLC to confirm if they lead to infeasible or unsafe recoveries.
Recovery Margin Quantification
The framework builds a Multi-layer Constraint Model (MCM) from system policies. This model allows RISK to calculate key metrics like the minimum remaining slack, the time it will take to reach a critical boundary, and how fast recovery actions can restore that margin.
Recovery-Infeasible/Unsafe
These are two specific failure modes identified by RISK. Recovery is 'infeasible' if the deployed procedure cannot meet its constraints after detection. It is 'unsafe' if executing the recovery actions causes new, cascading violations or hazards before the recovery goal is achieved.

Terminology

Summary

The security of industrial control systems (ICS) requires attention to recovery after detection, as existing efforts often focus only on detection. This paper addresses this gap by defining and auditing too-late-to-recover (TLTR) vulnerabilities, which allow an attack to drain the available recovery margin before detection, rendering subsequent recovery procedures unable to restore the system safely.

The gist

RISK is an automated framework that discovers and validates possible TLTR attack scenarios by holistically modeling and analyzing the PLC control logic, attack detection policies, recovery procedures, and operational behaviors of an ICS to generate TLTR attack scenarios with concrete attack parameters.

How it works

RISK operates in four sequential steps to audit a subject ICS for TLTR vulnerabilities:

  1. It first extracts process-specific constraints from PLC programs as well as detection policies, recovery procedures, and, where available, protection policies, organizing them into a unified multi-layer model that characterizes how the process is governed across the control, detection, and supervisory layers.

  2. Next, it analyzes operational traces to capture how the process evolves relative to constraints during normal operation, specifically quantifying how quickly the safety margin shrinks and how effective recovery actions are.

  3. It then uses this model to explore a space of feasible process manipulations that remain within control and safety limits, yet progressively exhaust recovery margins prior to detection. For this exploration, it leverages a large language model (LLM) agent that suggests candidate TLTR attacks by specifying which process variables to manipulate, when to manipulate them, and by how much, with these attacks evaluated using a virtual PLC (vPLC), a software replica of physical controllers.

  4. Finally, it analyzes the resulting execution traces to confirm TLTR attacks, classifying them as recovery-infeasible or recovery-unsafe.

TLTR Vulnerability Formulation

The paper defines an ICS exhibiting a TLTR vulnerability when "adversarial manipulations can drain the available physical or temporal recovery margin before detection, such that, once detected, the deployed recovery procedures can no longer restore the plant to its intended safe operating state due to the insufficient margin. This vulnerability is characterized by attacks that only needs to drive the process to the point, beyond which the deployed recovery procedure will not have sufficient margin to succeed." TLTR vulnerabilities lead to two specific failure modes:

  1. Recovery can become infeasible, where the deployed recovery procedure can no longer restore the process to its intended safe operating state within the applicable recovery constraints.

  2. Recovery can become unsafe, where executing recovery actions triggers cascading violations or hazards, leading to multiple failures.

Constraint-Guided Trace Analysis

RISK builds a Multi-layer Constraint Model (MCM) and uses it for constraint-guided trace analysis. The MCM is a directed graph where nodes represent constraints annotated with their triggering condition, enforcement action, and regulated variables. Edges capture relationships between constraints, such as implication between activation conditions, shared regulated variables, and interactions across control, detection, and supervisory logic. This model allows RISK to quantify the available recovery margin by deriving three quantities from trace segments:

(Di)

The minimum remaining slack to the relevant recovery-target.

(Ti)

The minimum estimated time to reach that boundary while the process is approaching it.

(Ri)

The maximum observed rate at which slack is restored within the segment.

TLTR Scenario Discovery and Evaluation

RISK discovers candidate TLTR scenarios by exploring admissible process manipulations that reduce the recovery margin at detection relative to the constraint-centric patterns identified above. The LLM agent proposes attack scripts based on these low-margin regions, and these scripts are validated against constraint-derived bounds before being executed in a vPLC environment. The final classification uses Algorithm 1 to determine recoverability:

(Recovery-Infeasible)

If a target constraint is violated after detection and before recovery completes, or if the recovery condition is not reached within the specified interval.

(Recovery-Unsafe)

If execution of the deployed recovery logic is associated with a secondary constraint violation before the recovery condition is reached.

Evaluation and Findings

RISK was evaluated on three ICS testbeds (Fischertechnik manufacturing plant, water treatment testbed, and chemical processing testbed) as well as a real-world fertilizer plant. Across these three testbeds, a total of 392 TLTR attacks are generated and confirmed, with 76% resulting in TLTR conditions. Among these cases, 222 are recovery-infeasible scenarios and 170 are recovery-unsafe scenarios. Furthermore, the paper demonstrates that existing ICS vetting tools like AttackLLM and GENICS only identify a small fraction of TLTR attacks (7% and 8%, respectively), showing that "existing solutions are not TLTR-sensitive.

Improvements for AI systems

Here are specific improvements to AI systems based on the methodology and findings of the RISK framework described in the paper, along with what these improved systems can achieve:


  1. Develop a new class of Recoverability-Aware or Resilience-Optimized AI Agents for critical infrastructure control.

  2. The improved system will be able to proactively identify and mitigate Too-Late-To-Recover (TLTR) vulnerabilities in PLC logic, detection policies, and recovery procedures.

  3. Specifically, the system can perform the following functions:

  4. Identify TLTR Scenarios: The AI agent will systematically search for adversarial manipulations (within physical constraints) that drain the available physical or temporal recovery margin before a detection event occurs.

  5. Characterize Recovery Failure Modes: The system will not only detect if an attack is successful but will explicitly classify the resulting state as either Recovery-Infeasible (where the recovery procedure cannot reach its target constraint within its interval) or Recovery-Unsafe (where executing the recovery procedure triggers cascading secondary constraint violations).

  6. Enhance Detection Thresholds: The system can recommend specific, data-driven adjustments to alarm thresholds and detection policies based on predicted TLTR susceptibility, ensuring that alarms trigger early enough for recovery to succeed.

  7. Optimize Recovery Procedures: By simulating various attack trajectories against the deployed recovery logic, the system can suggest modifications to existing recovery procedures—such as adjusting actuation rates or timing—to ensure they are robust against margin exhaustion attacks.

  8. Provide Proactive Security Auditing: The AI system will serve as a continuous, automated auditor that integrates PLC code analysis (static/dynamic) with operational trace monitoring to continuously assess the real-time recoverability margin of the ICS.

  9. Compare and Contrast Attack Generation Methods: The AI system can be used to benchmark existing attack generation tools (like AttackLLM or GENICS) against its own method, providing quantifiable evidence on whether a tool is sensitive to post-detection recoverability constraints, thereby guiding better security research and tool development.

Related papers