Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

arXiv:2608.11498 · cs.CV, cs.ET, cs.LG · Submitted 2026-08-11 · Read on arXiv

Aditya Humnabadkar, Huaizhong Zhang, Ardhendu Behera

Edge Hill University

cs.CV, cs.ET, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-13

Comments: Accepted manuscript: Workshop on Emerging Behaviors in Embodied AI for Achieving Robust Autonomy as part of European Conference on Computer Vision (ECCV) 2026

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: The paper introduces Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), for threat-aware control in safety-critical driving scenarios.

Terminology

Summary

The paper introduces Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), for threat-aware control in safety-critical driving scenarios. The framework addresses the question: what does an ego policy learn when its training distribution is structured by language-specified intent, but the language prompt and semantic actor roles are unavailable during policy execution?

The pipeline converts natural-language descriptions into schema-valid traffic configurations and executable interaction programs (cut-in, sudden-braking, overtaking, and tailgating) using the Gemini API with schema-constrained few-shot prompting. The parser extracts lane index, longitudinal position, speed, behaviour style (cautious, normal, aggressive), actor role (ego, adversarial, background), and interaction type. Crucially, the ego-policy observation excludes both the original prompt and privileged actor-role labels—language structures the policy's training experience rather than serving as a direct policy input.

ERQ-Net couples dynamic traffic-graph reasoning with value-based action selection. At each timestep, the traffic scene is represented as a graph with k-NN connectivity (k=5, d max=100m), where each vehicle has observation-only features: normalized positions, velocities, and continuous heading representation. The architecture uses two graph-attention layers with four heads, producing a 128-dimensional ego-centric representation, followed by a two-layer MLP Q-value head (128→64→5, ReLU) over five manoeuvres: LEFT, IDLE, RIGHT, FASTER, SLOWER. The graph encoder and Q-value head are jointly optimized via a temporal-difference objective, so attention is learned by its impact on ego action selection rather than as a generic scene representation.

Scenario Instantiation Reliability: The parser produces valid executable configurations for 94.2% of prompts on the first attempt and 98.5% after one retry, with 1.5% using template fallback. Mean API latency is 1.2 seconds.

Effect of Language-Structured Training: Across 2,500 safety-critical scenarios (2,000 training, 500 testing, balanced by traffic density), language-structured training improves test success from 49–52% (random control) to 55–58%, and increases adversary-focused attention from 1.2× to 2.1×. Since prompts and role labels are hidden from ERQ-Net, "this preferential attention is induced by coherent interaction patterns in the training distribution rather than by inference-time supervision, supporting the emergence of threat-aware representations from behaviour alone."

Relational-Encoder Ablation: Flattening vehicle features into an MLP yields 44–47% success; uniform graph aggregation (GCN) improves to 48–51%; single-head GAT reaches 51–54%; transformer set encoder achieves 52–55%; ERQ-Net achieves 55–58%, showing a modest but consistent three-point advantage over the strongest alternative.

Recognition–Control Gap and Policy Collapse: Despite the emergent threat awareness, trained policies (55–58%) remain comparable to always selecting SLOWER (57%), while clearly exceeding random action (31–34%) and untrained networks (45–49%). The union of 12 diverse policies achieves 76% success, giving a recognition–control gap of ∆RC = 76−58 = 18 percentage points. The authors formalise this as: language-structured training enables ERQ-Net to identify the relevant threat, but the learned Q-function does not consistently convert that representation into scenario-dependent control. Reward reweighting (reduced speed reward) and margin shaping do not eliminate the attractor, instead reducing success to 48–53% and 48–52%, respectively.

Decision-Frequency Ablation: Increasing the decision frequency from 1 Hz to 2 Hz improves success to 59–62%, to 5 Hz achieves 64–68%, and 10 Hz reaches 66–69%. The recognition–control gap shrinks from 18 to 7 percentage points, indicating that temporal resolution is important but not the only cause of policy collapse.

Scenario Quality: Realism R = 0.94 (94% of timesteps satisfy kinematic and road-validity constraints), criticality C = 0.75 (75% of scenarios induce collision against a non-reactive reference ego), semantic fidelity F = 0.86, composite Q = 0.85. Shifting 0.1 weight between realism and criticality changes the composite only from 0.83 to 0.87.

State-Interface Transfer to CARLA: Without fine-tuning, CARLA achieves 72.6% average success (2.7 percentage points below HighwayEnv), with 99.0% mean trajectory consistency across clear, heavy-rain, and dense-fog conditions. The authors emphasize this evaluates state-interface transfer, not perception-level domain adaptation or real-world sim-to-real generalisation.

The paper's contributions are: (1) introducing ERQ-Net for joint learning of actor relevance, ego-state representations, and action values; (2) developing a controlled language-structured training protocol where semantic intent shapes interactions while prompts and role labels remain hidden; (3) showing emergent threat-focused attention and formalising the recognition–control gap; (4) defining a reproducible policy-collapse attractor via constant-action and policy-portfolio analyses; and (5) evaluating scenario realism, criticality, semantic fidelity, and zero-shot state-interface transfer to CARLA.

The study focuses on four types of controlled highway interaction with simulator-derived kinematic observations and a discrete 1 Hz action space. It does not cover perception uncertainty, complex urban interactions, or continuous vehicle control. Future work will extend ERQ-Net to multimodal perception, higher-frequency and continuous control, and more compositional multi-actor scenarios.

Improvements for AI systems

Improvements to AI Systems:

  1. Threat-Aware Attention Without Inference-Time Supervision: Enhance AI systems with a training protocol where natural-language descriptions structure the training distribution (e.g., specifying adversarial vs. background roles) but are hidden during execution. This forces the model to learn threat-relevant representations purely from interaction patterns, improving attention to critical entities (2.1× adversary focus) without needing privileged labels at runtime.

  2. Recognition–Control Gap Mitigation via Temporal Resolution: Increase decision frequency from 1 Hz to 10 Hz in sequential decision-making systems. This reduces the recognition–control gap from 18 to 7 percentage points, enabling the model to convert identified threats into effective actions. The improved system can achieve 66–69% success in safety-critical scenarios, up from 55–58% at lower frequencies.

  3. Relational Graph Encoders for Action-Conditioned Representation Learning: Replace flat or uniform aggregation encoders with multi-head graph-attention networks (e.g., 4 heads, k-NN connectivity) where attention is optimized jointly with the action-value objective. This yields a consistent 3-point success advantage over transformer set encoders and 7–11 points over MLP baselines, improving the system’s ability to reason about dynamic multi-agent interactions.

  4. Language-Schema-Constrained Scenario Generation: Use schema-constrained few-shot prompting with an LLM API (e.g., Gemini) to convert free-text intents into executable traffic configurations (cut-in, sudden-braking, etc.) with 94–98% validity. This enables automated generation of diverse, realistic, and critical training scenarios, improving generalization and reducing manual scenario engineering.

  5. Zero-Shot State-Interface Transfer: Design state-based policies that transfer across simulators (e.g., HighwayEnv to CARLA) without fine-tuning, achieving 72.6% success and 99.0% trajectory consistency under varied weather conditions. This improves deployment readiness by decoupling policy learning from simulator-specific dynamics.

  6. Policy-Collapse Detection and Portfolio Analysis: Implement diagnostic tools that compare trained policies against constant-action baselines and policy portfolios to identify attractors (e.g., always SLOWER). This enables early detection of degenerate solutions and guides architectural or training changes (e.g., frequency adjustments) before deployment.

What the Improved AI System Can Do:

  • Operate in safety-critical driving with emergent threat awareness, focusing on adversarial actors (e.g., cut-in vehicles) without explicit labels, improving success rates from 50% to 58% at 1 Hz and 69% at 10 Hz.

  • Generate and validate thousands of realistic, critical scenarios from natural-language descriptions, enabling scalable training and testing across diverse interaction types.

  • Transfer learned control policies across simulation environments with minimal performance loss, robust to weather and kinematic variations.

  • Avoid policy collapse by detecting when the Q-function defaults to a single action, and adapt decision frequency or reward shaping to maintain scenario-dependent control.

  • Reason over dynamic traffic graphs with learned attention, outperforming generic encoders in multi-agent coordination tasks.

Abstract

Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured data leads to truly adaptive control policies. We propose Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), which jointly learns inter-vehicle relevance and action values from dynamic traffic graphs. Language descriptions define surrounding-vehicle behaviours during training, while prompts and semantic actor roles are hidden from the policy. ERQ-Net must therefore infer threat relevance solely from observable kinematics and interactions. Across 2,500 safety-critical scenarios, language-structured training improves test success from 49-52% to 55-58% and increases adversary-focused attention from 1.2x to 2.1x, demonstrating emergent threat awareness. However, this representational gain does not consistently translate into adaptive control: trained policies perform similarly to the best constant action, while a portfolio of simple policies solves 76% of scenarios. We formalise this discrepancy as a recognition-control gap and show that reward reweighting and margin shaping do not eliminate the resulting policy collapse. Evaluations of realism, criticality, semantic accuracy, and transfer of state-interface representations to CARLA further highlight both the strengths and the constraints of language-structured relational policy learning in safety-critical driving scenarios.

Sources

Related papers