Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards

arXiv:2608.11451 · cs.RO, cs.AI, cs.SY, eess.SY · Submitted 2026-08-11 · Read on arXiv

Simón Patiño Idarraga, Erick Silva, Rehana Yasmin, Ali Shoker

Universidad de Antioquia · King Abdullah University of Science and Technology

cs.RO, cs.AI, cs.SY, eess.SY

Submitted: 2026-08-11

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces a neuro-symbolic safety guard for end-to-end autonomous driving.

Terminology

Summary

The paper introduces a neuro-symbolic safety guard for end-to-end autonomous driving. The guard is a lightweight module that attaches to the final command interface of an already-trained agent. Immediately before a command reaches the vehicle, it checks the command against explicit safety rules and, only when necessary, replaces it with the nearest safe alternative. Each intervention is directly executable and traceable to the rule that triggered it, while the guard itself requires no retraining and adds no learned component.

The core problem addressed is that end-to-end driving agents "can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guarantee safe driving, leaving their decision-making process opaque and safety constraints unenforced. The paper argues that the safety requirement violated in these cases is not hidden or subtle: it can be stated as a simple physical condition on how the vehicle may move."

The guard operates in three stages: it reads the scene signals the network already exposes, reasons over them with explicit rules, and restricts the command to the safe set they define. The learned policy still performs perception, planning, and nominal control; the guard does not choose routes or replace the planner. Its role is narrower: it prevents the vehicle from executing a command that would move it toward an unrecoverable and dangerous state.

The guard uses five safety rules, each derived in closed form from an established safety or vehicle-dynamics model and reduces to a single bound on throttle, brake, or steering. The rules are:

  • R1 (Planned-Path Collision Safeguard): Prevents the vehicle from following its intended waypoint path into a solid obstacle. The rule first suppresses forward acceleration and escalates to braking when the remaining gap becomes unsafe.

  • R2 (Speed-Limit Compliance): Keeps the vehicle inside the urban speed envelope by reducing positive longitudinal command and, if necessary, requesting mild braking until the speed returns to a safe legal range.

  • R3 (Red-Light Stop Compliance): Forces the command toward a full stop when a red light is active and the remaining distance no longer supports safe continuation through the intersection.

  • R4 (Pedestrian Right-of-Way Protection): Enforces early yielding whenever a pedestrian occupies, or is about to enter, the forward crossing corridor in front of the ego vehicle.

  • R5 (Speed-Conditioned Steering Stability): Shrinks the admissible steering range as speed increases, keeping the executed command within a safe lateral-acceleration envelope.

The guard reads a structured scene state S(o) = (K, B, V, W) from the same forward pass that produces the command, consisting of CenterNet-style object detections K (position, class, confidence), a bird’s-eye-view (BEV) semantic map B, radar returns V (radial velocity and range), and the policy’s own planned waypoints W.

The executed command is the Euclidean projection of the nominal proposal onto the feasible set defined by the rules, solving the quadratic program: x* = argmin x∈C(o) ½∥x − x0∥22, where C(o) = x ∈ U: Dx ≤ g(o). When several rules over-constrain the command, the most restrictive admissible bound wins, resolving the conflict in favor of safety.

The guard is evaluated on the Fail2Drive benchmark, a CARLA v2 benchmark of 200 short routes in Town13 covering 17 rare-hazard scenario classes, with matched pairs defining in-distribution and generalization splits. The state-of-the-art TransFuser v6 (TFv6) is used as a case study.

Key results on the generalization split (where hazards are staged with unfamiliar objects and layouts):

  • The guard lifts every metric of the frozen policy it wraps: SR by 10.6 points (+15.0%) and HM by 5.8 (+7.8%), without lowering Driving Score.

  • TFv6 alone loses 18.4% in harmonic mean between splits, and the guard reduces this to 9.8%, the smallest loss of any learned agent and second only to a privileged expert that reads ground-truth simulator state.

  • Collisions fall in every category: pedestrian collisions from 6.3% of routes to 3.0%, layout collisions from 15.0% to 7.7%, and vehicle collisions from 4.7% to 3.3%. The abstract states safety-critical collisions fall by up to 53%.

  • The guarded policy attains the highest Harmonic Mean of any learned agent (81.3 HM on generalization split).

On the in-distribution split, the guard costs 2.5% HM — this is expected because the guard can only narrow the set of allowed commands, never widen it. The paper notes: "A layer that helped on both splits would simply be a better policy, and one that hurt on both would be a worse one. Helping only where the driving is unfamiliar is what a safety constraint should do: it acts when the policy is about to err and stays inactive when the policy is right."

What increases instead is lost time: min-speed infractions from 72.3% to 92.0%, and timeouts by 5–8 points. Because the guard can only slow the car, every intervention is paid for in delay. The paper notes a limitation: "When the planner routes into the oncoming lane, no allowed command remains and the guard stops the car, taking a timeout instead of a collision. This is where the added timeouts come from: the guard can block an unsafe command, but it cannot invent a safe trajectory the planner never proposed, and it acts only on what perception reports."

The paper's three major contributions are:

  1. A neuro-symbolic safety guard for end-to-end driving that attaches to the command interface of a frozen neural agent.

  2. A set of grounded traffic-safety rules derived in closed form from responsibility-sensitive safety for longitudinal margins and the kinematic bicycle model for cornering.

  3. Measured robustness and safety gains — task success rises by 15%, safety-critical collisions fall by up to 53%, and competence lost on unfamiliar scenes is nearly halved, at an unchanged Driving Score and without retraining.

The paper concludes: These results argue for judging an agent not by the distance it covers but by the conditions its commands are guaranteed to satisfy.

Improvements for AI systems

Improvements to AI systems:

  1. Add a post-hoc rule-based safety layer to any frozen learned policy.
  • The improved system attaches a lightweight, non-learned guard to the final command interface of an already-trained agent (e.g., an autonomous driving policy).

  • It reads the policy’s existing outputs (detections, BEV map, radar, waypoints) and enforces explicit physical safety rules (collision avoidance, speed limits, red-light stops, pedestrian right-of-way, steering stability) by projecting the nominal command onto the nearest safe command set via a quadratic program.

  • Result: The system can guarantee that every executed command satisfies basic safety constraints, regardless of the policy’s training distribution or blind spots.

  1. Improve robustness to distribution shift and rare hazards without retraining.
  • The improved system wraps a frozen policy and, on unfamiliar or out-of-distribution scenes, automatically intervenes to prevent safety violations.

  • In the paper’s evaluation, this raises task success by 15%, reduces safety-critical collisions by up to 53% (pedestrian collisions from 6.3% to 3.0%, layout collisions from 15.0% to 7.7%), and nearly halves the performance drop between in-distribution and generalization splits (from 18.4% to 9.8% loss in harmonic mean).

  • The system achieves the highest harmonic mean of any learned agent on the generalization split (81.3) while keeping Driving Score unchanged.

  1. Provide interpretable, traceable safety interventions.
  • Each time the guard modifies a command, the system records which rule triggered the intervention and the exact bound applied.

  • The improved system can output a human-readable explanation (e.g., “R1: braking due to imminent collision with obstacle at 12m”) alongside the corrected command, enabling auditability, debugging, and trust in deployment.

  1. Enable safe operation in novel or adversarial scenarios where the learned policy fails.
  • The guard acts as a hard safety net that cannot be bypassed by statistical shortcuts.

  • For example, if the planner proposes a trajectory into the oncoming lane, the guard stops the vehicle (taking a timeout) rather than allowing a collision.

  • The system can be deployed in environments with unseen object types, layouts, or traffic situations, where the learned policy alone would violate basic rules.

  1. Trade off safety and progress explicitly.
  • The improved system can be configured to prioritize safety over speed: it will slow down or stop when rules are violated, accepting increased timeouts or min-speed infractions (e.g., from 72.3% to 92.0% in the paper) in exchange for zero safety-critical collisions.

  • This makes the system suitable for safety-critical applications (e.g., urban delivery, passenger transport) where avoiding crashes is more important than minimizing travel time.

  1. Serve as a plug-in safety module for any end-to-end driving agent.
  • The guard requires no retraining, no additional learned parameters, and attaches to the command interface of any policy that exposes a structured scene state (detections, BEV map, radar, waypoints).

  • The improved system can be applied to existing commercial or research driving stacks to instantly enforce safety rules, without altering the perception or planning modules.

What the improved AI system can do:

  • Guarantee that every executed driving command satisfies closed-form physical safety constraints (no collisions, speed limits, red-light stops, pedestrian yielding, stable steering).

  • Operate reliably in unfamiliar or rare-hazard scenarios where the base learned policy fails, improving task success and reducing collisions without retraining.

  • Provide transparent, rule-based explanations for every safety intervention.

  • Be deployed as a drop-in safety layer on top of any frozen end-to-end driving agent, with predictable trade-offs between safety and efficiency.

Abstract

Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guarantee safe driving, leaving their decision-making process opaque and safety constraints unenforced. We introduce a neuro-symbolic safety guard, a lightweight module that attaches to the final command interface of an already-trained agent. Immediately before a command reaches the vehicle, it checks the command against explicit safety rules and, only when necessary, replaces it with the nearest safe alternative. Each intervention is directly executable and traceable to the rule that triggered it, while the guard itself requires no retraining and adds no learned component. Evaluated on the long-tail benchmarks Fail2Drive and Bench2Drive using the state-of-the-art TransFuser v6 (TFv6) as a case study, the guard improves Success Rate by 15% and reduces safety-critical collisions by up to 53%, while preserving the original Driving Score.

Sources

Related papers