Can Attribution Predict Risk? From Multi-View Attribution to Planning Risk Signals in End-to-End Autonomous Driving
Le Yang, Haijun Liu, Jiawei Liang, ShangQuan Sun, Xiaochun Cao
Sun Yat-sen University · University of Chinese Academy of Sciences · Nanyang Technological University
cs.LG
Submitted: 2026-08-15
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: The paper investigates whether attribution can serve as a predictive signal for planning risk in end-to-end autonomous driving, rather than merely as a post-hoc explanation tool.
Terminology
Summary
The paper investigates whether attribution can serve as a predictive signal for planning risk in end-to-end autonomous driving, rather than merely as a post-hoc explanation tool. The authors state: We investigate whether attribution can go beyond post-hoc explanation and serve as a predictive signal for planning risk in end-to-end autonomous driving.
End-to-end autonomous driving models generate future trajectories from multiview inputs, improving system integration but introducing opaque decisions and hard-to-localize risks. The authors note: Existing methods either rely on auxiliary monitoring models or generate textual explanations, but are decoupled from the planning process and fail to reveal the visual evidence underlying trajectory generation.
They further explain that "planning differs from image classification by taking six-view camera images as input and predicting continuous multi-step trajectories, requiring attribution to capture both critical views and regions and their influence on outputs. Moreover, whether attribution maps can support risk identification remains underexplored."
The authors propose a hierarchical attribution framework for end-to-end planning. Specifically, "using L2 consistency with the original trajectory as the objective, we design a coarse-to-fine region attribution strategy that searches candidate regions across the full six-view input and refines attribution within them."
The attribution method treats six camera views as a unified attribution space. The objective function combines two criteria: sufficiency, defined as Fsuf(S) = −∥ŷ(S) − ŷ∥2
which measures whether keeping only regions in S still produces a trajectory close to the full-input trajectory, and necessity, defined as Fnec(S) = ∥ŷ(V S) − ŷ∥2
which measures whether removing S causes substantial deviation. The combined objective is F(S) = λsuf Fsuf(S) + λnec Fnec(S).
The coarse-to-fine search first merges spatially adjacent subregions into groups and runs the coarse-stage greedy search over these groups,
then for refinement, each group from the coarse stage now enters refinement... it uses the union of all preceding groups as a conditional prefix.
The authors derive three complementary statistics from the saliency tensor T ∈ R(C×H×W):
-
Attribution entropy:
The global concentration of attribution reflects whether a planning decision relies on a small portion of the full multi-view visual space,
defined as H = −Σ p(c,u,v) log p(c,u,v).A smaller H indicates that the planner relies on a limited set of critical visual locations, suggesting stronger global over-reliance.
-
Within-camera spatial variance:
Conditioned on each camera view, the spatial spread of attribution reflects how localized the planner's visual reliance is within the image plane,
defined as σ2 sp = Σ c (m c / Σ c' m c') · (1/m c) Σ u,v M(c) u,v ‖r u,v − r̄(c)‖2.A smaller σ2 sp indicates that attribution is concentrated around localized regions within individual views.
-
Cross-camera Gini coefficient:
For a multi-camera planner, attribution allocation across views reflects whether the planning decision depends disproportionately on particular cameras,
defined as Gini cam = (Σ c Σ c' m c − m c') / (2C Σ c m c).A higher Gini cam indicates that attribution is concentrated in a few views.
Experiments are conducted on the validation split of the nuScenes dataset with three representative end-to-end autonomous driving models: BridgeAD, UniAD, and GenAD.
Two planning-risk proxies are used: average displacement error (ADE) as a continuous risk proxy, and a binary collision indicator flagging whether the ego footprint E(ŷ t) at any predicted waypoint intersects an obstacle bounding box.
Attribution statistics track planning risk: The joint predictive strength is nearly identical across the three planners, with ρ ADE of 0.310 / 0.299 / 0.307 and AUROC of 0.768 / 0.770 / 0.765 for BridgeAD / UniAD / GenAD, respectively.
The authors note: Which attribution axis carries the strongest signal is architecture-dependent; that such a signal exists is not.
Signal is not a stand-in for object-layout complexity: Collision AUROC under controls is at or near chance, ranging from 0.55 to 0.64 across the three planners... Attribution statistics raise both metrics far above this on every planner.
Generalization to held-out scenes: Joint attribution statistics on held-out scenes give ρ ADE of 0.298 / 0.322 / 0.343 for BridgeAD / UniAD / GenAD and AUROC of 0.763 / 0.782 / 0.779, close to the in-domain numbers.
At k = 10% budget, Recall and Precision coincide, both reach 0.40 to 0.42 across the three planners, four times the random baseline.
Robustness across attribution methods: Replacing the proposed method with RISE yields joint-model ρ ADE values of 0.205 / 0.226 / 0.292 and collision AUROC values of 0.642 / 0.651 / 0.728,
showing the diagnostic signal survives a different attribution algorithm.
Faithfulness and efficiency: On every planner and every faithfulness metric, our method clears RISE by a wide margin and trails EAGLE by a narrow one.
The proposed method is the fastest on all three planners, taking roughly a fifth of EAGLE's wall-clock time, and slightly faster than RISE.
The authors conclude: "Experiments on nuScenes with BridgeAD, UniAD, and GenAD show that these attribution-based signals correlate with trajectory error, predict collision risk, and generalize to unseen samples. These results suggest that attribution patterns can provide useful warning signals for decision-level risk in end-to-end driving planners."
The main limitation noted is attribution efficiency: Although the proposed hierarchical search improves over exhaustive subset selection, it still requires multiple planner forward passes and cannot yet support real-time risk monitoring.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement:
Implementation: Build a coarse-to-fine hierarchical attribution module that:
-
Treats all six camera views as a unified attribution space (not per-camera)
-
Uses L2 trajectory consistency (not class confidence) as the attribution objective
-
Performs two-stage search: global group ranking across all views, then subregion refinement within top groups
-
Produces per-pixel saliency tensors at 270s/sample (vs 1290s for exact methods)
Resulting capability: The system can now identify exactly which visual regions in any of six cameras drove a specific trajectory prediction, with faithfulness (Insertion AUC 0.77–0.78) close to exact-search baselines (0.81–0.88) at 5× lower compute.
Implementation: Extract three statistics from the saliency tensor:
-
Attribution entropy (H): global concentration over the joint six-view pixel space
-
Within-camera spatial variance (σ2 sp): spatial dispersion within each view
-
Cross-camera Gini coefficient (Gini cam): imbalance of reliance across cameras
Fit these with ridge regression (for continuous ADE) and logistic regression (for binary collision).
Implementation: Train the risk model on 80% of scenes, evaluate on held-out 20%. Use the predicted risk score to rank samples and flag top-k as high-risk.
Implementation: Include three object-configuration controls (object count, spatial spread, cross-camera Gini of objects) plus ten extended scene controls (ego speed, near-field density, radar support, interaction dynamics) to verify the attribution signal is not a proxy for scene layout.
Implementation: Replace the hierarchical attribution with RISE (random sampling) while keeping the same statistics and evaluation protocol.
-
Explain any trajectory in terms of specific visual regions across six cameras, with per-pixel saliency that is faithful to the planner's actual decision process.
-
Flag risky planning decisions before they occur using only the planner's own visual-attribution pattern, achieving AUROC 0.77 for collision prediction across three different planner architectures.
-
Triage unseen driving scenes by risk level, enabling efficient allocation of human review or closed-loop simulation resources to the most dangerous 10–20% of samples.
-
Distinguish planner-induced risk from scene-induced risk, so that a busy intersection is not automatically flagged as dangerous unless the planner's visual reliance pattern indicates over-concentration.
-
Monitor planners across versions or training runs by tracking how attribution statistics shift, providing an input-side early warning that complements output-side trajectory-error evaluation.
-
Operate on any end-to-end planner without modification, since the attribution and statistics are post-hoc and planner-agnostic (validated on BridgeAD, UniAD, GenAD).
Abstract
End-to-end autonomous driving models generate future trajectories from multi-view inputs, improving system integration but introducing opaque decisions and hard-to-localize risks. Existing methods either rely on auxiliary monitoring models or generate textual explanations, but are decoupled from the planning process and fail to reveal the visual evidence underlying trajectory generation. While attribution offers a direct alternative, planning differs from image classification by taking six-view camera images as input and predicting continuous multi-step trajectories, requiring attribution to capture both critical views and regions and their influence on outputs. Moreover, whether attribution maps can support risk identification remains underexplored. To address this, we propose a hierarchical attribution framework for end-to-end planning. Specifically, using L2 consistency with the original trajectory as the objective, we design a coarse-to-fine region attribution strategy that searches candidate regions across the full six-view input and refines attribution within them. We further extract three attribution statistics as predictive signals for planning risk, including attribution entropy to measure how concentrated the planner's reliance is over the joint visual space, within-camera spatial variance to characterize how spread out the attribution is within each view, and cross-camera Gini coefficient to quantify how unevenly attribution is distributed across the six cameras. Experiments on BridgeAD, UniAD, and GenAD show that these statistics correlate with planning risk, achieving Spearman correlations of 0.30 plus or minus 0.07 with trajectory error and AUROC of 0.77 plus or minus 0.04 for collision detection. The signal generalizes to held-out scenes with negligible degradation and remains stable under an alternative attribution baseline.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks