AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving

arXiv:2601.01762 · cs.RO, cs.CV · Submitted 2026-01-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving".

Dev: AlignDrive proposes a novel cascaded planning paradigm designed to explicitly align longitudinal motion reasoning with surrounding agent behavior along the intended driving path,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're looking at the paper titled "AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving," which sounds pretty technical, but the basic idea is that it tackles a coordination issue where current state-of-the-art parallel planning architectures struggle to link speed decisions with what's happening around the vehicle along its intended path.

Dev: It does sound intricate, Rosa, and I'm curious about the authors. They are a team from XJTU, Horizon Robotics, and UCASL—they seem like they bring a lot of solid expertise in both robotics and AI systems to this problem.

Taro: I've been looking at the paper structure, and it seems they are trying to solve that coordination failure by moving away from parallel planning toward a cascaded framework where longitudinal planning is explicitly conditioned on the predicted lateral drive path.

Rosa: Exactly, Taro, so instead of treating speed and path as completely separate things predicted in parallel, they're making the speed prediction dependent on the spatial context of where the car is supposed to go next.

Dev: That conditional dependency sounds promising for handling those tricky cut-ins or sudden changes in surrounding agent behavior that we see in real driving.

Taro: I think it's a significant structural shift because it reframes longitudinal planning from being an independent prediction task into a process that is directly influenced by the path itself, which should help with robustness when things go wrong.

The paper's summary: Rosa: To get into what they are actually doing, the core summary of "AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving" is that they introduce a cascaded framework and an anchor-based regression design to condition longitudinal prediction on the lateral drive path.

Dev: The summary suggests that by treating longitudinal planning as 1D displacement prediction along the path, they are trying to reduce geometric uncertainty and make the model focus more sharply on interaction-driven dynamics instead of just predicting trajectories in two dimensions independently.

Taro: That reduction in degrees of freedom sounds smart because it simplifies the problem into a one-dimensional prediction along a specific structure—the drive path—which should inherently couple the longitudinal motion with the spatial context of surrounding agents.

Rosa: And they also mention that on the data side, they introduce a planning-oriented data augmentation strategy where they programmatically insert agents and relabel those 1D displacement targets to enforce collision avoidance logic rather than just memorizing patterns.

Dev: That sounds like a very deliberate way to train the model to learn causal collision avoidance, which is much more robust than just relying on typical expert patterns found in training data.

Taro: If they can successfully teach the AI this causal logic through that augmentation, it means the system won't just be good at nominal driving scenarios but should handle rare or unexpected interactions much better when things misbehave.

The paper's improvements: Rosa: Thinking about the specific improvements, the authors point out three main contributions: first, proposing that cascaded planning paradigm where longitudinal planning is explicitly conditioned on a predicted lateral drive path.

Dev: They also reformulate the task as a simpler 1D displacement prediction problem along the drive path, which I think is key because it cuts down on complexity compared to predicting full 2D trajectories independently.

Taro: And third, they introduced an effective, planning-oriented data augmentation strategy by modifying only the 1D displacement labels in response to inserted agents to enforce collision avoidance.

Rosa: The implication of these improvements is that they achieve a driving score of eighty-nine point zero seven and a success rate of seventy-three point one eight percent on Bench2Drive, and they show strong generalization on Fail2Drive where it consistently outperforms prior RGB-only methods under the Generalization setting.

Dev: I see what you mean; the ablation studies also confirm that the path-conditioned design achieves a higher overall driving score and reduces the collision rate significantly when compared to parallel formulations.

Taro: That performance on Fail2Drive, where it beats prior RGB-only methods in the Generalization setting, really tells us that this coupling approach is more effective at handling unseen interactive scenarios than methods that treat path and speed as separate entities.

Conclusion: Rosa: So wrapping up the AlignDrive paper on "Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving," it seems the main implication is that explicitly conditioning longitudinal planning on a predicted lateral drive path leads to better coordination between steering and speed decisions.

Dev: It moves the task from independent prediction to path-conditioned reasoning, which sharpens the model's focus on interaction dynamics by reformulating speed as 1D displacement prediction along that specific path.

Taro: I think it points toward a future where AI systems don't just generate trajectories but actively manage the coupling between what the car does laterally and how fast it goes longitudinally based on the immediate surroundings.

Rosa: And they used a planning-oriented data augmentation strategy to make this robust, which suggests we can train these systems to handle edge cases by simulating collision avoidance logic through those modified 1D displacement labels.

Dev: The system architecture itself, with its Drive Path Predictor refining queries and the Longitudinal Planning module using anchor-based offset regression, shows a clear path toward lower latency while maintaining high planning ability.

Taro: If we can keep that low latency while achieving that level of coordination against misbehaving agents, it really opens up possibilities for truly reliable autonomous driving in complex, real-world environments.

Yanhao Wu, Haoyang Zhang, Fei He, Rui Wu, Yanhu Shan, Congpei Qiu, Liang Gao, Wei Ke Tong Zhang

School of Software Engineering, XJTU Horizon Robotics Shenzhen Loop Area Institute University of Chinese Academy of Sciences

cs.RO, cs.CV

Submitted: 2026-01-05

Updated: 2026-09-29

Comments: NeurIPS 2026

Project page: https://yanhaowu.github.io/AlignDrive

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: AlignDrive proposes a novel cascaded planning paradigm designed to explicitly align longitudinal motion reasoning with surrounding agent behavior along the intended driving path, addressing a

Key concepts

Cascaded Planning Paradigm
This framework moves away from parallel planning by making longitudinal planning dependent on the predicted lateral drive path. It reframes longitudinal motion as a process directly influenced by the spatial context of where the car is supposed to go next, rather than treating speed and path as separate predictions.
1D Displacement Prediction
The paper reformulates longitudinal planning as predicting 1D displacement along the specific drive path. This simplification reduces geometric uncertainty and forces the model to focus more sharply on interaction-driven dynamics instead of predicting full two-dimensional trajectories independently.
Planning-Oriented Data Augmentation
This strategy involves programmatically inserting agents into training data and relabeling the 1D displacement targets. This modification enforces collision avoidance logic during training, teaching the AI causal avoidance rather than just memorizing typical driving patterns.

Terminology

Summary

AlignDrive proposes a novel cascaded planning paradigm designed to explicitly align longitudinal motion reasoning with surrounding agent behavior along the intended driving path, addressing a critical coordination failure in state-of-the-art parallel planning architectures. This method is significant because it transforms longitudinal planning from an independent prediction task into a path-conditioned reasoning process, which is essential for robust generalization and safety in complex, interaction-heavy autonomous driving scenarios where speed decisions critically depend on the surrounding environment.

Model Design: Cascaded Framework

The core innovation lies in a cascaded framework that conditions longitudinal planning on the lateral drive path via an anchor-based regression design. Instead of predicting full 2D trajectory waypoints independently, the model first establishes the drive path as a structured geometric prior, and then predicts longitudinal motion as 1D displacement values along that path. This design offers two key benefits: first, anchoring longitudinal reasoning to the drive path reduces geometric uncertainty and provides a coarse but explicit alignment between longitudinal motion decisions and the spatial context of surrounding agents. Second, by reformulating planning as 1D displacement prediction, it removes unnecessary geometric degrees of freedom, sharpening the model’s focus on interaction-driven dynamics.

Longitudinal Planning Reformulation

The paper reframes longitudinal planning as a simpler problem: 1D displacement prediction along the path. This is achieved by defining anchors that represent sequences of longitudinal displacements for the current step and future steps. The final displacements are obtained by adding predicted offsets to the anchors, which couples the drive path geometry with agent interactions. This process ensures that each longitudinal planning query incorporates both the geometry of the drive path and its anchor-based temporal reference.

Data Augmentation Strategy

To improve generalization, AlignDrive introduces a planning-oriented data augmentation strategy. This involves programmatically inserting agents and relabeling 1D displacement targets to enforce collision avoidance. Specifically, when a virtual agent is inserted that would collide with the ego vehicle, the ground-truth displacement sequence is scaled by a factor β = Dsafe/Dorig to ensure the displacement at each step is consistently reduced to avoid collisions. This structural clarity allows the model to learn causal collision-avoidance logic rather than merely memorizing expert patterns, as this formulation untangles variables that parallel 2D models fail to learn from.

Key Contributions and Evaluation

The authors detail three primary contributions:

  1. Proposing a novel cascaded planning paradigm where longitudinal planning is explicitly conditioned on a predicted lateral drive path.

  2. Reformulating the task as a simpler 1D displacement prediction problem along the drive path.

  3. Introducing an effective, planning-oriented data augmentation strategy by modifying only the 1D displacement labels in response to inserted agents.

The method was evaluated on several benchmarks, achieving a driving score of 89.07 and a success rate of 73.18% on Bench2Drive and demonstrating strong generalization on Fail2Drive, where it consistently outperforms prior RGB-only methods under the Generalization setting. Ablation studies confirm that the path-conditioned design (Variant C) achieves a higher overall driving score and reduces the collision rate significantly compared to parallel formulations.

System Architecture

The AlignDrive system consists of three main components:

  1. The Drive Path Predictor, which refines queries via cross-attention with image features to encode drive paths, maps, and dynamic agents through iterative refinement blocks.

  2. The Planning-oriented Data Augmentation module, which decodes agent queries into bounding boxes and re-encodes them as structured features to enable synthetic agent insertion with consistent supervision.

  3. The Longitudinal Planning module, which predicts displacements along the drive path using anchor-based offset regression, incorporating Path-aware Interaction and Contextual Interaction through cross-attention with agent and map features.

The final output involves a hierarchical selection strategy that chooses the candidate based on confidence scores to yield the final trajectory for downstream control via PID controllers. The model achieves the best overall performance, reaching a Merging score of 75 in multi-ability tests. The method also demonstrates lower latency than models like DriveTransformer and VAD while maintaining superior planning ability.

Limitations and Future Work

The authors note that their agent insertion relies on minimal rule-based constraints and does not explicitly use road information, suggesting a promising future direction to constrain inserted agents according to road elements. Furthermore, they acknowledge that existing generative models often focus on trajectory-level generation without modeling the coupling between longitudinal decisions and surrounding agent dynamics along the driving path, viewing integrating their formulation into these paradigms as a promising direction for future work. The open-loop evaluation confirms the lowest collision rate in this category.

Implementation Details Summary

The model uses 900 agent queries, 100 map queries, 6 drive path queries, and 5 longitudinal queries.

Improvements for AI systems

As a fastidious researcher, I have analyzed the AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving paper. The core innovation lies in explicitly coupling longitudinal planning (speed/displacement) with lateral planning (path) via a cascaded anchor-based regression design, and enhancing generalization through a planning-oriented data augmentation strategy.

Here are the specific improvements that can be made to AI systems by implementing the AlignDrive framework:


  1. The improved AI system will transition from decoupled or parallel planning architectures to a cascaded, path-conditioned reasoning process.

  2. It will move from predicting full 2D trajectories independently to predicting sequences of 1D longitudinal displacements conditioned on the predicted lateral drive path.

Specific capabilities enabled by these improvements:

  1. The system will achieve significantly improved coordination between steering and speed decisions, preventing the unsafe outcomes where a geometrically accurate path leads to an unsafe speed profile (e.g., failing to brake for a cut-in).

  2. The model will reduce geometric uncertainty by anchoring longitudinal reasoning to the explicit spatial context of surrounding agents along the driving path, leading to more collision-aware longitudinal planning.

  3. By reformulating speed as 1D displacement prediction, the model will sharpen its focus on interaction-driven dynamics rather than redundantly encoding static geometry, resulting in more precise and efficient motion control.

  4. The system will exhibit superior generalization to rare edge cases and safety-critical events (like sudden cut-ins or pedestrian crossings) because the data augmentation strategy programmatically simulates these events by inserting agents and relabeling longitudinal targets to enforce collision avoidance, rather than relying solely on memorizing expert patterns.

  5. The resulting system will be highly robust across distribution shifts (e.g., from Bench2Drive to Fail2Drive), maintaining state-of-the-art performance even in unseen interactive scenarios where parallel formulations typically fail.

  6. The AI system will provide a higher Merging score (up to 75 compared to previous bests of 50) in multi-ability tasks, demonstrating enhanced capability in handling complex, consecutive lane changes and cut-ins.

  7. The final output will be a set of candidate drive paths and corresponding longitudinal displacement sequences, which are then selected based on confidence scores that penalize collision risks with predicted agent motions before being executed by independent PID controllers for steering and speed control.

Sources

Related papers