Closed-Loop Refinement and Execution for Learned Driving Planners

arXiv:2610.00992 · eess.SY, cs.RO, cs.SY · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Closed-Loop Refinement and Execution for Learned Driving Planners".

Dev: Learning-based driving planners are typically trained and evaluated in open loop against logged trajectories,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So, looking at this paper, "Closed-Loop Refinement and Execution for Learned Driving Planners," the main idea seems to be tackling those issues where learned planners fail when they have to actually drive. The authors introduce Closed-Loop Refinement and Execution, or CLRE, which is a hierarchical receding-horizon control framework designed to fix problems like stalling or abrupt braking that happen when small errors in the plan cause things to go wrong during execution. It claims this works without needing any new learned models, just by refining the existing plan and adding a safety layer.

Dev: That sounds promising because it directly addresses the problem of error accumulation when actions change subsequent observations, which is something we see all the time in closed-loop systems. I’m wondering how robust this refinement step actually is when we are talking about real-time performance and loop rates—the core of my concern as a control engineer.

Taro: It’s interesting that they keep the upstream planner frozen while adding this refinement layer. That suggests the learned planner is doing a lot of heavy lifting on the long-term route, and CLRE is just polishing it up locally to handle immediate dangers. I’m curious what happens when the world gets really weird and misbehaves in ways that the initial plan didn't anticipate.

Rosa: Exactly, Taro; they are essentially using an optimal control problem in the upper layer to balance keeping on track with predicting how surrounding agents will interact. They generate a candidate set from several initializations and then filter that set using a prediction-conditioned oriented bounding-box, or OBB, feasibility test to keep only the safest options.

Dev: Filtering by an OBB clearance threshold sounds like a solid way to manage immediate conflicts without having to impose hard collision constraints directly inside the main optimization problem, which is smart for computational efficiency. But how does this refinement process translate into actual vehicle movement speed and smoothness when we look at the execution layer?

Taro: That’s where I want to know more about the execution layer—what specific safety measures they put in place to handle those sudden deviations from the refined plan. If a candidate fails, what is the fallback behavior when no feasible option remains?

Rosa: The paper details a forward-range speed bound and a backup policy for execution, which acts as a safety net if no refined candidate passes the feasibility test. They even have a saturated proportional braking law that ensures safe longitudinal execution by capping the commanded speed at whatever the safety layer dictates.

Paper summary: Dev: Capping the speed based on closing rates to obstacles and required stopping distances sounds like a necessary constraint, but I need to know about latency here; how quickly can this entire refinement and execution cycle actually run in practice without causing noticeable lag or jitter in the vehicle's control inputs?

Taro: The sensitivity analysis mentioned in their work suggests that the progress and commitment weights are highly sensitive terms, which implies that a small change there could drastically alter how the system prioritizes route advancement versus sticking to a path. That points to where we might need deep investigation when dealing with unpredictable real-world behavior.

Rosa: It seems the authors are trying to find a sweet spot between aggressively following the learned plan and being reactive enough to handle immediate, unexpected hazards. This whole CLRE framework is designed specifically to improve closed-loop behavior when dealing with imperfect learned models.

Dev: So, to wrap up this summary, the core contribution of "Closed-Loop Refinement and Execution for Learned Driving Planners" is proposing a training-free method that uses trajectory refinement and an OBB test to improve the performance of a frozen learned planner in closed-loop driving scenarios.

Taro: I think the real implication here is that we can get significantly better closed-loop performance just by adding this hierarchical control structure without needing to retrain the entire underlying planner model. It opens up possibilities for deploying these systems in environments where they need to react quickly but still maintain a high-level route plan.

Rosa: It really is exciting because it moves the focus from just training a perfect planner to building a robust system around an imperfect one that can handle real-world driving uncertainties. We've seen significant improvements in benchmarks, with the driving score going up substantially from forty-three point four one to fifty-six point four two and route completion improving from fifty-seven point two seven to seventy-two point two three.

Dev: The reduction in collision events, dropping from seventy to fifty-three, is a concrete result that speaks directly to the safety aspect of this framework. My main question remains how long we can expect this system to operate reliably outside of the controlled lab environment before these kinds of failure modes reappear.

Taro: The transferability studies they conducted, showing that CLRE works with different upstream planners like UniAD and DriveTransformer, suggest that the refinement mechanism itself is quite general and doesn't rely on a specific learned model structure. That’s a big deal for real-world deployment because it means we don't have to build a whole new learning pipeline every time we want to improve closed-loop safety.

Paper summary: Rosa: It sounds like the authors are really pushing the idea that this refinement approach is a viable way forward for making autonomous driving more dependable in complex, dynamic situations. This paper shows how we can augment existing systems rather than starting from scratch.

Dev: So, the title "Closed-Loop Refinement and Execution for Learned Driving Planners" points to a clear focus on improving the execution phase of systems that already have a learned planner. We need to keep an eye on those sensitivity results regarding the progress and commitment weights, as those are clearly critical parameters for tuning this system effectively.

Taro: I think the future work they hint at will be crucial for truly pushing this out into unpredictable environments where failures aren't just minor stalls but major safety risks. We need to see how these constraints hold up when the environment deviates significantly from the logged trajectories used for training.

Rosa: It seems the authors have laid out a very practical and targeted approach to making these AI systems more reliable in real driving situations. This framework offers a tangible path toward better closed-loop performance without requiring massive retraining efforts.

Dev: We’ve covered the core idea and the immediate results, but we still have to figure out the practical deployment hurdles regarding latency and loop stability for our control system design. It's not just about getting a high score; it's about ensuring that high score is achieved within acceptable real-time constraints.

Taro: I think the broader implication is that this kind of layered, safety-oriented control structure could become standard practice for integrating learned planners into more complex, safety-critical driving tasks. It’s about building resilience into the decision-making process itself.

Rosa: That's a good summary of where we are with this paper, focusing on how CLRE improves reliability through refinement and execution layers. We’ll keep watching this space for more research on these types of hierarchical control methods.

Dev: Indeed, the focus on mitigating those specific failure modes—stalling and conflict motion—makes this paper very relevant to our work in ensuring stable, low-latency vehicle control.

Taro: We’ve covered what the CLRE framework is and why it's important for closing the loop on learned planners. The potential to generalize this refinement technique across different planner architectures is a significant point for future autonomy research.

Rosa: That's all we have for this discussion on "Closed-Loop Refinement and Execution for Learned Driving Planners".

Conclusion: Rosa: So, we've looked at the technical details of Closed-Loop Refinement and Execution for Learned Driving Planners, and now we need to get to where this all leads us in terms of what it actually means for us out there on the road.

Dev: Yeah, I think the real takeaway is that they’ve found a way to make those pre-trained AI planners behave much more predictably when they're actually driving, which is a huge deal for control engineering.

Taro: I agree with Dev; it addresses the fundamental problem of error accumulation in closed-loop systems, especially when actions change based on new observations.

Rosa: Exactly, and we should talk about the authors and what their framing of this problem suggests about how we build these complex autonomous systems moving forward.

Dev: The authors are focusing heavily on that hierarchical control structure they introduced, which seems to be their main contribution for ensuring stability during execution.

Taro: I think the implication here is that we might not need to retrain the whole planner just to fix execution issues; we can build a refinement layer on top of it.

Rosa: That's a big idea, and I wonder if this means autonomous driving will become much more robust in messy, real-world environments than we currently expect.

Dev: If the loop rate and latency constraints can be met with this refinement process, then we could see much tighter control over vehicle dynamics in complex traffic scenarios.

Taro: We need to look closely at those results showing collision events dropping from seventy to fifty-three; that is a very tangible improvement for safety metrics.

Rosa: I think the long-term impact will be seeing these systems deployed in more unpredictable settings, not just perfectly scripted routes.

Dev: I'm still thinking about the robustness outside the lab; how long can we rely on this refinement layer to keep things stable when faced with unexpected sensor noise or novel obstacles?

Taro: That's exactly what we need to investigate next, because if it holds up under those real-world stresses, it could genuinely make a lot of autonomous driving systems safer.

Rosa: Right, so the question for next time is whether this refinement mechanism can keep our vehicles safe when they encounter situations that weren't perfectly represented in the training data.

Cornell University

eess.SY, cs.RO, cs.SY

Submitted: 2026-10-01

Updated: 2026-10-03

Comments: 8 pages, 6 figures, 3 tables. Submitted to the 2027 American Control Conference (ACC 2027)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

The gist: Learning-based driving planners are typically trained and evaluated in open loop against logged trajectories, but this approach fails to guarantee reliable closed-loop execution because planning

Key concepts

Closed-Loop Refinement and Execution (CLRE)
A hierarchical system where an upper layer optimizes a nominal trajectory based on a frozen planner's output, and a lower layer executes the refined path. This process involves generating candidate trajectories and testing them for safety before execution.
Trajectory Refinement Objective Jt(xt, ut)
This objective guides the trajectory toward progress while respecting constraints. It includes terms for tracking the nominal path (Huber loss), committing to a geometric path, and minimizing distance to the route target point to ensure forward movement.
Prediction-Conditioned Oriented Bounding-Box (OBB) Feasibility Test
A safety check applied during refinement that filters candidate trajectories. It retains only those candidates whose minimum predicted clearance from obstacles over the planning horizon meets a required threshold, ensuring predicted spatial safety.
Saturated Proportional Braking
The execution law for longitudinal control. It sets the commanded speed as the minimum of the reference speed and a calculated safe speed based on obstacle closing rates. Full braking is triggered when collision thresholds are met.

Terminology

Summary

Learning-based driving planners are typically trained and evaluated in open loop against logged trajectories, but this approach fails to guarantee reliable closed-loop execution because planning errors accumulate when actions change subsequent observations. This research introduces Closed-Loop Refinement and Execution (CLRE), a hierarchical receding-horizon control framework that mitigates failure modes like stalling, conflict motion, and abrupt braking by refining the plan while keeping the upstream planner frozen.

How it works

CLRE operates as a hierarchical system where an upper layer refines a nominal trajectory using optimal control, and a lower layer executes the resulting candidate. The upper layer treats the frozen learned planner's output as a reference and solves a finite-horizon optimal control problem that balances route progress against predicted agent interaction. This refinement process involves several steps:

  1. Solving the optimization problem from several initializations to generate a candidate set.

  2. Applying a prediction-conditioned oriented bounding-box (OBB) feasibility test to retain only candidates whose minimum predicted OBB clearance over the horizon meets a threshold.

Trajectory Refinement Objective

The refinement objective, denoted as Jt(xt, ut), is composed of several terms designed to guide the trajectory toward progress while respecting constraints. Key components include:

Point-wise Nominal Tracking,

which uses a Huber loss function (jpt) to anchor the refinement to the nominal trajectory, penalizing deviations quadratically but allowing departure when other objectives favor different motion.

Geometric Path Commitment,

which uses a term (jct) that penalizes spatial departure from the nominal path while allowing advancement at different rates, controlled by a commitment gate γt.

Forward Progress,

represented by jprog, which minimizes the distance to the route target point (rt), rewarding terminal displacement in the route direction.

Interaction and Constraint Terms

The refinement objective incorporates terms related to surrounding agents and road boundaries:

  1. An interaction cost (jint) is calculated based on predicted agent motion modes, using softplus penalties for both outer and inner interaction ellipses defined by ego footprints. This term shapes candidates away from predicted conflicts without imposing a hard collision constraint inside the refinement problem.

  2. A road-boundary penalty (jbnd) penalizes proximity to sampled lane boundaries, ensuring the trajectory respects a prescribed boundary-proximity margin (mbnd).

  3. A lateral corridor regulation term (jcor) limits large lateral excursions from the navigation direction, preventing the optimizer from exploiting large lateral excursions in pursuit of route progress.

Execution and Safety Layer

The lower layer executes the selected refined trajectory through a tracking controller augmented by specific safety measures:

  1. A Backup Policy is implemented: if no refined candidate is feasible, the system generates waypoints from the navigation route centerline and latches the measured ego speed until a feasible candidate appears.

  2. A Forward-Range Speed Bound (vsafe) uses a forward-facing radar to set a ceiling on longitudinal speed, calculated based on closing rates to obstacles and required stopping distances.

  3. A Saturated Proportional Braking law is used for longitudinal execution, where the commanded speed (v cmd t) is the minimum of the reference speed (v ref t) and vsafe t, with full braking applied when collision thresholds are met.

Evaluation and Results

The CLRE framework was evaluated on 126 Bench2Drive routes using VAD as the upstream planner. The results demonstrated significant improvements:

CLRE raises the driving score from 43.41 to 56.42 and route completion from 57.27 to 72.23, and reduces collision events from 70 to 53.

A key ablation study showed that removing the progress term substantially reduced forward motion, confirming its importance for maintaining route advancement. Furthermore, removing the OBB feasibility test increased collisions by 70%, proving it provides an additional check beyond the soft interaction term (11). The framework showed transferability to other upstream planners like UniAD and DriveTransformer.

Solver Sensitivity

The refinement problem is solved using Adam optimization on a nonconvex objective. Sensitivity analysis indicated that the progress and commitment weights are the most sensitive terms, with winner retention being highly dependent on these parameters. The system's performance was also tested across different iteration budgets, showing that while candidate selection can be sensitive to budget increases, feasibility verdicts remain stable, though deviations concentrate near the horizon end. The implementation currently runs at a speed limited by CPU computation time for the feasibility test but shows potential for GPU-parallel trajectory optimization.

The gist

CLRE is a training-free framework that improves the closed-loop behavior of a frozen learned planner through trajectory refinement, OBB feasibility testing, and execution-level intervention.

Improvements for AI systems

As a diligent researcher, I have analyzed this paper, Closed-Loop Refinement and Execution for Learned Driving Planners (CLRE), which introduces a novel framework for improving closed-loop execution of learned driving planners.

Here are the specific improvements that can be made to AI systems by implementing the CLRE framework:

  1. Replacement of Open-Loop Planners with Closed-Loop Execution:

  2. Mitigation of Stalling and Route Incompletion Failures:

  3. Proactive Conflict Avoidance via Prediction-Conditioned Screening:

  4. Preservation of Learned Planning Intent during Refinement:

  5. Robust Execution under Uncertainty and Dynamic Conditions:


Specific Capabilities of the Improved AI System (CLRE):

  1. The system can execute learned driving trajectories in real-time, ensuring continuous vehicle motion rather than stalling when the upstream planner generates an insufficient plan.

  2. It can maintain route progress even when the nominal trajectory is overly conservative or stalls by actively refining the path to prioritize forward displacement (via the Forward Progress term).

  3. The system can proactively avoid predicted conflicts by using an Oriented Bounding Box (OBB) feasibility test that screens candidate trajectories against predicted agent motion, ensuring a minimum clearance threshold is met before execution.

  4. It can operate on top of existing learned planners (like VAD or UniAD) without requiring retraining of the upstream model or adding new learned components, making it a training-free improvement mechanism.

  5. The system can handle dynamic obstacles by incorporating interaction costs based on high-probability motion modes and using a multi-disk ego footprint representation to accurately model longitudinal extent during optimization.

  6. During execution, the system can dynamically adjust speed constraints based on forward-facing radar data to ensure safe stopping distances (Time to Collision calculations) and apply saturated proportional braking for smooth, controlled deceleration when necessary.

Sources

Related papers