PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning

arXiv:2605.09860 · cs.AI · Submitted 2026-05-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning".

Tom: Detailed Research Summary: PACE - Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning This research introduces PACE (Policy-Native Adaptive Decision Timing),

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back everyone to the radio show. Today we're diving into something quite fascinating that just hit arXiv: "PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning". Jane, can you give us the quick rundown of what this paper is all about?

Jane: Absolutely, Tom. Essentially, PACE tackles a core problem in long-horizon reasoning where the system has to decide not just what action to take next but also how many steps it should commit to before looking at new information. The central idea here is that instead of picking a fixed number for this commitment depth, the paper treats it as something the policy learns and adapts based on the current state. It claims this adaptive approach is better than having a set depth because the best commitment length can actually depend on where you are in the task.

Lu: That state-conditioned variable idea really opens up some creative avenues for how AI systems can plan over long sequences. It suggests that instead of rigid planning structures, we could have policies that dynamically adjust their scope of execution based on environmental cues. I'm really excited about the potential for this kind of flexible reasoning in complex, open-ended scenarios.

Meng: From an engineering standpoint, I’m curious how they manage the trade-off between making decisions too frequently and committing too long without a clear plan. That balancing act sounds tricky to implement reliably in practice.

Lalam: I see this paper as having implications for how we structure the decision-making processes within large language models. If the model can learn its own optimal execution length, it could lead to much more efficient and purposeful generation of complex outputs, improving overall AI culture by making decisions more deliberate rather than just reactive.

Tom: Exactly what I mean with that efficiency, Meng. So, this PACE paper suggests that fixing a commitment depth is suboptimal when the task demands different levels of planning at different points. Jane, can you elaborate on the main claim they are pushing here?

Jane: Well, the core theoretical foundation they lay out is a budget-constrained optimization problem where success means reaching the goal while staying within a training budget for total primitive actions. The paper proves that adaptive commitment depth strictly outperforms any fixed-depth commitment strategy whenever the locally optimal depth isn't constant across different states. This is formally shown as Prop. one in their work on PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning.

Lu: That formalization of the problem as a budget constraint is very rigorous; it moves this from just an observation to a mathematical proof of its superiority over static methods. It really grounds the intuition behind why we need state-conditioned decisions.

Paper summary: Meng: Rigor is good, but I still want to know what this means for deploying these models in real-world systems where we have strict latency requirements. Can a model that learns depth on the fly actually run fast enough?

Lalam: If it can learn the right commitment length, it means the AI won't waste time executing unnecessary steps just because a pre-set rule told it to go deeper or shallower than needed for that specific moment. That translates directly into more efficient computation and better resource management for the system.

Tom: Right, so we’re looking at a model that learns its own temporal abstraction during training, which is pretty ambitious stuff. So, if we look at their proposed architecture in "PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning", what are they actually proposing to build?

Jane: They propose a unified model-native Vision-Language Model architecture designed to do both things simultaneously: predict the next action and predict how long it should execute that action for. They integrate a dedicated depth head right alongside the standard action generation mechanism.

Lu: The separation into an explicit depth head and an autoregressive decoder for actions is a clever architectural move; it keeps the prediction of *what* to do separate from *how long* to do it, which makes the learning process cleaner.

Meng: I see how that structure addresses the problem we discussed earlier—it separates those two concerns into distinct learned components. But what kind of depth choices are they allowing? Are these discrete steps or continuous adjustments?

Jane: The paper specifies a set of possible commitment depths, H, which is defined as one two four eight. This output from the depth head is state-conditioned and learned by the model itself.

Lalam: Having those specific choices makes it tangible; it's not just some abstract variable but a set of meaningful temporal choices that the AI can select from depending on its understanding of the situation. That structure allows for much more predictable behavior in complex planning tasks.

Tom: That’s a neat distinction, Jane. So, this depth head outputs a categorical distribution over those four levels, and that’s conditioned on the state representation derived from the backbone model. How does that condition work in practice?

Jane: The backbone model first encodes the current state into a hidden representation we call z tk, which comes from the final token position of a fixed task prompt. This representation then feeds into the depth head to guide its choice of commitment depth.

Paper summary: Lu: That linkage between the state encoding and the depth selection is what makes it policy-native; it’s not just an external controller dictating timing, but an intrinsic property learned by the policy itself based on how well it perceives the current situation.

Meng: It sounds like a tight coupling between perception and planning that needs careful training. Does this method require a massive amount of data to learn these state-conditioned depths effectively?

Jane: The training uses a single GRPO objective that simultaneously optimizes both action selection and depth selection. The reward structure is designed around per-step progress signals, like the optimal-solution distance, and an episode reward based on how much that distance shrinks.

Lalam: Training with a progress-shaped reward is smart because it directly incentivizes the policy to choose commitment depths that lead to rapid and meaningful reductions in the remaining solution distance. It steers the learning toward efficient planning trajectories.

Tom: So, if we look at the results they present in "PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning", what kind of performance gains are they seeing when comparing their adaptive policy against those fixed-depth baselines?

Jane: On both the Sliding Puzzle and Sokoban tasks, the adaptive policy strictly Pareto-dominates every non-degenerate fixed-depth baseline. They report achieving up to twelve point five percentage points higher solve rate while using about twenty-five percent fewer primitive actions per episode compared to those fixed depth policies.

Lu: That quantitative comparison is compelling; seeing that performance leap on both tasks, regardless of the complexity of the task, really validates the hypothesis that adaptivity is superior here. It’s not just a small gain on one benchmark.

Meng: A twenty-five percent reduction in primitive actions is significant because it directly translates to lower computational load during execution, which addresses my earlier concern about efficiency and latency. That's a real practical win for any system we build.

Lalam: From the perspective of cultural impact, this suggests that future AI agents won't just follow rigid scripts; they will exhibit a level of temporal judgment in their planning that allows them to be far more economical with their computational resources while still achieving complex goals. That's a subtle but important improvement for how we design and interact with these systems.

Tom: So, looking at the title and authors of "PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning", what are the broader implications we should be considering right now?

Paper summary: Jane: In simple terms, PACE suggests that for any AI tasked with a long sequence of actions, we shouldn't pre-program a fixed plan for how deep to dive into every single decision point. Instead, the system should learn exactly how much commitment is appropriate at each stage based on what it sees right then and there.

Lu: The implication here is that we might move away from purely hierarchical planning structures towards more fluid, integrated reasoning where temporal abstraction becomes a learned property of the policy rather than an externally imposed constraint. That opens up possibilities for tackling much messier, less predictable problems.

Meng: I wonder if this means that future engineering efforts should focus less on designing perfect fixed plans and more on designing robust learning signals that allow the policy to discover its own optimal temporal strategy? That shifts the focus of development quite a bit.

Lalam: I think this points toward a future where AI systems are not just executors of commands, but rather autonomous reasoning agents capable of self-regulating their internal planning depth based on the immediate environmental feedback. That level of autonomy in decision timing is something we should definitely be focusing on for our next generation of models.

Tom: That's a big shift in thinking, moving from rigid planning to learned temporal adaptation, and it’s been really interesting to hear how the PACE paper lays out this mechanism. We've seen how it beats fixed strategies on Sliding Puzzle and Sokoban, which shows the theory holds up empirically. What should we think about next as we digest all of this?

Jane: We should definitely keep an eye on how researchers start applying this concept to more dynamic environments where the state changes in unpredictable ways, because that's where the adaptive nature of PACE could really shine.

Lu: I think exploring these non-degenerate oracle distributions mentioned in page one is key; understanding *how* the model learns to select those depths across different states will give us a lot of insight into its internal reasoning process.

Meng: I’m still focused on the practical side: how do we make sure this learned depth head doesn't just lead to overly erratic behavior when deployed in a critical system? We need stability as much as performance.

Lalam: Stability and efficiency go hand-in-hand; if the system learns to be efficient by cutting down unnecessary computation, that inherently reduces the chance of introducing errors through overly long, unproductive commitments. That synergy is really promising for real deployment.

Tom: It sounds like we’ve got a lot of exciting material here about this PACE paper, and it really highlights how crucial temporal decision-making is for advanced AI reasoning. We'll keep digging into these concepts as we go.

Conclusion: Tom: So, to wrap up our deep dive today, we're focusing on the title and authors of PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning and what it actually means for how we think about AI planning.

Jane: Basically, this paper is all about moving away from using a single, fixed plan for long tasks to instead letting the AI learn how deep it should commit to an action based on its current situation.

Lu: The authors are doing something really interesting by framing this as a budget-constrained optimization problem, which shows they’ve thought through the computational trade-offs very carefully.

Meng: I think the real impact here is in efficiency; if we can get an AI to decide how many steps it needs to take rather than just blindly following a script, that means less wasted computation overall.

Lalam: For me, this suggests a cultural shift where we start designing AI agents that have internal judgment about when to commit resources and when to seek new information.

Tom: It really boils down to this idea: instead of pre-programming the entire journey, the AI learns its own best way to navigate it step by step.

Jane: Exactly, and the core concept is that this adaptive timing can lead to better overall performance than sticking with a standard depth setting.

Lu: The way they’ve structured the model with that dedicated depth head alongside the action generator really showcases a clean way to implement this learned decision-making process.

Meng: From an engineering standpoint, it’s exciting because it suggests we can build systems that are more responsive and less brittle when things go unexpectedly during execution.

Lalam: I think this capability could fundamentally improve how we approach complex, open-ended problems in various domains beyond just puzzles or games.

Tom: Speaking of those domains, the paper shows strong results on established tasks like Sokoban, which gives us a solid starting point for seeing this in action.

Jane: It’s great that they achieved higher solve rates while using fewer primitive actions on those benchmarks, which backs up their main claim pretty well.

Lu: That quantitative evidence is what really helps bridge the gap between the theory and practical application of this adaptive timing mechanism.

Meng: I'm still thinking about how this learned depth head behaves when we push it into a situation where the state changes in a way we haven't seen before; that’s where I need to see more stability.

Lalam: That stability is vital, because if the AI starts making wild, erratic decisions about its commitment length, then the benefit of adaptivity disappears quickly.

Tom: So, we’ve seen the theory and the results: adaptive timing beats fixed timing on these tasks by being smarter about how much to commit.

Jane: And I think what this implies is that future AI agents will be much better at managing their own internal planning scope during execution.

Lu: The next step in research should probably be looking at how this applies when the state representation is even more abstract, like in truly multi-modal reasoning scenarios.

Meng: I agree, and I think we need to see more papers that address the practical deployment challenges of training these depth heads on real-world datasets.

Lalam: It makes me feel hopeful about our future development because this points toward agents that exhibit a more deliberate and efficient form of reasoning across all their operations.

Chen Li, Zhantao Yang, Fangyi Chen, Han Zhang, Anudeepsekhar Bolimera, Marios Savvides

Carnegie Mellon University

cs.AI

Submitted: 2026-05-11

Updated: 2026-09-28

Code: https://github.com/mpSchrader/gym-sokoban

Project page: https://stellar-neuron.github.io/recommit/Preprint

Importance score: 92/100

The gist: This research introduces PACE (Policy-Native Adaptive Decision Timing), a novel approach to long-horizon reasoning that treats temporal abstraction—specifically, the decision of how deeply to

Key concepts

Commitment Depth
This refers to how many subsequent primitive actions a policy executes before it takes a new observation and decides whether to replan. The paper argues that this depth should not be a fixed number but should change depending on the current situation to be optimal.
Budget-Constrained Optimization
The research frames the problem as finding the best policy within a limit on total actions used during training. It proves that adapting commitment depth to different states always outperforms using a single, fixed commitment depth because the ideal depth varies by state.
Depth Head
This is an added component in the model architecture specifically designed to output a categorical choice from a set of possible commitment depths (like 1, 2, 4, or 8). This head learns to condition its output on the current state representation to select the most appropriate planning horizon.
GRPO Objective
This is the specific reinforcement learning objective used to train PACE. It simultaneously optimizes two things: selecting the next action and choosing how long to commit (the depth). The reward function encourages trajectories that make rapid, meaningful progress toward solving the task.

Terminology

Summary

This research introduces PACE (Policy-Native Adaptive Decision Timing), a novel approach to long-horizon reasoning that treats temporal abstraction—specifically, the decision of how deeply to commit to a sequence of actions before re-evaluating the state—as a dynamic control variable rather than a fixed architectural hyperparameter. The core hypothesis is that fixing this commitment depth is suboptimal; instead, the system should learn an optimal, state-conditioned depth for each decision point.

Long-horizon reasoning fundamentally requires not only selecting the next action but also deciding how many subsequent primitive actions to execute before taking another observation and replanning. This concept is formalized as commitment depth. The central challenge addressed by PACE is balancing the computational cost of frequent replanning against the risk of compounding execution errors from overly long commitments.

The paper establishes a theoretical foundation demonstrating that fixed-depth commitment strategies are inherently suboptimal when the locally optimal depth varies across different states. Specifically, they formalize this as a budget-constrained optimization problem:

pi E s 0 about rho 0, pi [1[goal reached] K used(pi; s 0) K train]

where K used(pi; s 0) is the total number of primitive actions used by the policy pi starting from state s 0, and K train is a training budget constraint. They identify the commitment depth surrogate within this optimization framework, proving that adaptive, state-conditioned depth strictly dominates any fixed-depth commitment whenever the locally optimal depth is not constant across states (Prop. 1).

PACE proposes a unified model-native Vision-Language Model (VLM) architecture designed to jointly predict what to execute and for how long. This is achieved by integrating a dedicated depth head alongside the standard action generation mechanism.

The proposed policy consists of three main components:

  1. Backbone: A Qwen2.5-VL-7B model that encodes the current state into a hidden representation, z tk, derived from the final token position of a fixed task prompt.

  2. Depth Head (pi h): This head outputs a categorical distribution over a set of possible commitment depths, H = 1, 2, 4, 8. This output is state-conditioned and learned by the model.

  3. Action Head (pi a): An autoregressive decoder that generates an action sequence of length h k (the chosen depth) conditioned on the state representation z tk and only the prior actions—crucially, it does not re-encode the state for every step within the commitment.

The policy is trained using a single GRPO objective that simultaneously optimizes action selection and depth selection. The reward structure is designed to encourage efficient progress:

  • Per-step Progress Signal (d(s)): Defined as the optimal-solution distance (the shortest path/minimum remaining steps to the goal).

  • Per-episode Reward (R(tau)): A combination of a binary success reward and a dense progress reward: R(tau) = 1[solved] + lambda (d), where d is the mean per-step reduction in the optimal distance.

This objective ensures that the policy prioritizes trajectories that yield rapid, meaningful reductions in solution distance, effectively training it to choose appropriate commitment depths based on how much progress can be made before a state change necessitates replanning.

The empirical validation strongly supports the theoretical claims:

  • Pareto Domination: The adaptive policy strictly Pareto-dominates every non-degenerate fixed-depth baseline across both the Sliding Puzzle and Sokoban tasks, regardless of training budget or task complexity.

  • Superior Performance: On both tasks, the adaptive model achieves up to 12.5 percentage points higher solve rate while utilizing approximately 25% fewer primitive actions per episode.

  • State-Conditioned Depth Mechanism: The mechanism is mechanistically traceable: the depth selection head directly tracks state and task structure, leading to straighter trajectories that match or exceed fixed baselines on every measured metric.

  • Competitive Benchmarking: Despite being a 7B backbone, the adaptive policy outperforms large frontier models like GPT-5.5 and Claude Sonnet on both tasks.

Improvements for AI systems

Based on the provided research, here are specific, high-impact improvements for AI systems derived from Chen Li et al.'s work:


The core improvement is shifting commitment depth from a hand-tuned scalar to a learnable, state-conditioned variable within a unified model architecture. This leads to an adaptive policy that optimizes the trade-off between replanning cost and compounding execution errors dynamically.

Here are the specific improvements and what the improved system can achieve:

  1. The AI system will adopt a single, model-native Vision-Language Model (VLM) with two specialized heads: a depth head and an action head, sharing one backbone.

  2. This VLM will jointly predict two things at every decision point:

Choose the optimal commitment depth from the set of admissible values, specifically from the learned distribution over commitment depths, denoted as a categorical distribution over H = [1, 2, 4, 8].

Generate an action sequence of that chosen length (the commitment) using an autoregressive decoder.

  1. The system will be trained via a GRPO-style Reinforcement Learning objective using a dense reward signal that combines task success with state-dependent progress metrics. The reward function will be:

Reward = 1[goal reached] + λ tanh(∆¯d), where ∆¯d is the mean per-step reduction in the optimal solution distance (a positive value indicates moving closer to the goal).

  1. The depth head will learn a state-conditioned commitment depth, meaning it will dynamically select a longer commitment when it senses high reliability in progress (near-goal states) and a shorter one when progress signals are noisy or uncertain (far-from-goal states).

  2. The action head will learn to generate the correct sequence of primitive actions for the chosen length, conditioned only on the current state encoding and the required commitment length, without needing to re-encode intermediate states within that commitment.

The improved AI system can perform the following specific capabilities:

  1. It will solve complex long-horizon visual reasoning tasks (like Sliding Puzzle and Sokoban) with significantly higher success rates (up to 12.5 percentage points better than fixed-depth baselines).

  2. It will achieve this efficiency gain while using approximately 25% fewer primitive actions per episode compared to the best fixed-depth policies on those same tasks, leading to faster execution and lower computational cost.

  3. It will demonstrate superior robustness against budget constraints (both tight and loose evaluation budgets), maintaining a competitive solve rate even when the decision budget is severely restricted.

  4. It will exhibit task-dependent reasoning: it learns that the optimal commitment depth varies based on the specific dynamics of the task (e.g., committing longer on Sokoban where planning-dominant errors are catastrophic, and shorter on Sliding where uncertainty dominates).

  5. It will be significantly more competitive than current frontier closed-source models (like GPT-5.5 and Claude Sonnet) and open-weight VLMs when evaluated under the same commitment interface, proving that scale alone does not solve the problem—state-conditioned commitment depth is the key mechanism.

Sources

Related papers