From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

arXiv:2608.21897 · cs.AI · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "From Solver Feedback to Faithful Plans".

Tom: Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: The title itself is pretty descriptive: "From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning." It really tells us that the focus here isn't just getting a program to run, but ensuring the resulting plan is actually faithful to the original natural language request.

Jane: That’s right, Tom; they are addressing a specific tension where optimizing only for solver success can lead to plans that are easy to solve but semantically shifted away from the human intent. The authors argue that this creates a conflict between learning without human labels and maintaining semantic faithfulness.

Lu: I think the authors hit on a crucial point when they mention that optimizing only for solver success can reward underspecified or semantically shifted formalizations, which is a problem we see everywhere in complex AI systems; it's like the model finds a shortcut to pass the test without truly understanding the original challenge.

Meng: I wonder how they balance that conflict during training; if the reward structure isn't tuned perfectly, the Actor might learn to generate cheap, easily verifiable code instead of complex, correct code. We need robust mechanisms for that.

Lalam: What really strikes me is their proposal to use a Judge not just as a simple pass/fail check but as something calibrated against what we consider true task-level correctness; that suggests they are trying to move beyond superficial success metrics.

The paper's summary: Tom: So, the core idea of the "From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning" paper is this solver-grounded framework where a single LLM acts as an Actor, Judge, and Editor to learn how to create PDDL specifications from natural language.

Jane: That's the gist; the Actor generates the initial PDDL specification based on what we ask for in natural language, and then the Judge evaluates that specification by checking if a deterministic solver can actually run it successfully.

Lu: And if that verification fails, which is common, the Editor steps in to perform a bounded number of refinements guided by diagnostic information from the solver to fix those specific errors. It’s this loop—generate, verify, and repair—that is central to how they propose learning without human annotations.

Meng: The paper mentions on PlanBench results showing an improvement in average success from thirty-five point five percent for LLM+P up to seventy point eight percent, but the real story is how they achieved that fidelity score of sixty-six point three percent faithful success while reducing semantic drift by six point four percent <ref:2608.21897#pg0,average success from 35.5% for LLM+P>. That's a big claim for annotation-free methods to make.

Lalam: That reduction in semantic drift is significant because it shows that their method doesn't just produce syntactically correct plans, but they are actually getting closer to the intended meaning of the original instruction, which is what we need for practical deployment.

The paper's improvements: Tom: Now let’s talk about the specific mechanisms they suggest for improvement; they propose using role-specific reward signals. The Actor learns by maximizing a reward that combines initial solver success with a term from the Judge’s quality score, essentially teaching it to aim for executable code early on.

Jane: And the Editor gets its learning signal through a reward function that prioritizes successful repairs while actively penalizing the number of edits made; this encourages it to be efficient and targeted when fixing errors instead of just making random changes.

Lu: The Judge is trained as a calibrated binary predictor of solver success, which is important because it's grounded in the actual outcome from the deterministic solver rather than some abstract human score. They also discuss Theorem C.two showing how this calibration error in the Judge relates to task-level correctness under certain assumptions.

Meng: The paper separates roles because they argue that a single final policy can leave an initial specification weakly constrained, and separating Actor and Editor allows reward-hacking gains to scale better with the model size if they are fully separate, though their shared backbone design limits that scaling.

Lalam: I see the benefit in this separation; it’s like having specialized cognitive modules within the same AI structure, where each one is optimized for a very specific type of learning challenge—global generation versus local correction.

Conclusion: Tom: So, wrapping up on "From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning," the main implication is that organizing solver feedback into generation, verification, and repair roles creates a much more robust way to learn natural language to PDDL formalization.

Jane: It means we can expect systems that are not only capable of producing plans but are also significantly more faithful to the original intent because they learn what it means to be executable and correct simultaneously.

Lu: The work points toward a new direction for symbolic AI where the learning signal is tightly coupled with the execution oracle, which is a very powerful concept for moving beyond simple pattern matching toward genuine reasoning.

Meng: Practically speaking, this framework could drastically cut down on the time spent debugging LLM-generated plans in real planning scenarios because the system learns to fix its own errors systematically rather than requiring human intervention for every failure.

Lalam: This paper shows that by structuring how an AI learns—giving it distinct roles like Actor, Judge, and Editor—we can build systems that are more reliable and less prone to those tricky semantic drifts we see in current models.

Tom: Indeed, "From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning" offers a compelling roadmap for building planning systems that are both scalable and actually useful in handling complex natural language tasks. We’ve got some heavy hitters on the other papers today, but this one really sets a strong standard for how we ground symbolic learning.

cs.AI

Submitted: 2026-08-22

Updated: 2026-09-28

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 83/100

The gist: Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit

Key concepts

Actor
The Actor's role is to propose the initial formal planning specification in PDDL format based on natural language instructions. It learns by maximizing a reward that combines initial solver success with a quality signal from the Judge, aiming for globally good specifications.
Judge
The Judge acts as a calibrated binary predictor that assesses whether a generated PDDL specification is executable by the deterministic solver. It is trained to provide a reliable quality signal based on actual solver outcomes rather than human labels, ensuring consistency with executability.
Editor
The Editor handles failures by performing bounded, diagnostic-conditioned repairs on the specification. It learns to maximize a reward that favors successful repairs while penalizing excessive edits, allowing it to refine flawed plans iteratively until the solver succeeds or the repair budget is exhausted.

Terminology

Summary

Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways. This paper proposes a solver-grounded multi-role reinforcement learning framework where a single language model acts as an Actor, Judge, and Editor for generation, verification, and repair.

How it works

The proposed framework is a zero-annotation approach for natural-language-to-PDDL planning that organizes learning around three coordinated roles: the Actor that proposes formal specifications, the Judge that provides a solver-calibrated quality signal, and the Editor that performs bounded diagnostic-conditioned refinement. This design preserves scalability by using a single LLM backbone conditioned on role embeddings to specialize in generation, verification, and repair.

The training process follows an episode structure where the Actor generates an initial PDDL specification, which is then queried against a deterministic PDDL solver for a binary success label and structured diagnostics. If the specification fails (solver returns 0), the Editor performs at most Hmax repair steps, each conditioned on the task, the current specification, and the latest diagnostic. The episode terminates when the solver succeeds or the repair budget is exhausted.

Role-Specific Learning Signals

The roles learn distinct objectives through specific reward signals designed to address different learning challenges. The Actor learns global specification generation by maximizing a reward combining initial solver success and a Judge-based shaping term: RA = b0 + λJ sJ(x, y0), where sJ is the Judge's quality score. The Judge is trained as a calibrated binary predictor of solver success, using rewards that enforce consistency with the solver outcome. The Editor learns diagnostic-conditioned repair by maximizing a reward that favors successful repairs while penalizing excessive edits: RE = rsolver(yT) − β · T, where T is the number of repair steps.

Judge Calibration and Correctness Connection

The Judge is trained against solver outcomes rather than human semantic labels, targeting solver-verified executability. The connection to task-level correctness depends on the agreement between solver success and a reference correctness notion on the policy-induced distribution. Theorem C.2 proves that under Assumption C.1 (Bounded solver–correctness disagreement), any Judge with calibration error δJ satisfies: E(X,Y) sJ(X,Y) − C(X,Y) ≤ ε + δJ. This demonstrates that the Judge provides a reliable shaping signal when solver success and task-level correctness have limited disagreement on generated samples.

Actor–Editor Decomposition and Reward Hacking

The separation of Actor and Editor roles is motivated by the fact that a single final-outcome policy can leave the initial specification weakly constrained once local repair becomes strong. The analysis shows that fully separate role models allow reward-hacking gains to scale with parameter count (Theorem 3.1), whereas the shared backbone design restricts this gain to O(σ√h), where h is the dimension of the Actor head, rather than scaling with the total backbone size p. This decomposition ensures that the Actor handles global formalization, the Judge supplies a calibrated quality signal, and the Editor performs bounded local correction.

Experimental Results and Conclusion

Across PlanBench domains (BlocksWorld, Mystery BlocksWorld, Logistics, Gripper) and zero-shot transfer benchmarks (ProntoQA, NATURAL PLAN), the method achieves strong in-domain planning success. The results show that organizing solver feedback into generation, verification, and repair roles improves both planning success and semantic faithfulness, rather than merely increasing solver executability. The method achieves 70.8% average success on PlanBench with a reduced semantic drift of 6.4%, demonstrating that this framework enables more scalable and faithful annotation-free symbolic planning. The Judge’s ability to distinguish between solved-but-unfaithful and truly faithful outputs further confirms the value of this multi-role structure.

The gist

A single LLM plays Actor, Judge, and Editor roles with shared parameters, grounded by deterministic PDDL solver feedback. This enables learning reliable natural-language-to-PDDL formalization through generation, verification, and repair roles. Across PlanBench domains, our method improves average success from 35.5% for LLM+P to 70.8%, achieves 66.3% faithful success, and reduces semantic drift to 6.4%. These results show that organizing solver feedback into generation, verification, and repair roles enables more scalable and faithful annotation-free symbolic planning.

Key Contributions

  1. We identify a key obstacle in annotation-free symbolic planning: raw solver success is a necessary but insufficient learning signal because it can diverge from semantic faithfulness to the original natural-language task.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the Solver-Grounded Multi-Role Reinforcement Learning for Symbolic Planning paper. The core innovation is moving from brittle, annotation-dependent Natural Language (NL) to executable Planning Domain Definition Language (PDDL) formalization by leveraging deterministic solver feedback within a coordinated reinforcement learning framework.

Here are the specific improvements that can be made to AI systems and what those improved systems will be capable of:


Based on this research, I propose the following concrete improvements to existing LLM-based planning systems:

  1. [Actor-Judge-Editor] Multi-Role Reinforcement Learning Architecture

  2. Incorporation of Solver Feedback into an Integrated Learning Loop

  3. Calibrated Quality Signal Generation for Specification Verification

  4. Diagnostic-Conditioned, Bounded Repair Mechanism

Here is a detailed breakdown of what these improvements enable the resulting AI system to do:

  1. [Actor-Judge-Editor] Multi-Role Reinforcement Learning Architecture

The improved system will utilize a single Large Language Model (LLM) backbone conditioned into three specialized roles:

  • The Actor: Responsible for global generation, translating high-level natural language instructions directly into candidate PDDL specifications.

  • The Judge: Responsible for learning a solver-calibrated quality signal that predicts the probability of solver success based on the generated specification. This moves beyond simple text plausibility to predict executability.

  • The Editor: Responsible for bounded, diagnostic-conditioned refinement of failed specifications using structured feedback from the PDDL solver (e.g., syntax errors, type mismatches, precondition failures).

  1. Incorporation of Solver Feedback into an Integrated Learning Loop

The system will be trained via a closed-loop Reinforcement Learning process driven by deterministic solver outcomes:

  • The Actor is rewarded based on both raw solver success and the Judge's quality score (shaping the generation towards executable specifications).

  • The Editor is rewarded based on successful repairs, penalized by the number of edits performed, ensuring it learns efficient error correction strategies.

  1. Calibrated Quality Signal Generation for Specification Verification

Instead of relying solely on binary solver success (which can reward easy but semantically weak formalizations), the Judge will provide a continuous scalar score calibrated against the true task-level correctness metric. This allows the system to learn a nuanced quality gradient that distinguishes between merely executable and truly faithful specifications.

  1. Diagnostic-Conditioned, Bounded Repair Mechanism

When an initial PDDL specification fails solver verification, the Editor will use structured diagnostic feedback (e.g., precondition unsatisfied, type mismatch) to guide its repair steps. This ensures repairs are targeted at specific symbolic errors rather than random text edits, leading to more stable and semantically faithful final plans. The repair process is strictly bounded (e.g., a maximum of 5 steps), preventing unbounded search loops and ensuring tractability.

By implementing these improvements, the resulting AI system will be capable of:

  • Generating executable PDDL specifications directly from complex natural language instructions without requiring costly human annotation or supervised fine-tuning on planning data.

  • Achieving significantly higher planning success rates (e.g., up to 70.8% on PlanBench) and substantially lower semantic drift compared to current LLM+P approaches, even in challenging domains with obfuscated predicates (like Mystery BlocksWorld).

  • Stabilizing the relationship between symbolic generation and task semantics by learning to avoid solver-facing shortcuts (like goal weakening or object type distortion), leading to plans that are not only syntactically valid but also logically sound and faithful to the original intent.

  • Demonstrating superior budget efficiency, often solving problems with significantly fewer total solver calls than prompting-based methods, as the learned roles efficiently guide the search toward executable solutions early in the process.

Sources

Related papers