SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration

arXiv:2609.39645 · cs.CL, cs.AI, cs.LG · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration".

Jane: SEPAL introduces a method for reliable large language model collaboration by assigning three private Actor–Critic teams to direct reasoning, evidence grounding, and verification,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving on to the specifics of this paper, we have to talk about who wrote it and what they're calling their method. The full title is "SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration," and the authors are Ren, Zhang, Li, Qi Hengyi Zhang, Naibo Wang.

Jane: Those authors are clearly coming from a strong group of research institutions in China, which often means they're dealing with very rigorous standards when it comes to building reliable systems.

Lu: I find the collaborative nature of their approach really interesting; they aren't just looking at one model or one technique, but designing a whole system around how multiple agents can interact reliably.

Meng: It’s compelling because it tries to solve the inherent difficulty in making AI teams actually work together effectively on complex tasks instead of just talking over each other.

Lalam: The title itself makes me feel optimistic about achieving a level of reliability we haven't seen before in how LLMs tackle complex question answering.

Tom: So, when you break down "Separated Expert Pairs with Answer-Level Fusion," what does that mean in plain language for our listeners?

Jane: Simply put, it means instead of one big group debating everything together, they set up three separate teams to handle different parts of the task—one for solving, one for checking facts, and one for double-checking—and then only combine their final results after each team has done its private work.

Lu: It’s about giving each part a dedicated focus so that the whole process becomes much more manageable and less prone to getting tangled up in conflicting suggestions.

Meng: It sounds like they are moving away from a messy, shared discussion where everyone gets confused by conflicting inputs, toward something more modular and controlled.

Lalam: That modularity is what really stands out; it suggests we can build much more dependable AI collaboration tools for real-world applications where accuracy matters.

Tom: So the authors are essentially proposing a way to structure the debate so that correction happens locally within each expert group before anyone sees the final result. Where does this lead us next?

Jane: It sets up a clear blueprint for how we can achieve reliable consensus in multi-agent scenarios by emphasizing isolation during the refinement steps.

Lu: It opens up a lot of avenues for future research into how these different roles might best complement each other in complex problem-solving situations.

The paper's summary: Tom: Now that we know the structure, let's get into the actual substance of what SEPAL is proposing. The paper summarizes this system as assigning three private Actor-Critic teams—Direct, Evidence, and Verification—to handle reasoning, grounding, and checking.

Jane: The summary emphasizes that while multi-agent discussion can improve things by letting agents question arguments and revise answers, the existing methods often struggle with the trade-off between getting corrective feedback and keeping diverse alternatives for a final decision.

Lu: They point out that in shared history debates, an early mistake can easily leak into several later candidates because everyone is looking at the same flawed path.

Meng: That leakage of errors across candidates is a major issue they are trying to solve by giving each team its own private feedback path during its revision rounds.

Lalam: It really highlights that the current gap in research is between having feedback that can fix mistakes and actually keeping those different corrected candidates separate for a final vote.

Tom: So, the paper's main summary is about introducing this specific mechanism to manage those two benefits simultaneously: getting corrections while maintaining diversity for voting.

Jane: That’s right; the system allows revision within each team in isolation, and then they only combine the results after a fixed process to make a final decision based on a majority vote.

Lu: This framework seems to give us an explicit way to control exactly what information each agent has access to during its work, which is pretty powerful.

Meng: It moves the focus from just "how many agents are there" to "how do we structure their interaction" in a way that minimizes negative side effects.

Lalam: I think this summary really captures the essence of what makes SEPAL distinct—it’s not just more agents; it's smarter, structured collaboration.

The paper's improvements: Tom: Let's look at what they actually suggest as the improvements over prior methods. The paper highlights that their method improves accuracy by showing that this approach can boost mean accuracy by one point eight one percentage points over a matched single Actor-Critic pair across five different backbones and five question-answering benchmarks.

Jane: That improvement is quite substantial, especially considering it happens consistently across all those different model sizes and task types they tested, which is a strong indicator of robustness for the method.

Lu: The fact that this performance gain isn't just an isolated win on one specific model or task; it suggests the structural improvements are fundamentally sound.

Meng: I’m also interested in the component contrasts they show, specifically how Critic feedback produces the largest improvement, which tells us where we should focus our training efforts for maximum impact.

Lalam: That insight about Critic feedback being most valuable is really useful because it tells us that focusing on improving the quality of those internal review processes yields the best return on investment in terms of accuracy gains.

Tom: And they also found that the same trained Actors gain substantial points from round zero through round four, meaning iterative refinement is a key driver of improvement, not just a single pass.

Jane: So, it’s not just about one good answer; it's about the continuous cycle of local revision leading to better results over multiple rounds before we settle on one final answer.

Lu: This confirms that the system isn't just static; it’s designed to leverage iterative refinement as a core part of its strength, which is something I was thinking about when considering how these models might tackle more complex reasoning problems.

Meng: It shows that the training and preference learning process is carefully designed to prioritize those revisions that lead to "more correct Actor revisions," which gives us a clear target for optimizing our training objectives.

Lalam: That whole process, from initialization through inference, seems incredibly well-designed because it systematically builds up the solution step-by-step rather than relying on a single lucky guess.

Conclusion: Tom: So, to wrap up this discussion on SEPAL, the authors have summarized that the main implication is that we can achieve reliable LLM collaboration by revising candidates locally and fusing those completed answers.

Jane: The overall message is that by assigning specialized roles and enforcing isolation during revision, we can prevent errors from spreading and still capture the benefits of iterative refinement.

Lu: It suggests a path forward for designing more sophisticated multi-agent systems where agents have clearly defined mandates to avoid the confusion of uncontrolled interaction.

Meng: From an engineering side, it gives us a clear recipe: revise candidates locally and fuse completed answers, which is something we can immediately start implementing in our next project design cycle.

Lalam: I feel really good about this approach because it provides a robust framework for building systems that are truly dependable in their reasoning.

Tom: This SEPAL framework is a solid step forward in moving the conversation toward more structured, expert-driven collaboration, and it really shows how much structure helps with reliability.

Jane: It’s been a fascinating deep dive into how specialized roles can lead to tangible accuracy gains when handled correctly.

Lu: I'm looking forward to seeing how researchers build on this foundation to explore even more complex ways these three roles might interact in the future.

Meng: We should definitely keep an eye on the diagnostics they provide regarding when to stop revising, as that will be crucial for setting good operational rules for deployment.

Lalam: I'm just happy we have this paper to look at because it lays out such a clear path toward building more trustworthy AI systems.

Weijie Ren, *, Yanwen Zhang, *, Hao Li, *, Zhuolin Qi, Hengyi Zhang, Naibo Wang

Zhejiang University · University of Electronic Science and Technology of China

cs.CL, cs.AI, cs.LG

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 22 pages, 4 figures. Code: https://github.com/zhansan114514/SEPAL

Code: https://github.com/zhansan114514/SEPAL

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 90/100

The gist: SEPAL introduces a method for reliable large language model collaboration by assigning three private Actor–Critic teams to direct reasoning, evidence grounding, and verification, allowing them to

Key concepts

Direct Team
This team focuses on achieving a decisive derivation of the answer. Its Actor proposes solutions, and its Critic guides revisions specifically to ensure the proposed answer is sound and conclusive.
Evidence Team
The Evidence team's goal is to ground its answers in factual support. Its process involves proposing answers and having a Critic guide revisions based on relevant definitions, facts, or domain principles from provided evidence.
Verification Team
This team acts as an independent checker. It re-solves the question and checks alternatives to ensure robustness. Its Critic guides revisions focused on validating the proposed solution against other possibilities before finalizing its output.

Terminology

Summary

SEPAL introduces a method for reliable large language model collaboration by assigning three private Actor–Critic teams to direct reasoning, evidence grounding, and verification, allowing them to revise answers in isolation before combining only their final results through a majority vote. This approach addresses the limitations of shared discussion by preventing errors from propagating across candidates while retaining revision gains.

How it works

The core mechanism involves three isolated role pairs: Direct (for decisive solve), Evidence (for grounded support), and Verification (for checking alternatives). Each team consists of an Actor that proposes an answer and a Critic that guides its revision process. The system follows a fixed protocol where each team undergoes four private revisions guided by its specific Critic, preventing feedback from carrying errors across candidates.

The three roles are assigned distinct reasoning objectives:

  1. Direct: Emphasizes a decisive derivation.

  2. Evidence: Grounds the answer in relevant definition, fact, passage evidence, or domain principle.

  3. Verification: Independently re-solves the question and checks alternatives before stating the final answer.

Training and Preference Learning

The teams share a backbone and training questions but learn separate adapters. The initialization involves a sequence: Role SFT / Critic DPO / Actor DPO Trained separately for each pair. Specifically, role-specific supervised training initializes the Actors, followed by Critic and Actor preference learning. The Critic's feedback is valued during training when it leads to more correct Actor revisions, utilizing an ordered rule where a retained pair is kept if its reward gap meets a threshold (e.g., if its gap ≥ 0.6).

Inference and Fusion

During inference, the process is structured as follows:

  1. In parallel, each Actor produces an answer and its paired Critic returns feedback.

  2. For rounds 1 through 4, each Actor revises using only its previous answer and feedback from its own pair. Crucially, Do not expose a t i, c t i, or z i to another pair during Steps 1–3.

  3. After round four (R4), the final answer is parsed as zi = g(a4 i).

  4. The final decision uses a fixed rule: If a valid answer occurs at least twice among them, return it; otherwise the Direct answer is the fixed fallback.

Evaluation and Results

SEPAL was evaluated across five open-weight backbones (2B to 8B parameters) and five question-answering benchmarks. Across these settings, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor–Critic pair, with improvements observed across all five backbones. The decision diagnostics show that the vote reaches an accuracy of 76.31% on average, exceeding the strongest individual role in 21 of 25 cells, and Pairwise agreement covers 78.70–79.22% of examples.

Key Findings

The analysis indicates that Correction and regression counts provide a second target: deciding when to stop revising. The component contrasts show that Critic feedback produces the largest improvement, with the same trained Actors gaining substantial points from R0 to R4. Furthermore, the oracle-any-role gap reaches 10.83 for Mistral, quantifying an opportunity for future calibrated routers or verifiers. The overall recipe is summarized as: revise candidates locally and fuse completed answers. The final decision records trace both repairs and selection errors to individual questions.

The gist: SEPAL assigns three private Actor–Critic teams to direct reasoning, evidence grounding, and verification, allowing them to revise answers in isolation before combining only their final results through a majority vote. This approach addresses the limitations of shared discussion by preventing errors from propagating across candidates while retaining revision gains. Across five backbones, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor–Critic pair. The component contrasts show that Critic feedback produces the largest improvement, with the same trained Actors gaining substantial points from R0 to R4. The overall recipe is to revise candidates locally and fuse completed answers. The final decision records trace both repairs and selection errors to individual questions.

The three roles are assigned distinct reasoning objectives:

  1. Direct: Emphasizes a decisive derivation.

  2. Evidence: Grounds the answer in relevant definition, fact, passage evidence, or domain principle.

Improvements for AI systems

Based on the SEPAL (SEPARATED EXPERT PAIRS WITH ANSWER-LEVEL FUSION) framework described in this paper, here are specific improvements that can be made to existing LLM question-answering systems:


  1. The system can perform high-stakes, complex reasoning by utilizing three specialized expert agents simultaneously for every query.

  2. It prevents the propagation of errors across candidates during the revision phase by isolating each expert team's feedback loop (private history).

  3. The system can achieve a significant accuracy boost (averaging 1.81 percentage points over single pairs) by tailoring reasoning objectives for different tasks: one for direct derivation, one for evidence grounding, and one for verification/alternative checking.

  4. It allows researchers to systematically analyze the impact of different collaborative mechanisms by isolating components (e.g., comparing SFT-only actors vs. trained critics vs. full revision rounds) to determine whether initial quality or iterative refinement is the primary source of gain for a specific model backbone or task type (e.g., BoolQ, MMLU, BBH).

  5. The system can provide a robust decision-making mechanism by using a fixed majority vote on final answers, which is highly reliable and avoids the pitfalls of shared discussion where errors can synchronize multiple agents.

  6. It offers a quantifiable metric for selection gap (Oracle-any-role accuracy), allowing researchers to identify questions where even the best three experts disagree, signaling a need for better external validation or routing mechanisms.

  7. The system provides diagnostic data on revision dynamics, showing exactly when and how much accuracy is gained from the first critical feedback exchange versus subsequent rounds, enabling researchers to set optimal stopping rules for iterative refinement.

  8. The system can be configured to exploit role complementarity, where the final decision is made by combining distinct reasoning paths (Direct vs. Evidence vs. Verification), potentially leading to more nuanced and accurate answers than a single expert could provide alone.

This improved AI system can function as a highly reliable, multi-stage reasoning engine for question answering, capable of generating answers that are not just individually plausible but have been vetted through three distinct logical lenses before being fused into a final decision. It moves beyond simple self-correction to structured, role-specific expertise.

Abstract

Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while single-pair Actor-Critic collaboration refines only one candidate. We introduce SEPAL, which assigns three private Actor-Critic teams to direct reasoning, evidence grounding, and verification. Role-specific training gives the teams different reasoning objectives beyond sampling variation. Each Critic guides revisions within its own team, preventing feedback from carrying errors across candidates. Once revision ends, majority voting combines only the final answers, keeping the reasoning histories separate until the decision. Across five open-weight backbones and five question-answering benchmarks, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor-Critic pair, with improvements across all five backbones. Code is available at https://github.com/zhansan114514/SEPAL.

Sources

Related papers