Efficient LLM Collaboration via Planning

arXiv:2506.11578 · cs.AI · Submitted 2025-06-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Efficient LLM Collaboration via Planning".

Jane: The paper was written by Byeongchan Lee, Kyungjoon Park, Jonghoon Lee, Dongyoung Kim, Dongjun Lee et al. from KAIST and Yonsei University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title & Authors: Tom: So, since the title is "Efficient LLM Collaboration via Planning," what exactly does that imply about the approach?

Jane: It suggests they' aren't just running one huge model, but rather using a structured way to make smaller and larger models work together.

Meng: That makes sense when you look at the current constraints; we can’t afford to run GPT-4o for every single request anymore in many practical applications.

Lu: The title implies that the planning step is the critical bridge, allowing us to structure how these different capacities interact without needing a massive, monolithic model.

Tom: Exactly! It' not just about speed, but about cooperation between two distinct capabilities we have available right now.

Jane: It’s a sophisticated way of saying that they are building an architecture designed for collaboration rather than just one single powerful AI.

Meng: I agree, it addresses the core trade-off between how much a model can do and how much money we have to spend to make it work.

Lalam: And I see this as a massive step toward making advanced AI accessible, not just confined to massive cloud servers.

Summary/Abstract: Tom: That’s a great starting point, Jane. So, what is the fundamental problem they are solving in the abstract of "Efficient LLM Collaboration via Planning"?

Jane: They're showing that large models are amazing but extremely expensive to run frequently, and small models are cheap but limited in their reasoning power for complex tasks.

Meng: The core idea is finding a way to use those strengths of small and big models simultaneously without paying the full cost every time we need a result.

Lu: The paper proposes COPE, which is short for Collaborative Planning and Execution, as the solution to this structural limitation.

Tom: COPE sounds like it's not just dumping the task on a big model if the first attempt fails, right?

Jane: No, that’s where it differs from previous methods; they allow a smaller model to plan and guide the process before full escalation to a bigger model.

Meng: This is highly practical because it means we can use cheap models for ninety percent of the tasks and only involve the expensive ones when necessary.

Lalam: It allows us to be adaptive in our AI usage, which is a massive win for sustainability and resource management in the future.

Improvements/Results: Tom: The results section really shows how effective this collaboration is, especially when we look at the benchmarks they tested.

Jane: We saw some really impressive numbers on MATH-five hundred where COPE actually achieved seventy-five point eight percent accuracy compared to GPT-4o’s seventy-five point two percent.

Meng: But what's even more compelling is the cost reduction; they managed to do that while cutting the inference API cost by nearly forty-five percent.

Lu: And I loved seeing how it scales on difficulty, especially in Table ten where they achieved a massive improvement of fifteen percent on those most challenging Level five problems.

Tom: That suggests the benefits of COPE grow as the problem gets harder, which is a really interesting finding.

Jane: It’s not just good for easy tasks; it performs better when the complexity demands that collaboration we discussed earlier.

Meng: And looking at code generation on MBPP, Table five shows similar trends where COPE delivered higher accuracy and lower costs than the competition.

Lalam: This proves that this framework isn's ability to guide execution is not limited to math; it applies to creative and structured problem-solving too.

Conclusion/Wrap-up: Tom: So, we’ve seen how "Efficient LLM Collaboration via Planning" works, from the initial concept through the incredible results across different tasks.

Jane: It seems like a robust framework that truly balances capability and computational cost without having to sacrifice performance for efficiency.

Lu: I think we can look forward to this being applied in agentic workflows where long-term planning is necessary but expensive execution is limited.

Meng: For me, the practical implication here's that it makes advanced AI deployment scalable across a huge range of real-world scenarios.

Lalam: It fundamentally changes how we define "efficient" AI—it’s not just about speed, it's about smart resource allocation.

Tom: Absolutely. As we wrap up our discussion on this fantastic paper, I want to thank the authors for their work in "Efficient LLM Collaboration via Planning."

Jane: And thank you to Lu, Meng, and Lalam for sharing your insights with us today.

Lu: I’m excited to see how this impacts more than one specific domain of research.

Meng: We're ready to implement this framework at scale for a massive user base.

Lalam: This is a huge step toward making AI truly collaborative, moving it toward the future of human-AI interaction.

Byeongchan Lee, Kyungjoon Park, Jonghoon Lee, Dongyoung Kim, Dongjun Lee, Jinwoo Shin, Jaehyung Kim

KAIST · Yonsei University

cs.AI

Submitted: 2025-06-13

Updated: 2026-08-25

Importance score: 84/100

The gist: This paper introduces COPE (Collaborative Planning and Execution), a test-time collaboration framework designed to bridge the trade-off between the high performance of large language models (LLMs)

Key concepts

COPE (Collaborative Planning and Execution)
COPE is the solution proposed in the paper. It is an architecture that allows a smaller model to plan and guide a complex task before escalating to a larger, more expensive model. This process structures how different AI capacities interact.
LLM Collaboration
This refers to using multiple models—both small and large—together rather than relying on one massive model. It involves structuring the workflow so that different models work cooperatively, balancing capability with computational cost.
Inference API Cost
This is the expense incurred when running a large language model (LLM) to generate responses. The paper demonstrates that COPE can significantly reduce this cost by limiting the use of expensive, high-capacity models.

Terminology

Summary

This paper introduces COPE (Collaborative Planning and Execution), a test-time collaboration framework designed to bridge the trade-off between the high performance of large language models (LLMs) and the low computational cost of smaller models. By utilizing planning as a lightweight, transferable intermediate to guide execution, the framework enables efficient cross-model collaboration, making it a scalable solution for cost-aware inference in realistic deployment scenarios.

The Motivation for Planning

The authors identify a critical trade-off in LLM deployment: while large models achieve remarkable results, they incur substantial monetary inference cost, whereas smaller models are easier to deploy but have limited capacity on complex tasks. Existing cost-aware methods often rely on independent delegation through multi-stage cascades, which limits the ability of models to jointly perform complex tasks in a structured and interactive manner. COPE addresses this by using planning to allow models to scaffold each other’s thinking.

The research is driven by three key observations regarding model interaction:

  1. Larger planners help smaller executors, such as when GPT-mini planning for Llama-3B increases accuracy.

  2. Smaller planners degrade larger executors, because low-quality plans generated by smaller models can hinder the execution ability of larger models.

  3. A model benefits from plans aligned with its capacity, suggesting that while large models can successfully scaffold their own execution, small models may require simpler goal types rather than complex guidelines.

How it works

COPE operates through a multi-stage cascade where small and large models alternate roles as planner and executor. The process is triggered by task-specific confidence measures, such as the consensus ratio (the fraction of samples agreeing on the most frequent answer) for reasoning tasks, test case pass rates for coding, or perplexity for open-ended generation. The framework proceeds through three distinct stages:

  • Stage 1: A small model acts as both planner and executor, sampling n plans and solutions. If the consensus exceeds a threshold tau 1, the answer is accepted.

  • Stage 2: If Stage 1 fails, a large model generates a new guideline-type plan, which is passed to the small model to attempt execution again. If the consensus exceeds tau 2, the task is complete.

  • Stage 3: If the small model still fails in Stage 2, the large model takes full control of both planning and execution to directly perform the task.

Experimental Results and Impact

The framework was evaluated across diverse benchmarks, including mathematical reasoning, code generation, open-ended tasks, and agent tasks. The results demonstrate that COPE achieves performance comparable to large proprietary models, while drastically reducing the inference API cost.

Key performance highlights include:

  • Mathematical Reasoning: On the MATH-500 dataset, COPE achieved 75.8% accuracy (surpassing GPT-4o’s 75.2%) while reducing cost by nearly 45%. It also showed significant gains on the more challenging AIME-2024 dataset.

  • Code Generation: On the MBPP benchmark, COPE improved accuracy to 66.4% compared to GPT-4o’s 64.0%, while cutting inference cost by nearly 75%.

  • Open-ended and Agent Tasks: COPE proved effective using perplexity as a confidence signal for open-ended generation and showed success in multi-step decision-making for agent tasks, where cost-efficiency is critical due to the long sequence of actions.

Improvements for AI systems

1. Dynamic Role-Swapping Inference Engine (Multi-Stage Cascade)

  • Improvement: Replace monolithic single-model inference with a three-stage collaborative architecture where models alternate between Planner and Executor roles based on task difficulty.

  • Capabilities: The system will drastically reduce API operational costs (by approximately 45% to 75%) while maintaining or exceeding the accuracy of top-tier proprietary models (e.g., GPT-4o). It will handle easy queries using local, free models and selectively escalate complex reasoning to high-cost cloud models only when necessary.

2. Dual-Layer Planning Interface (Granularity Scaling)

  • Improvement: Implement a tiered planning strategy that adjusts the abstraction level of the plan based on the model's capacity: use Goal-based planning (high-level objectives) for small models and Guideline-based planning (detailed strategic instructions) for large models.

  • Capabilities: This prevents the performance degradation typically seen when small models provide low-quality, overly complex instructions to large executors. The system will provide just enough cognitive scaffolding—simple goals for small executors to prevent confusion, and detailed heuristics for larger executors to maximize reasoning depth.

3. Task-Specific Consensus-Based Confidence Scoring

  • Improvement: Replace generic confidence metrics with task-optimized consensus ratios: majority voting for mathematical/symbolic reasoning, pass-rate verification for code generation, and perplexity thresholds for open-ended text generation.

  • Capabilities: The system will achieve highly reliable early exit capabilities. It will autonomously decide whether a solution is mathematically sound or syntactically correct before spending additional compute, ensuring that escalation only occurs when the internal logic is genuinely uncertain.

4. Plan-Augmented Self-Correction (Stage 2 Memory Retention)

  • Improvement: Modify the escalation protocol so that in Stage 2, the executor receives both its original failed plan and the new high-level plan from a larger model as joint context.

  • Capabilities: This enables the system to perform comparative reasoning, where the small model can leverage the discrepancy between its initial failed strategy and the large model's new guidance to correct its execution path without requiring explicit fine-tuning or supervised training.

5. Cost-Efficient Agentic Decision Loops

  • Improvement: Integrate the planning/execution split into multi-step agentic workflows (e.g., embodied AI or tool-use agents), where the planner generates a high-level action sequence and the executor performs individual steps.

  • Capabilities: The system will enable long-horizon, multi-step agent tasks to run at a fraction of current costs. It will allow agents to execute routine, low-stakes actions using local models while reserving expensive System-2 reasoning for critical decision points or high-uncertainty environmental changes.

Sources

Related papers