Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
Junzhi Li, Peng He, Qirui Ji, Wei Wang, Lixiang Liu, Chuxiong Sun
University of Chinese Academy of Sciences · Institute of Software, Chinese Academy of Sciences · Beijing University of Posts and Telecommunications
cs.MA, cs.AI
Submitted: 2026-08-14
Updated: 2026-08-17
Comments: 18 pages, 4 figures. Submitted to AAAI 2027
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper proposes E2-Explainer, a model-agnostic post-hoc framework for identifying compact, task-preserving communication subgraphs from communication topologies generated by arbitrary LLM-based
Terminology
Summary
The paper proposes E2-Explainer, a model-agnostic post-hoc framework for identifying compact, task-preserving communication subgraphs from communication topologies generated by arbitrary LLM-based multi-agent system (MAS) topology designers. The work is motivated by the observation that the communication topologies generated by existing methods still contain substantial redundancy,
and that existing topology designers provide little insight into which communication substructures are truly responsible for successful collaboration.
The core formulation treats topology explanation as a causal attribution problem: we formulate topology explanation as a causal attribution problem that identifies compact communication subgraphs supported by edge-level evidence of task preservation.
The approach is inspired by Granger causality: if the information transmitted through an edge improves the cooperative performance of the MAS, that edge can be regarded as causally important to the resulting collaboration.
The method has four main steps. First, for each edge in a candidate graph, E2-Explainer performs single-edge removal interventions while keeping all other factors fixed
and measures the change in a task evaluation signal. The task-level causal contribution is defined as ∆task e = Q(x, G) − Q(x, G−e),
retaining only nonnegative values. Second, because the final task outcome can be too coarse to distinguish edge masks that receive the same task score,
the method adds an auxiliary semantic signal using result-level semantic entropy, computed as H̄(Y) = −Σ p(c) log p(c)
over semantic equivalence classes from M=5 stochastic executions. These are combined into a unified edge utility: ue = clip[0,1](α·ûtaske + β·ûseme)
with weights (α, β) = (0.8, 0.2).
Third, the edge utilities induce a preservation ordering, and budget-aware subgraph selection is performed: Eb⋆ = TopK(πG, kE(b))
followed by an executability-constrained projection
Rvalid that only selects edges from the original edge set E and never introduces new links,
preserving acyclicity when the original graph is a DAG. Applying this across a budget set yields an explanation family F⋆(x, G) = Hb⋆: b ∈ B.
Fourth, because directly applying this procedure to a newly generated topology at test time is computationally expensive,
the method distills the extracted subgraphs into an amortized explainer: we train a parameterized explainer Fθ to approximate the graph-to-subgraph mapping induced by the causal extraction procedure.
The explainer is trained with binary cross-entropy losses on edge and node retention labels, with objective L = Ledge + λnode Lnode + λwd ∥θ∥22.
At test time, it directly predicts a compact causal subgraph from a newly generated topology, avoiding repeated edge-masking evaluations.
Experiments are conducted on six benchmarks: AQuA, GSM8K, MultiArith, and SVAMP for mathematical reasoning, MMLU for knowledge-intensive reasoning, and HumanEval for code generation,
using Qwen3-8B as the backbone. The explainer is trained only on G-Designer-generated graphs and applied without retraining to ARG-Designer, OFA-MAS, and AgentPrune.
Key results include: Across all four optimizers, E2-Explainer reduces weighted token usage by 20.1%–25.6%, while preserving or improving their average accuracy.
Specifically, it improves average accuracy of ARG-Designer, OFA-MAS, AgentPrune, and G-Designer by 0.44, 0.53, 0.58, and 1.09 points respectively. G-Designer+E2-Explainer achieves the highest average accuracy of 90.71 while consuming 25.1% fewer tokens than the original G-Designer.
Across 24 optimizer–dataset combinations, E2-Explainer reduces token usage in every case, with dataset-level reductions ranging from 10.6% to 44.0%,
while accuracy improves in 16 combinations and remains unchanged in one.
The paper also demonstrates cross-scale generalization: training the explainer on 5-agent G-Designer graphs and directly evaluating it on 6-agent graphs
improves MMLU accuracy from 79.43 to 81.70 and HumanEval accuracy from 89.26 to 90.08, with token reductions of 10.16% and 19.38%. Component ablation shows that removing the causal signal leads to a larger degradation,
while the full model outperforms causal-only, indicating the result-level semantic entropy proxy complements causal supervision by densifying the refinement signal when final-task feedback is sparse.
Additional experiments show transfer to hand-crafted topologies (complete, random, layered, chain, star), where E2-Explainer improves or matches accuracy in 26 of the 30 dataset–topology combinations while reducing token usage across every topology and benchmark.
A budget sweep reveals that E25+N20 provides the most favorable overall trade-off, improving average accuracy from 89.63 to 90.71 while reducing weighted token usage by 25.10%,
while more aggressive node pruning (E25+N40) reduces tokens by 42.63% but decreases average accuracy to 88.71.
The paper concludes: "E2-Explainer distills offline edge-level evaluations into an amortized explainer, enabling efficient test-time generation of compact communication subgraphs that preserve collaborative task performance without requiring ground-truth answers or repeated edge-level evaluations. Experiments across multiple topology optimizers show that it reduces communication costs while maintaining task performance, and transfers across graph sources and agent scales."
Improvements for AI systems
Improvements to AI Systems:
-
Cost-Aware Multi-Agent Communication Pruning: Integrate E2-Explainer’s causal edge-attribution method (single-edge removal + task-score delta) as a post-hoc pruning layer for any LLM-based multi-agent system. The improved system dynamically removes redundant communication edges at inference time, reducing token consumption by 20–25% while preserving or improving task accuracy, enabling deployment in token-constrained or latency-sensitive environments.
-
Amortized Topology Optimization via Distilled Explainer: Train a lightweight graph-to-subgraph predictor (the amortized explainer Fθ) on offline causal evaluations, then deploy it to instantly generate compact communication subgraphs for newly designed topologies. This eliminates the need for expensive per-graph edge-masking trials, allowing real-time adaptation of agent communication structures without ground-truth answers.
-
Semantic-Entropy-Augmented Causal Attribution: Enhance the edge-importance scoring by combining task-level causal contribution (Δtask) with result-level semantic entropy (H̄(Y)) from stochastic executions. This dual-signal utility function (α=0.8, β=0.2) provides denser supervision when final-task feedback is sparse, enabling more precise identification of causally critical edges in complex reasoning tasks.
-
Budget-Aware Subgraph Selection with Executability Constraints: Implement the TopK-based budget sweep (e.g., E25+N20) to automatically select the optimal communication subgraph size per task. The improved system balances token savings (up to 25% reduction) against accuracy retention (e.g., 89.63→90.71), with a fallback to more aggressive pruning (E25+N40) for extreme cost reduction when slight accuracy loss is acceptable.
-
Cross-Scale and Cross-Architecture Generalization: Use the explainer’s demonstrated transferability (trained on 5-agent G-Designer graphs, applied to 6-agent graphs and other optimizers like ARG-Designer, OFA-MAS, AgentPrune) to build a universal communication pruner. The improved system can be trained once on one topology designer and then applied to unseen graph generators, agent counts, and hand-crafted topologies (complete, chain, star) without retraining, ensuring robust performance across diverse MAS configurations.
-
Interpretable Collaboration Debugging: Leverage the extracted compact subgraphs as explainable artifacts—showing which agent-to-agent communication links are causally responsible for task success. The improved system can highlight redundant or critical edges, enabling developers to audit and redesign MAS architectures for efficiency and reliability, and to detect failure modes where important information is being pruned.
Sources
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- SurrealDriver: Designing LLM-powered Generative Driver Agent Framework based on Human Drivers' Driving-thinking Data
- Adaptive Graph Pruning for Multi-Agent Communication
- Qwen3 Technical Report
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning