Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry
cs.CL, cs.AI
Submitted: 2025-10-29
Updated: 2026-09-17
Comments: COLM 2026
Code: https://github.com/Roihn/EinsteinPuzzles
License: http://creativecommons.org/licenses/by/4.0/
The gist: While Large Language Model (LLM) agents are often approached from the angle of action planning/generation to accomplish a goal (e.g., given by language descriptions), their abilities to collaborate
Terminology
Abstract
While Large Language Model (LLM) agents are often approached from the angle of action planning/generation to accomplish a goal (e.g., given by language descriptions), their abilities to collaborate with each other to achieve a joint goal are not well explored. To address this limitation, this paper studies LLM agents in task collaboration, particularly under the condition of information asymmetry, where agents have disparities in their knowledge and skills and need to work together to complete a shared task. We extend Einstein Puzzles, a classical symbolic puzzle, to a table-top game. In this game, two LLM agents must reason, communicate, and act to satisfy spatial and relational constraints required to solve the puzzle. We apply a fine-tuning-plus-verifier framework in which LLM agents are equipped with various communication strategies and verification signals from the environment. Empirical results highlight the critical importance of aligned communication, especially when agents possess both information-seeking and-providing capabilities. Interestingly, agents without communication can still achieve high task performance; however, further analysis reveals a lack of true rule understanding and lower trust from human evaluators. Instead, by integrating an environment-based verifier, we enhance agents' ability to comprehend task rules and complete tasks, promoting both safer and more interpretable collaboration in AI systems. https://github.com/Roihn/EinsteinPuzzles
Sources
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
- Training Verifiers to Solve Math Word Problems
- Agent AI: Surveying the Horizons of Multimodal Interaction
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- The Llama 3 Herd of Models
- Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
- Autonomous Agents for Collaborative Task under Information Asymmetry
- LLM Critics Help Catch LLM Bugs
- Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- Tree Search for Language Model Agents
- Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
- Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
- Sibyl: Simple yet Effective Agent Framework for Complex Real-world Reasoning
- OASIS: Open Agent Social Interaction Simulations with One Million Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering