Reinforcement Learning of Communication in a Mesh of Small Language Models
cs.LG
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement Learning
- Training Verifiers to Solve Math Word Problems
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System
- Learning to Communicate with Deep Multi-Agent Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- LoRA: Low-Rank Adaptation of Large Language Models
- Byzantine-Robust Decentralized Coordination of LLM Agents
- Robust Multi-Agent LLMs under Byzantine Faults
- SwarmSys: Decentralized Swarm-Inspired Agents for Scalable and Adaptive Reasoning
- Let's Verify Step by Step
- Self-Organizing Agent Teams Learn to Reason Together
- Scaling Discovery through Test-Time Communication
- Benchmarking LLMs' Swarm intelligence
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LLM-based Multi-Agent Reinforcement Learning: Current and Future Directions
- QueenBee Planner: Skill-Evolving Communication Topologies for Token-Efficient LLM Multi-Agent Systems
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- AgentNet: Decentralized Evolutionary Coordination for LLM-based Multi-Agent Systems
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks