LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
cs.AI
Submitted: 2026-06-19
Updated: 2026-09-02
Code: https://github.com/galaxyChen/LivingArena
License: http://creativecommons.org/licenses/by/4.0/
The gist: Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures.
Terminology
Abstract
Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We ask whether LLMs can instead discover one another's weaknesses and turn those observations into an evaluation process. To study this question, we introduce LivingArena, an automated peer-probing framework in which models take turns testing one another. Using the interaction history, each questioner identifies potential weaknesses of its opponent and constructs targeted, verifiable questions to probe them. A 3,600-round tournament of ten models reveals a clear role asymmetry: strong answerers are not always reliable questioners, because they may generate internally inconsistent tests or fail to verify their own reference answers. After a questioner exposes an answerer's failure, it is more likely to pursue the same capability domain, while the answerer's weakness recurs on independently generated questions, including questions written by different models. These findings show that peer probing can reveal persistent model-specific weaknesses while separately evaluating answering and reliable test construction. By automating this process and allowing test difficulty to evolve with model capabilities, LivingArena provides a "living" benchmark for model development, red-teaming, and capability-aware multi-agent coordination. We publicly release our code: https://github.com/galaxyChen/LivingArena
Sources
- AutoBench: Automating LLM Evaluation through Reciprocal Peer Assessment
- QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks
- PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review
- CodeArena: A Collective Evaluation Platform for LLM Code Generation
- Agent-as-a-Judge
- Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection