Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
cs.AI, cs.IR
Submitted: 2026-09-10
Updated: 2026-09-22
Comments: Project site: https://benchmark-radar.org/ Code: https://github.com/ktwu01/benchmark-radar
Code: https://github.com/ktwu01/benchmark-radar
Project page: https://benchmark-radar.org
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings
Terminology
Abstract
Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.
Sources
- Contextual Information Policy Optimization for Search Agents
- MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
- ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
- Position: Evaluation Scores Are Perishable Knowledge Claims
- Measuring Massive Multitask Language Understanding
- Measuring Mathematical Problem Solving With the MATH Dataset
- FinanceBench: A New Benchmark for Financial Question Answering
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- The Semantic Scholar Open Data Platform
- EXP-Bench: Can AI Conduct AI Research Experiments?
- Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
- Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection