WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
cs.AI, cs.LG
Submitted: 2026-09-23
Updated: 2026-09-28
Terminology
Sources
- Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design
- CAFE: A Compound-AI Factorial Evaluation Framework
- BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
- Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab
- HPOBench: A Collection of Reproducible Multi-Fidelity Benchmark Problems for HPO
- Large Language Models to Enhance Bayesian Optimization
- SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
- Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
- Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- Science sandboxes measure the scientific capability of AI agents
- DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
- Hierarchical Experimentalist Agents
- Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
- DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- ResearchGym: Evaluating Language Model Agents on Real-World AI Research
- InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection