EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
cs.CL, cs.AI
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/TransformerLensOrg/TransformerLens
Terminology
Sources
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
- ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
- TORAX: A Fast and Differentiable Tokamak Transport Simulator in JAX
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- On the sensitivity of different galaxy properties to warm dark matter
- A foundation model of vision, audition, and language for in-silico neuroscience
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Training AI Scientists to Replicate Research
- BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
- The DREAMS Project: Disentangling the Impact of Halo-to-Halo Variance and Baryonic Feedback on Milky Way Dark Matter Density Profiles
- Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
- Accumulating Context Changes the Beliefs of Language Models
- MaD Physics: Evaluating information seeking under constraints in physical environments
- DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
- LLM-AutoSciLab: Closed-Loop Scientific Discovery via Active Experimentation with LLMs
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
- Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents
- EXP-Bench: Can AI Conduct AI Research Experiments?
- Linking Warm Dark Matter to Merger Tree Histories via Deep Learning Networks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering