BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks
cs.AI, q-bio.QM
Submitted: 2026-08-26
Updated: 2026-09-22
Code: https://github.com/UKGovernmentBEIS/inspect_ai
Terminology
Sources
- SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
- Verifiable Benchmarking of Long-Horizon Spatial Biology
- scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
- GenoTEX: An LLM Agent Benchmark for Automated Gene Expression Data Analysis
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- Benchmarking AI scientists for omics data driven biological discovery
- Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- Kosmos: An AI Scientist for Autonomous Discovery
- HCAST: Human-Calibrated Autonomy Software Tasks
- BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection