Science sandboxes measure the scientific capability of AI agents
q-bio.QM, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/asr2210/science-sandbox
Project page: https://openai.com/index/introducing-gpt-5-5
Terminology
Sources
- Empowering Biomedical Discovery with AI Agents
- DiscoveryBench: Towards Data-Driven Discovery with Large Language Models
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- Iterative Foundation Model Fine-Tuning on Multiple Rewards
- Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design
- Life After Benchmark Saturation: A Case Study of CORE-Bench
Related papers
- A likelihood-based framework for simultaneously learning both noise and growth dynamics using biologically-informed neural networks
- Automated Lesion Segmentation of Stroke MRI Using nnU-Net: A Comprehensive External Validation Across Acute and Chronic Lesions
- Resolving satellite-in situ mismatches in Net Primary Production using high-frequency in situ bio-optical observations in the subpolar Northwest Atlantic
- easyplater: The easy way to generate microplate designs deconvolved from multivariate clinical data
- Essential Workers at Risk: An Agent-Based Model (SAFE-ABM) with Bayesian Uncertainty Quantification
- OmniBioTwin: A System-of-Twinned-Systems Framework for Health Digital Twins