D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
cs.AI, cs.LG
Submitted: 2026-04-30
Updated: 2026-09-01
Code: https://github.com/OSU-NLP-Group/D3-Gym
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world
Terminology
Abstract
Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments representing real-world scientific tasks. To fill this gap, we introduce D3-Gym, the first automatically constructed dataset with verifiable environments for scientific Data-Driven Discovery. D3-Gym comprises 565 tasks from 239 real scientific repositories across four disciplines, each with a natural language instruction, an executable environment with pre-installed dependencies, dataset previews, a reference solution, and an automatically synthesized evaluation script. Our evaluation scripts achieve 87.5% agreement with human-annotated gold standards and strong alignment in domain-specific evaluation logic. Training on trajectories sampled from D3-Gym yields consistent gains across Qwen3 models on ScienceAgentBench, boosting Qwen3-32B by 7.8 absolute points and shrinking the gap with strong proprietary models. We further illustrate, through case studies, how D3-Gym environments can serve as a testbed for studying agentic optimization loops such as Autoresearch on real scientific workflows. We open-source D3-Gym, its creation workflow, sampled trajectories, and training scripts at https://github.com/OSU-NLP-Group/D3-Gym.
Sources
- LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
- HardTests: Synthesizing High-Quality Test Cases for LLM Coding
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- Training a Scientific Reasoning Model for Chemistry
- Language agents achieve superhuman synthesis of scientific knowledge
- Qwen3 Technical Report
- Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection