AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
summary
The gist
Building AI co-scientists that assist in open-ended scientific discovery remains challenging due to a scarcity of high-quality, large-scale data for training and evaluation.
In short
AutoSDT is an automatic pipeline that collects high-quality coding tasks from real-world scientific workflows to build a massive dataset called AutoSDT5K. This process uses LLMs to search for relevant code, filter it, and format it into instruction-code pairs across four scientific fields. Fine-tuning models on this data significantly boosts their performance in complex discovery benchmarks.
Key concepts
- AutoSDT Pipeline
- This is an automated system designed to gather high-quality coding tasks for scientific discovery. It uses LLMs to search code repositories, filter relevant files based on scientific criteria, and then adapt the code into executable programs with clear instructions.
- AutoSDT5K Dataset
- This is the largest open dataset created by AutoSDT, containing 5,404 coding tasks spanning Bioinformatics, Chemistry, GIS, and Psychology. It serves as a high-quality training resource for developing AI co-scientists focused on data-driven scientific discovery.
- AutoSDT-Search Stage
- This initial stage uses LLMs to find relevant code by starting with keywords and expanding them. It then checks repositories on platforms like GitHub and PapersWithCode to ensure the found code actually relates to the target scientific discipline based on its documentation.
- Open AI Co-scientists
- These are AI models fine-tuned using high-quality, domain-specific data derived from real scientist code. The goal is to create powerful open models capable of assisting in open-ended scientific discovery tasks.
Terminology used across episodes
This episode discusses
- AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists · Paper Radio
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Accelerating scientific discovery with Co-Scientist
- The Llama 3 Herd of Models · Paper Radio
- Qwen2.5-Coder Technical Report
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
- RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
- SWE-smith: Scaling Data for Software Engineering Agents
The paper
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists · Read on arXiv
Department of Computer Science and Engineering · College of Pharmacy · Department of Psychology · Department of Biomedical Informatics · Department of Geography · Chemistry Department at University of Wisconsin–Madison (Cisco Research) · The Ohio State University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists".
Jane: Building AI co-scientists that assist in open-ended scientific discovery remains challenging due to a scarcity of high-quality, large-scale data for training and evaluation.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to wrap up this part, we’ve looked at how AutoSDT automatically builds a massive dataset from real scientific code tasks to help train AI co-scientists.
Jane: Exactly, and the authors are calling this AutoSDT-5K "the only" automatically collected and largest open dataset for data-driven scientific discovery.
Lu: I think the real power of this paper lies in how they’re using LLMs to create that training material, moving beyond just consuming existing code. It opens up a whole new way to fuel model training with domain-specific examples.
Meng: From an engineering standpoint, having this massive collection pipeline means we can finally stress-test how robust these AI agents are when they actually have to handle complex, multi-step scientific workflows. It moves us closer to practical deployment.
Lalam: I see this as a huge cultural shift; by providing open models with training data derived from actual scientist-authored code, we're democratizing access to sophisticated AI tools for scientific reasoning. It makes these co-scientists accessible to everyone working in those fields.
Tom: That cultural shift is huge, Jane; imagine researchers not having to spend months manually labeling data anymore because the AI can do it automatically.
Jane: It really is a massive reduction in manual labor, Tom; they've shown that this scaling of high-quality data directly translates into tangible performance improvements for models.
Lu: The authors’ focus on open-weight models is smart because it shows the potential for these kinds of large datasets to benefit the entire community, not just a few proprietary systems.
Meng: I think the impact is that we can finally see measurable gains in areas like hypothesis matching on DiscoveryBench, which was a seventeen point four percent relative improvement over their base LLM. That's a concrete result we can talk about.
Lalam: And that concrete result helps build trust in these systems because the data quality has been validated by subject matter experts, who found ninety-three percent of the tasks were scientifically authentic. That level of rigor is essential for building reliable AI co-scientists.
Tom: It's exciting to see a pipeline that manages everything from searching repositories to adapting code into clear scientific instructions, all driven by LLMs.
Jane: And the conclusion is that AutoSDT-5K provides the necessary foundation for open AI co-scientists by feeding them high-quality, domain-specific training data.
Lu: It really pushes the envelope on what we thought was possible when it comes to scaling data creation in complex scientific domains.
Meng: So, the next thing we need to watch is how they address those limitations, specifically around generating effective long chain-of-thought rationales at scale.
Lalam: That’s where the future work gets interesting; figuring out how to automate the creation of instance-specific evaluation scripts will be key for making these tools truly usable in research settings.
Conclusion: Tom: So, we’re wrapping up our look at AutoSDT, and it really boils down to this paper titled "AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists" by those researchers who built this pipeline.
Jane: It’s a big concept because they’re essentially showing us how to take messy, real-world scientific code and automatically turn it into the perfect training material for next-generation AI co-scientists.
Lu: The authors are really emphasizing that the focus isn't just on making models write code; it’s on equipping them with a structured way to actually discover new scientific knowledge using that code, which is a much deeper level of capability.
Meng: From my side, this means we can finally start testing how these open models handle multi-step scientific reasoning in ways that are currently very hard to measure. It moves the goalposts for what we expect from AI agents in research roles.
Lalam: I see the cultural impact here as really democratizing high-level scientific reasoning; if this dataset works, it means less reliance on massive proprietary datasets and more access to tools that can learn directly from authentic scientific practice.
Tom: That’s the core idea: moving AI from being a code generator to being a true partner in the discovery process. The paper shows how they systematically collected thousands of tasks across different fields—like bioinformatics and chemistry—to achieve this goal.
Jane: And what I find really compelling is their validation process; they didn't just collect data and call it done; they had nine subject matter experts check the quality, which gave us a high degree of confidence in the scientific accuracy of this massive AutoSDT-5K dataset.
Lu: That level of domain-specific validation is what really pushes the envelope on what we thought was possible when it comes to creating training data at scale for complex scientific domains.
Meng: And honestly, having a dataset that's validated by experts makes sense because it gives us a solid benchmark to test how well different open models perform on these discovery benchmarks. It moves the discussion from theoretical possibility to measurable performance gains.
Lalam: Ultimately, this work lays the groundwork for building true AI co-scientists by providing them with domain-specific training data derived from naturally occurring scientist-authored code, which is a vital step toward making open models truly useful in scientific collaboration.
Tom: Exactly, so this paper isn't just about a new dataset; it’s about creating the infrastructure to train AI that can genuinely assist scientists in finding new things.
Jane: It shows that leveraging existing AI capabilities can solve data scarcity issues in complex scientific domains by automating the creation of high-quality training material.
Lu: This points toward systems that can perform complex, multi-step scientific discovery tasks using open models, which opens up avenues for much more practical applications in research settings.
Meng: So if we look at the results mentioned earlier—like the performance gains on those benchmarks—it proves that this data translates into tangible improvements for models trying to match hypotheses better.
Lalam: And from my perspective, this work contributes significantly to improving the culture of AI development by providing a pathway for creating high-quality training materials automatically, which democratizes access to complex scientific reasoning.
Tom: So, in short, AutoSDT provides the necessary foundation for open AI co-scientists by feeding them domain-specific training data that is rigorously vetted and automatically collected from real science.
Jane: Precisely, and it demonstrates how we can move past just generating code to enabling AI agents to actually derive scientific insights through processing and analysis.
Lu: It really pushes the envelope on what we thought was possible when it comes to scaling data creation in complex scientific domains that require expert-level understanding.
Meng: So, the next thing we need to watch is how they address those limitations, specifically around generating effective long chain-of-thought rationales at scale.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language