AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists

arXiv:2506.08140 · cs.LG, cs.CL · Submitted 2025-06-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists".

Jane: Building AI co-scientists that assist in open-ended scientific discovery remains challenging due to a scarcity of high-quality, large-scale data for training and evaluation.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to wrap up this part, we’ve looked at how AutoSDT automatically builds a massive dataset from real scientific code tasks to help train AI co-scientists.

Jane: Exactly, and the authors are calling this AutoSDT-5K "the only" automatically collected and largest open dataset for data-driven scientific discovery.

Lu: I think the real power of this paper lies in how they’re using LLMs to create that training material, moving beyond just consuming existing code. It opens up a whole new way to fuel model training with domain-specific examples.

Meng: From an engineering standpoint, having this massive collection pipeline means we can finally stress-test how robust these AI agents are when they actually have to handle complex, multi-step scientific workflows. It moves us closer to practical deployment.

Lalam: I see this as a huge cultural shift; by providing open models with training data derived from actual scientist-authored code, we're democratizing access to sophisticated AI tools for scientific reasoning. It makes these co-scientists accessible to everyone working in those fields.

Tom: That cultural shift is huge, Jane; imagine researchers not having to spend months manually labeling data anymore because the AI can do it automatically.

Jane: It really is a massive reduction in manual labor, Tom; they've shown that this scaling of high-quality data directly translates into tangible performance improvements for models.

Lu: The authors’ focus on open-weight models is smart because it shows the potential for these kinds of large datasets to benefit the entire community, not just a few proprietary systems.

Meng: I think the impact is that we can finally see measurable gains in areas like hypothesis matching on DiscoveryBench, which was a seventeen point four percent relative improvement over their base LLM. That's a concrete result we can talk about.

Lalam: And that concrete result helps build trust in these systems because the data quality has been validated by subject matter experts, who found ninety-three percent of the tasks were scientifically authentic. That level of rigor is essential for building reliable AI co-scientists.

Tom: It's exciting to see a pipeline that manages everything from searching repositories to adapting code into clear scientific instructions, all driven by LLMs.

Jane: And the conclusion is that AutoSDT-5K provides the necessary foundation for open AI co-scientists by feeding them high-quality, domain-specific training data.

Lu: It really pushes the envelope on what we thought was possible when it comes to scaling data creation in complex scientific domains.

Meng: So, the next thing we need to watch is how they address those limitations, specifically around generating effective long chain-of-thought rationales at scale.

Lalam: That’s where the future work gets interesting; figuring out how to automate the creation of instance-specific evaluation scripts will be key for making these tools truly usable in research settings.

Conclusion: Tom: So, we’re wrapping up our look at AutoSDT, and it really boils down to this paper titled "AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists" by those researchers who built this pipeline.

Jane: It’s a big concept because they’re essentially showing us how to take messy, real-world scientific code and automatically turn it into the perfect training material for next-generation AI co-scientists.

Lu: The authors are really emphasizing that the focus isn't just on making models write code; it’s on equipping them with a structured way to actually discover new scientific knowledge using that code, which is a much deeper level of capability.

Meng: From my side, this means we can finally start testing how these open models handle multi-step scientific reasoning in ways that are currently very hard to measure. It moves the goalposts for what we expect from AI agents in research roles.

Lalam: I see the cultural impact here as really democratizing high-level scientific reasoning; if this dataset works, it means less reliance on massive proprietary datasets and more access to tools that can learn directly from authentic scientific practice.

Tom: That’s the core idea: moving AI from being a code generator to being a true partner in the discovery process. The paper shows how they systematically collected thousands of tasks across different fields—like bioinformatics and chemistry—to achieve this goal.

Jane: And what I find really compelling is their validation process; they didn't just collect data and call it done; they had nine subject matter experts check the quality, which gave us a high degree of confidence in the scientific accuracy of this massive AutoSDT-5K dataset.

Lu: That level of domain-specific validation is what really pushes the envelope on what we thought was possible when it comes to creating training data at scale for complex scientific domains.

Meng: And honestly, having a dataset that's validated by experts makes sense because it gives us a solid benchmark to test how well different open models perform on these discovery benchmarks. It moves the discussion from theoretical possibility to measurable performance gains.

Lalam: Ultimately, this work lays the groundwork for building true AI co-scientists by providing them with domain-specific training data derived from naturally occurring scientist-authored code, which is a vital step toward making open models truly useful in scientific collaboration.

Tom: Exactly, so this paper isn't just about a new dataset; it’s about creating the infrastructure to train AI that can genuinely assist scientists in finding new things.

Jane: It shows that leveraging existing AI capabilities can solve data scarcity issues in complex scientific domains by automating the creation of high-quality training material.

Lu: This points toward systems that can perform complex, multi-step scientific discovery tasks using open models, which opens up avenues for much more practical applications in research settings.

Meng: So if we look at the results mentioned earlier—like the performance gains on those benchmarks—it proves that this data translates into tangible improvements for models trying to match hypotheses better.

Lalam: And from my perspective, this work contributes significantly to improving the culture of AI development by providing a pathway for creating high-quality training materials automatically, which democratizes access to complex scientific reasoning.

Tom: So, in short, AutoSDT provides the necessary foundation for open AI co-scientists by feeding them domain-specific training data that is rigorously vetted and automatically collected from real science.

Jane: Precisely, and it demonstrates how we can move past just generating code to enabling AI agents to actually derive scientific insights through processing and analysis.

Lu: It really pushes the envelope on what we thought was possible when it comes to scaling data creation in complex scientific domains that require expert-level understanding.

Meng: So, the next thing we need to watch is how they address those limitations, specifically around generating effective long chain-of-thought rationales at scale.

Department of Computer Science and Engineering · College of Pharmacy · Department of Psychology · Department of Biomedical Informatics · Department of Geography · Chemistry Department at University of Wisconsin–Madison (Cisco Research) · The Ohio State University

cs.LG, cs.CL

Submitted: 2025-06-09

Updated: 2026-09-27

Code: https://github.com/bndr/pipreqs

Project page: https://osunlp-group.github.io/AutoSDT

Importance score: 92/100

The gist: Building AI co-scientists that assist in open-ended scientific discovery remains challenging due to a scarcity of high-quality, large-scale data for training and evaluation.

Key concepts

AutoSDT Pipeline
This is an automated system designed to gather high-quality coding tasks for scientific discovery. It uses LLMs to search code repositories, filter relevant files based on scientific criteria, and then adapt the code into executable programs with clear instructions.
AutoSDT5K Dataset
This is the largest open dataset created by AutoSDT, containing 5,404 coding tasks spanning Bioinformatics, Chemistry, GIS, and Psychology. It serves as a high-quality training resource for developing AI co-scientists focused on data-driven scientific discovery.
AutoSDT-Search Stage
This initial stage uses LLMs to find relevant code by starting with keywords and expanding them. It then checks repositories on platforms like GitHub and PapersWithCode to ensure the found code actually relates to the target scientific discipline based on its documentation.
Open AI Co-scientists
These are AI models fine-tuned using high-quality, domain-specific data derived from real scientist code. The goal is to create powerful open models capable of assisting in open-ended scientific discovery tasks.

Terminology

Summary

Building AI co-scientists that assist in open-ended scientific discovery remains challenging due to a scarcity of high-quality, large-scale data for training and evaluation. This paper introduces AutoSDT, an automatic pipeline designed to scale data-driven discovery tasks by collecting high-quality coding tasks from real-world workflows, ultimately leading to the creation of AutoSDT5K, the largest open dataset for this domain.

The gist

AutoSDT is an automatic pipeline that collects highquality coding tasks in real-world data-driven discovery workflows by leveraging the coding capabilities and parametric knowledge of LLMs to search for diverse sources, select ecologically valid tasks, and synthesize accurate task instructions and code solutions.

Data Collection Pipeline (AutoSDT)

The AutoSDT pipeline is designed to overcome the limitations of manual annotation by automating the process through three main stages:

  1. AutoSDT-Search: This stage initiates by searching for code repositories containing programs for data-driven discovery tasks, starting with user-provided Seed keywords. The system expands these keywords using an LLM2 to generate a comprehensive set of related search queries, significantly improving coverage. It then queries GitHub and PapersWithCode APIs and uses an LLM to judge whether each repository hosts code related to the targeted discipline based on the README.md file.

  2. AutoSDT-Select: After crawling Python files from identified repositories, rule-based filtering removes files exceeding 1,000 lines or located in directories unlikely to contain substantive tasks (e.g., “config” and “tests”). LLM-based filtering then assesses the relevance of the remaining source code files by checking if they meet three criteria: functionality related to scientific workflows (like model training), utilization of datasets as input, and generation of scientific outputs. Finally, LLMs analyze file content and repository structure to extract all necessary dependencies.

  3. AutoSDT-Adapt: This final stage creates the "" pairs by adapting the identified code snippets into independently executable programs and generating corresponding task instructions. Program adaptation involves prompting Claude-3.7-Sonnet to make minimal modifications to ensure executability without altering core functionality, followed by dependency extraction using pipreqs4. Task instruction generation prompts an LLM to back-translate the adapted program into a clear, domain-specific scientific language instruction detailing the goal, required inputs/models, and expected outputs.

Dataset Construction and Quality Assurance (AutoSDT-5K)

The pipeline is used to construct AutoSDT5K, a dataset comprising 5,404 coding tasks covering four scientific disciplines: Bioinformatics, Computational Chemistry, Geographical Information Science, and Psychology and Cognitive Neuroscience. This dataset covers 756 unique Python packages. The quality of the collected tasks was rigorously validated by engaging nine subject matter experts to examine a subset of 256 tasks. Experts reported that 93% of the tasks are scientifically authentic and 92.2% of the generated programs are functionally correct, validating the high quality of AutoSDT-5K. The dataset covers diverse subtasks, ranging from data transformation, model training, visualization to more advanced analytics.

Model Performance and Scaling

Training Qwen2.5-Coder-Instruct series models on AutoSDT-5K resulted in substantial performance gains on two challenging benchmarks: ScienceAgentBench and DiscoveryBench. Specifically, AutoSDT-Coder-32B reached the same level of performance as GPT4o on ScienceAgentBench with a success rate (SR) of 7.8%, doubling the performance of its base model. On DiscoveryBench, it lifted the hypothesis matching score to 8.1, bringing a 17.4% relative improvement over its base LLM and closing the gap between open-weight models and GPT-4o. The performance gains are evident across model sizes; while smaller models show saturation, the 32B model continues to benefit from increased training data.

Conclusion and Impact

AutoSDT-5K is presented as the only automatically collected and the largest open dataset for data-driven scientific discovery. By fine-tuning models on this dataset, AutoSDT can propel the advancement toward open AI co-scientists by providing them with high-quality, domain-specific training data derived from naturally occurring scientist-authored code. This work demonstrates how scaling high-quality data can significantly improve the capabilities of open models for scientific tasks.

Limitations and Future Work

The paper identifies several limitations, including the lack of evaluation scripts for each code solution, which limits usability in settings like reinforcement learning. Furthermore, generating effective long chain-of-thought (CoT) rationales at scale remains a challenge. Future work suggested includes implementing an automatic framework to generate instance-specific evaluation scripts and exploring agent frameworks such as OpenHands CodeAct and self-debug for further research.

Improvements for AI systems

Here are specific improvements that can be made to AI systems by leveraging the findings and methodology of AutoSDT, along with a description of what these improved systems can achieve:


  1. The development of a robust, scalable pipeline for automatically collecting high-quality data-driven scientific discovery tasks (AutoSDT) allows for the creation of training datasets orders of magnitude larger than current manual curation methods.

  2. This enables the training of AutoSDT-Coder series models (e.g., AutoSDT-Coder-32B) that achieve state-of-the-art performance on complex scientific benchmarks like ScienceAgentBench and DiscoveryBench, rivaling proprietary models like GPT-4o in specific tasks.

  3. The system can be fine-tuned to exhibit superior generalization across disciplines (e.g., cross-disciplinary performance), allowing a single model to tackle a wider range of scientific problems by leveraging shared computational tools and libraries across different fields (Bioinformatics, Chemistry, Geo. Info. Science).

  4. The pipeline provides automated methods for generating highly accurate task instructions and validating code solutions through multiple LLM adaptation rounds, ensuring the generated training data is both ecologically valid (scientifically authentic) and functionally correct (standalone executability).

A system improved by these techniques can perform the following specific tasks:

  1. It can function as a powerful, open-weight AI co-scientist capable of performing complex, multi-step scientific workflows autonomously.

  2. It can generate complete, executable Python programs for data analysis (e.g., processing genomic data, performing molecular simulations based on chemical fingerprints, or analyzing geospatial imagery) directly from high-level natural language prompts describing the scientific goal.

  3. It can perform sophisticated hypothesis generation and analysis in open-ended domains by matching scientific queries to relevant code solutions (as demonstrated on DiscoveryBench), significantly closing the gap with proprietary models in this area.

  4. It can serve as a reliable tool for researchers to rapidly prototype scientific methods by generating verified, domain-specific code solutions, drastically reducing the time spent on manual task annotation and debugging.

Sources

Related papers