An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific Process
cs.SE, cs.AI
Submitted: 2026-03-13
Updated: 2026-08-28
Comments: 24 pages (6 pages prose, 6 pages citations, 12 pages appendix), 9 figures, 2 tables, accepted as a short paper to the 22nd IEEE International Conference on eScience
Code: https://github.com/jgm/pandoc
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deep learning has achieved recognition for its impact within natural sciences, yet the prohibitive financial and technical cost of training models from scratch inhibit adoption.
Terminology
Abstract
Deep learning has achieved recognition for its impact within natural sciences, yet the prohibitive financial and technical cost of training models from scratch inhibit adoption. Following software engineering community guidance, natural scientists are reusing pre-trained deep learning models (PTMs) to amortize these costs. While prior works recommend PTM reuse patterns, we present the first empirical study of PTM reuse patterns in the natural sciences, quantifying the utilization and impact of PTM reuse within the scientific process across 17,718 peer reviewed, open access papers. Our results show that "Biochemistry, Genetics and Molecular Biology" has outpaced other natural scientific fields in PTM reuse, "adaptation" reuse is the most prevalent PTM reuse pattern identified across all natural science fields, and the "testing" stage of the scientific process has been most impacted by PTM integration.
Sources
- Oil & Water? Diffusion of AI Within and Across Scientific Fields
- Analysis of Failures and Risks in Deep Learning Model Converters: A Case Study in the ONNX Ecosystem
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- SciCat: A Curated Dataset of Scientific Software Repositories
- OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts
- PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
- De novo design of high-affinity protein binders with AlphaProteo
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties