S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation
cs.LG, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted by EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Spectroscopic structure elucidation is central to molecular analysis, but recent Large Language Model (LLM)-based methods mostly formulate it as direct spectrum-to-SMILES generation.
Terminology
Abstract
Spectroscopic structure elucidation is central to molecular analysis, but recent Large Language Model (LLM)-based methods mostly formulate it as direct spectrum-to-SMILES generation. Although this paradigm can leverage paired spectral data, it does not explicitly model the analytical workflow used by spectroscopists, such as diagnostic peak interpretation, fragment reasoning, formula constraints, and chemical consistency checking. In this paper, we introduce S3C-LLM, a skill-guided and code-grounded agentic LLM for spectrum-to-structure elucidation. Rather than directly predicting a molecule, S3C-LLM retrieves modality-specific spectroscopy skills, executes analysis code to instantiate these skills on the input spectra, and integrates the resulting peak-level evidence and formula constraints before generating SMILES. Specifically, we contribute a self-evolving spectroscopy skill library, a thinking-augmented skill-code trajectory construction pipeline, and a two-stage training strategy that teaches Qwen3-4B through supervised fine-tuning (SFT) followed by our proposed step-level reinforcement learning (RL). Experiments on diverse benchmarks show that S3C-LLM consistently outperforms current general LLMs and spectrum-specific models across spectra, while using less than 1/10th of SpectraLLM's training corpus.
Sources
- Accurate and efficient structure elucidation from routine one-dimensional NMR spectra using multitask machine learning
- IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra
- Toolformer: Language Models Can Teach Themselves to Use Tools
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DiffSpectra: Molecular Structure Elucidation from Spectra using Diffusion Models
- LUMIR: an LLM-Driven Unified Agent Framework for Multi-task Infrared Spectroscopy Reasoning
- Qwen3 Technical Report
- SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources
- NMRTrans: Structure Elucidation from Experimental NMR Spectra via Set Transformers
- SpecMol: A Spectroscopy-Grounded Foundation Model for Multi-Task Molecular Learning
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks