Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
The Hong Kong University of Science and Technology (Guangzhou) · Jilin University · University of Auckland · Shanghai Artificial Intelligence Laboratory · Hunan University · The Hong Kong University of Science and Technology
cs.LG, cs.AI
Submitted: 2026-08-14
Updated: 2026-08-17
Code: https://github.com/AIMS-Lab-HKUSTGZ/UltraIR
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: UltraIR is a foundation model for infrared (IR) spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules
Terminology
Summary
UltraIR is a foundation model for infrared (IR) spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. It learns a shared spectral representation by pretraining on approximately 60 million simulated IR spectra using three complementary pretraining objectives: wavelet-domain spectral reconstruction, molecular fingerprint similarity alignment, and functional-group prediction. The pretrained encoder is then adapted to each downstream objective using task-specific labels or targets paired with the corresponding IR spectra, typically with only a limited number of labeled spectra available.
The model uses a two-stage learning framework. In the first stage, a shared spectral encoder is pretrained on simulated spectra from three sources: approximately 1.5 million publicly released simulated spectra from IRtoMol, the multimodal spectroscopy dataset, and QM9S; approximately 7.5 million newly generated spectra from molecular-dynamics simulations; and approximately 51 million spectra generated using a machine-learning-based spectral predictor. Molecules present in the downstream molecular datasets were excluded from the pretraining pool to prevent molecule-level overlap. In the second stage, the pretrained encoder is paired with a newly initialized task-specific module, and all components are jointly optimized using labeled experimental IR spectra.
The encoder architecture combines a derivative-aware multi-channel input, hierarchical convolutional modules, and a patch-based Transformer. These components capture local line shapes and peak shifts, extract multiscale vibrational signatures, and model longer-range spectral dependencies, respectively. The pretraining objectives are designed to couple broad molecular exposure from simulated spectra with supervised adaptation using task-specific labeled experimental IR spectra.
UltraIR was evaluated across molecular and mixture benchmarks, as well as real-world chemical sensing applications. The benchmarks cover functional-group prediction, molecular structure elucidation, physicochemical property prediction, and mixture-component identification and quantification. The real-world applications include bacterial classification, medicinal-herb geographic origin traceability and constituent quantification using two newly generated in-house datasets, microplastics classification, and soil property prediction. Across these evaluations, UltraIR outperforms conventional machine-learning and task-specific deep-learning baselines.
For functional-group prediction, UltraIR consistently and substantially outperformed both task-specific deep-learning models (FCGFormer and IRAnalysis) and conventional machine-learning baselines (XGBoost, random forest, KNN, and logistic classifier) across NIST, SDBS, and USPTO datasets, evaluated using Micro-F1, Macro-F1, and exact match ratio (EMR). The most pronounced gains were in EMR on NIST and USPTO. UltraIR also consistently outperformed all competing methods at every training proportion on NIST and USPTO, and per-group F1-score rankings placed UltraIR consistently among the top-ranked methods across 17 functional groups.
For molecular structure elucidation, UltraIR achieved the strongest overall performance among the compared methods (IRtoMol, AISE, PBSA, and DLIR) across NIST, SDBS, and USPTO, evaluated using top-1, top-5, and top-10 accuracy. On the NIST test set, UltraIR produced candidates with higher, equal, and lower fingerprint-based Tanimoto similarity than IRtoMol for 44%, 41%, and 15% of spectra, respectively. Analysis of correct-case overlap showed that subsets including UltraIR accounted for 52.0% of test samples, whereas subsets excluding UltraIR accounted for only 2.9%. Representative examples showed UltraIR recovering exact structures with Tanimoto similarity of 1.00, while competing methods achieved similarities as low as 0.20.
For physicochemical property prediction, UltraIR achieved the lowest normalized MAE and normalized RMSE, together with the highest R2 among the evaluated methods (XGBoost regression, SVR, KNN regression, and PLSR) across NIST, SDBS, and USPTO for 11 structure-derived molecular descriptors. Property-wise R2 rankings placed UltraIR first in all 30 property–dataset combinations. For BertzCT prediction across cumulative molecular-complexity thresholds, UltraIR achieved the lowest mean relative error at every threshold.
For mixture analysis, UltraIR achieved strong targeted component-detection performance across NIST, SDBS, and USPTO, outperforming reverse match and HQI, and performing comparably to or better than DeepMIR. For quaternary mixtures, UltraIR established a clearer advantage in both accuracy and Macro-F1. For targeted fractional contribution estimation, UltraIR achieved the lowest MAE and RMSE and the highest R2 among all evaluated methods, with R2 values reaching 0.956 on NIST, 0.986 on SDBS, and 0.996 on USPTO. For mixture-level component quantification, UltraIR produced residuals most tightly concentrated around zero and achieved the highest R2 for each of the four quantified components: acrylonitrile, adiponitrile, propionitrile, and glycerol.
For bacterial classification, UltraIR achieved the highest accuracy, Macro-F1, and MCC among all compared methods (XGBoost, RF, KNN, and logistic classifier), and consistently outperformed the conventional baselines across all nine evaluated genera. For medicinal-herb characterization, UltraIR achieved the highest mean accuracy, Macro-F1, and MCC on both Jinyinhua and Shanyinhua for geographic origin traceability, outperforming all conventional baselines. For chemical constituent quantification, UltraIR achieved the highest R2 for every evaluated constituent in both Jinyinhua and Shanyinhua, with particularly large gains over the no-pretraining ablation.
For microplastics classification, UltraIR achieved the strongest performance among all evaluated methods (Softmax, DB-CNN-CBAM, XGBoost, RF, KNN, and logistic classifier) across accuracy, Macro-F1, and MCC. UltraIR corrected substantially more baseline errors than it introduced, and maintained consistently high performance across all 18 polymer categories. For soil property prediction, UltraIR achieved the strongest overall performance across the three aggregate regression metrics, and achieved the highest R2 for all ten soil properties. In training-data scaling experiments, UltraIR retained stronger performance than task-specific neural baselines as labeled supervision decreased.
The practical value of UltraIR’s simulation-to-real transfer learning was demonstrated under two demanding conditions: downstream adaptation with limited labeled experimental spectra and zero-shot inference for the same analytical task across FTIR spectrometers and laboratories. In the cross-instrument and cross-laboratory case study using soil property prediction, UltraIR achieved the highest R2 for all soil texture and acidity properties and for most soil chemical properties, and was the only method to maintain positive R2 across all evaluated properties.
UltraIR still requires supervised adaptation and a task-specific output module for each analytical objective, and a simulation-to-real domain gap remains because simulated IR spectra cannot fully reproduce experimental variation. Future adaptation strategies could exploit unlabeled experimental IR spectra to reduce the need for labeled experimental spectra. Overall, UltraIR establishes a scalable framework for simulation-to-real transfer learning in foundation modeling for IR spectroscopy, providing a route beyond collections of independently trained task-specific models toward adaptable and data-efficient IR-based chemical sensing systems built on reusable spectral representations.
Improvements for AI systems
Improvements to AI systems:
-
Add a wavelet-domain spectral reconstruction head to the pretraining objective, enabling the encoder to learn multi-resolution spectral features that are robust to noise, baseline drift, and peak overlap—improving generalization to experimental IR data.
-
Integrate molecular fingerprint similarity alignment as a contrastive or regression loss during pretraining, forcing the latent space to organize spectra by chemical structure similarity, which improves zero-shot retrieval and structure elucidation for unseen molecules.
-
Add functional-group prediction as an auxiliary task during pretraining, creating a chemically interpretable latent space that can be directly probed for explainability and used for partial-label learning when full molecular structures are unavailable.
-
Implement a derivative-aware multi-channel input layer (raw absorbance, first derivative, second derivative) to amplify subtle peak shifts and line-shape changes, improving sensitivity for mixture quantification and property prediction.
-
Use a patch-based Transformer encoder after hierarchical convolutions to capture long-range spectral dependencies, enabling the model to relate distant functional-group signatures—critical for complex mixtures and large molecules.
-
Adopt a two-stage fine-tuning protocol where the pretrained encoder is frozen initially and only the task-specific head is trained, then unfreeze the encoder with a low learning rate—reducing overfitting when only a few labeled experimental spectra are available.
-
Enable simulation-to-real transfer with domain adaptation by adding an adversarial discriminator that distinguishes simulated from experimental spectra during fine-tuning, reducing the domain gap and improving cross-instrument/cross-laboratory robustness.
-
Support zero-shot inference for analytical tasks by training the encoder with a shared spectral embedding space that aligns spectra from different instruments and labs, allowing direct comparison and classification without retraining—demonstrated by positive R2 across all soil properties in cross-lab tests.
-
Build a reusable spectral foundation backbone that can be plugged into any downstream task (classification, regression, retrieval, quantification) with a lightweight task-specific head, eliminating the need for task-specific architectures and reducing training data requirements by orders of magnitude.
-
Add a mixture decomposition module that outputs both component identities and fractional contributions simultaneously, using the pretrained encoder’s ability to separate overlapping vibrational signatures—achieving R2 > 0.95 for quaternary mixtures.
What the improved AI system can do:
-
Predict 17+ functional groups from IR spectra with higher exact-match ratios than task-specific deep learning models, even with only 10% of training data.
-
Elucidate molecular structures from IR spectra with top-1 accuracy exceeding existing methods, recovering exact structures (Tanimoto similarity = 1.00) where baselines fail (0.20).
-
Predict 11 physicochemical properties (e.g., BertzCT, molar refractivity) with the lowest error across all 30 property–dataset combinations, maintaining accuracy even for highly complex molecules.
-
Identify and quantify components in quaternary mixtures with near-perfect R2 (0.956–0.996) and tightly concentrated residuals, outperforming spectral library matching.
-
Classify bacteria into 9 genera, trace medicinal-herb geographic origins, quantify chemical constituents, classify 18 microplastic polymer types, and predict 10 soil properties—all with a single pretrained encoder and minimal labeled data.
-
Transfer across FTIR instruments and laboratories without retraining, maintaining positive predictive performance (R2 > 0) for all soil properties, a capability absent in all baselines.
-
Operate in data-scarce regimes: with as few as 10–50 labeled experimental spectra per task, it outperforms fully supervised task-specific models trained on hundreds of samples.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks