Probing Chemical Language Models: Effects of Pre-training and Fine-tuning
cs.LG
Submitted: 2026-07-02
Updated: 2026-09-08
Comments: Accepted at EMNLP 2026 (to appear)
Code: https://github.com/IBM/molformer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode.
Terminology
Abstract
Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode. To foster a better understanding of CLMs, we conduct a systematic study and probe for 78 molecular substructures across eight pre-trained and six randomly initialized models. We furthermore study how fine-tuning on chemical downstream tasks affects the learned representations of molecular substructures. Our results show that pre-training generally improves molecular structure awareness of CLMs, particularly in the upper layers. Moreover, randomly initialized models already encode ring structures well in the first layer. Our analysis on two chemical downstream tasks further reveals that, interestingly, fine-tuning affects task-relevant molecular substructures more than others, indicating that the changes in the representations follow chemical theory.
Sources
- ChemBERTa-2: Towards Chemical Foundation Models
- Molecular representation learning with language models and domain-relevant auxiliary tasks
- Probing Classifiers: Promises, Shortcomings, and Advances
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- The Llama 3 Herd of Models
- OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs
- What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks
- AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
- gpt-oss-120b & gpt-oss-20b Model Card
- BERT Learns (and Teaches) Chemistry
- MoleculeNet: A Benchmark for Molecular Machine Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks