MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification

arXiv:2605.19752 · cs.LG · Submitted 2026-05-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification".

Jane: Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research

50, 65: .

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're diving into the paper titled "MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification," and what they're doing is really tackling the core challenge of metabolite identification by merging two different machine learning approaches.

Jane: Exactly, Tom; it’s essentially proposing a unified framework that combines representation alignment with contrastive learning to solve the molecule retrieval task in metabolomics.

Lu: The authors are proposing MSAlign as their central innovation, which aims to learn a shared representation space by aligning two powerful foundation models: DreaMS for mass spectra and ChemBERTa for molecules.

Meng: I see; so they're leveraging these massive, already-trained models and just training small projection layers to make their outputs speak the same language; I have to ask how feasible that kind of fine-tuning is without needing immense computational power.

Lalam: It’s really about creating a common vocabulary for both modalities so the AI can understand them consistently, which will help improve our ability to process complex chemical information across different contexts in our future tools.

The paper's summary: Tom: Now that we know they're using alignment and contrastive learning, let's look at the paper’s summary of MSAlign in more detail; they introduce this method specifically to solve the molecule retrieval task by aligning DreaMS for spectra and ChemBERTa for molecules.

Jane: That means they are taking those two different data types—the raw spectra and the molecular structure representations—and mapping them into the same mathematical space so they can be compared meaningfully, which is about creating a common language between chemistry and physics.

Lu: The mechanism involves using lightweight MLP projections, specifically phi ms and phi mol, to map both modalities in that shared space, which is brilliant because it keeps the foundational models frozen while only training these small layers.

Meng: So, they're essentially taking those massive, already-trained models and just training small projection layers to make their outputs speak the same language; I have to ask how feasible that kind of fine-tuning is without needing immense computational power.

Lalam: It’s really about creating a common vocabulary for both modalities so the AI can understand them consistently, which will help improve our ability to process complex chemical information across different contexts in our future tools.

The paper's improvements: Tom: What makes this approach particularly impressive is their training objective; they found that candidate-based InfoNCE consistently beats other methods because it forces the model to learn from hard negatives during the search for a molecule, which is a big win for identification accuracy.

Jane: That’s smart thinking; using those hard negatives ensures the model isn't just learning to memorize easy matches but is actually learning to discriminate between true candidates and real distractors in that chemical search space.

Lu: Furthermore, they tackled a major issue in benchmarking by introducing a quantitative measure of distribution shift, which helps us understand the tension between data leakage and domain shift when splitting datasets for testing.

Meng: That distribution shift analysis is vital because if we don't understand how our data is split affecting generalization, we risk training models that work perfectly on the training set but fail spectacularly in a real lab environment.

Lalam: Understanding that tension between leakage and shift gives us a much better roadmap for designing experiments and building truly reliable tools for metabolite identification systems moving forward.

Conclusion: Tom: Wrapping things up, "MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification" shows incredible promise by creating a simple, fast framework that consistently outperforms existing methods across all benchmarks. It really confirms that aligning these two complex modalities using frozen foundation models isn't just theoretical; it’s yielding tangible improvements in the accuracy of metabolite identification.

Jane: It really confirms that aligning those two complex modalities using frozen foundation models isn't just theoretical; it’s yielding tangible improvements in the accuracy of metabolite identification. This means we can start expecting better results when we run our initial screening experiments.

Lu: The implications are huge for drug discovery and environmental analysis because if we can reliably identify molecules from raw spectra, the pace of discovery could accelerate dramatically, opening up entirely new avenues for research.

Meng: Practically speaking, the efficiency is also a big win here; by only training those lightweight MLP projection layers, they’ve managed to keep training times reasonable even when dealing with large datasets like Spectraverse.

Lalam: For our culture here, this paper reinforces the idea that combining powerful pre-trained knowledge with targeted learning is the way forward for building incredibly capable and reliable AI tools in science.

Paul Krzakala, Gabriel Melo, Camille Lançon, Charlotte Laclau, Rémi Flamary, Etienne Thévenot, Florence d’Alché-Buc

LTCI, Télécom Paris & CMAP, Ecole Polytechnique, Institut Polytechnique de Paris · CEA, INRAE, MetaboHUB, Université Paris-Saclay

cs.LG

Submitted: 2026-05-19

Updated: 2026-09-25

Code: https://github.com/HassounLab/JESTR1

Importance score: 89/100

The gist: Accurately identifying metabolites i.e.

Terminology

Summary

Accurately identifying metabolites i.e. small molecules from mass spectrometry data remains a core challenge in metabolomics, with broad applications in drug discovery, environmental analysis, and clinical research [50, 65]. The central challenge is the Molecule Retrieval task: given a mass spectrum and a set of candidate molecules, the goal is to retrieve the most likely molecular structure among the candidates.

The paper addresses this by proposing a unified framework encompassing recent approaches based on representation alignment and contrastive learning. The authors introduce MSAlign, which is inspired by multimodal alignment in vision-language models, aiming to learn a shared representation space by aligning two frozen foundation models: DreaMS for mass spectra and ChemBERTa for molecules. This is achieved through lightweight MLP projections trained with a candidate-based contrastive objective.

The contributions of the work are threefold:

  1. A unified framework encompassing recent approaches based on representation alignment and contrastive learning.

  2. MSAlign, which "learn[s] a shared representation space by aligning two frozen foundation models (DreaMS for mass spectra and ChemBERTa for molecules) through lightweight MLP projections trained with a candidate-based contrastive objective. MSAlign is described as being simple to implement, fast to train and consistently outperforms existing approaches across all benchmarks."

  3. An investigation into data splitting strategies, formalizing the tension between data leakage and domain shift by introducing a quantitative measure of distribution shift, which is used to evaluate splitting strategies in existing benchmarks.

The general framework for molecule retrieval is formulated as:

Given a dataset of paired observations (si, mi)N i=1 ∈ S × M, the goal of Metabolite Identification is to learn a model f: S → M that predicts the 2D molecular structure associated with a given spectrum. This problem is typically cast as Molecule Retrieval: given a mass spectrum and a set of candidate molecules, the goal is to retrieve the most likely molecular structure among the candidates. The scoring function and training loss are generally defined as:

We focus here on the large class of scoring functions ρθ defined as a similarity sim: Z × Z → R+ over a shared representation Z of spectra and molecules, i.e. ρθ (s, m) = sim(gθ (s), hθ (m))

The proposed MSAlign method utilizes the following components:

- Foundation models for MS/MS data:

On the spectral side, we use DreaMS [9], the first foundation model for mass spectra, pretrained on 24 million high-quality MS/MS spectra.

"On the molecular side, a large body of work has explored graph-based and text-based foundation models such as UniMol2 [35], MolE [45], ChemBerta [1], Grover [54] and GRALE [39]. ChemBERTa appears as a natural choice, both because it is simple to use and because it achieves strong performance on MoleculeNet benchmarks [66]."

- Architecture:

In MSAlign, Ems and Emol are the foundation models DreaMS and ChemBERTa, respectively, and ϕθms and ϕθmol are lightweight MLP projection layers that map both modalities in a shared space. This design leverages pretrained representations while remaining lightweight.

- Training Objective:

"The final key design choice is the training objective. Table 4, Figure 3 and Figure 5 show that candidate-based InfoNCE (5) consistently outperforms both regression-based approaches and inbatch contrastive learning, highlighting the importance of hard negatives for molecule identification." MSAlign makes this tractable by leveraging precomputed embeddings for fast similarity computation over candidate sets.

- Data Splitting Analysis:

The authors formalize the tension between data leakage and domain shift by introducing a quantitative metric: Shift(Dtrain, Dtest) = W (Dtrain, Dtest) / D2). They approximate the Wasserstein distance using the sliced Wasserstein with p = 100 projections and 5 random seeds, embedding each spectrum–molecule pair as E(s, m) = [Ems (s), Emol (m)] using DreaMS and ChemBERTa. The analysis on MassSpecGym showed that "MCES and Murcko splits induce large shifts (Shift > 4), while the Formula and 2D InChIKey splits provide a more balanced trade-off, with Shift values around 1-2," concluding that the Formula-split is sufficient to prevent data leakage while still reflecting a realistic deployment setting.

- Evaluation:

MSAlign consistently outperforms all baselines across datasets and metrics. Specifically, on MassSpecGym (Formula Split), MSAlign achieved R@1 scores of 16.2% and 53.8%, and on Spectraverse, it achieved "R@1 scores of 31.

Improvements for AI systems

Based on the provided research paper, here are specific improvements for AI systems and what those improved systems can achieve:


  1. Improving Metabolite Identification Accuracy via MSAlign:

  2. An AI system that uses MSAlign (aligning DreaMS for spectra and ChemBERTa for molecules) to perform molecule retrieval from mass spectrometry data.

  3. This system can identify the most likely molecular structure from a given mass spectrum, specifically when querying against large candidate databases like PubChem, outperforming existing methods by learning a shared representation space.

  4. Enhancing Robustness to Data Leakage and Domain Shift:

  5. An AI system that incorporates the proposed quantitative distribution shift measure (Wasserstein distance) into its training or evaluation pipeline.

  6. This improved system can intelligently select optimal data splitting strategies for molecule retrieval benchmarks, ensuring the model generalizes better to real-world deployment scenarios rather than overfitting to training data characteristics.

  7. Improving Computational Efficiency via Lightweight Foundation Models:

  8. An AI system that uses the MSAlign framework, leveraging frozen foundation models (DreaMS and ChemBERTa) and only training lightweight MLP projection layers (4M parameters).

  9. This system can achieve state-of-the-art performance (as shown in Table 7) while maintaining fast training times, making it suitable for rapid iteration on large datasets like Spectraverse.

  10. Optimizing Retrieval Loss via Hard Negative Mining:

  11. An AI system that employs the candidate-based InfoNCE loss (Equation 5), which explicitly uses hard negatives sampled from the candidate set, instead of simpler in-batch or regression losses.

  12. This improved system can achieve superior performance in metabolite identification by forcing the model to learn discriminative features between true matches and distractors within a constrained chemical search space.

  13. Integrating Spectral Metadata for Enhanced Prediction:

  14. An AI system that incorporates learnable embeddings for missing or variable metadata (like adduct type or collision energy) into the spectral encoder representation.

  15. This system can achieve modest but measurable performance gains (e.g., +2% improvement in R@1 on MassSpecGym) by better capturing fragmentation patterns influenced by experimental conditions, leading to more accurate identification under diverse MS/MS conditions.

  16. Enabling Expert-Based Inference:

  17. An AI system that utilizes the Mixture of Finetuned Experts strategy (Table 9) for inference, where different specialized models are selected at runtime based on the known adduct type of a given spectrum.

  18. This system can offer superior performance across different adduct classes by leveraging domain-specific knowledge learned during training, potentially yielding higher accuracy than a single monolithic model when operating in complex, real-world MS/MS environments.

  19. Improving Molecular Representation Choice:

  20. An AI system that utilizes ChemBERTa (Graph/Transformer-based) as the molecular encoder rather than simple fingerprints or binned representations.

  21. This improved system can achieve better recall across various metrics (Table 10) by capturing richer structural and contextual information directly from SMILES/graph structures, leading to more chemically nuanced predictions.

Sources

Related papers