CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy
cs.LG, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 9 pages
Code: https://github.com/littlepeachs/CoMPASS
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Accurate molecular property prediction requires both statistical reliability and chemical reasoning.
Terminology
Abstract
Accurate molecular property prediction requires both statistical reliability and chemical reasoning. Graph neural networks can be calibrated directly on labeled assays but remain limited by the coverage of their training data. Large language models (LLMs) can compare molecular evidence and articulate chemical rationales, yet are unreliable as standalone quantitative predictors. The central challenge is therefore to determine when an LLM should influence a calibrated model and by how much. Here we present CoMPASS, a retrieval-calibrated framework for small-large model collaboration. CoMPASS retains a graph attention network (GAT) as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreement-aware gate. Across six classification and two regression benchmarks, CoMPASS improves the GAT anchor in regions of correctable uncertainty while limiting LLM intervention in high-confidence regimes. Ablations show that the gains arise from validation-calibrated retrieval and bounded fusion rather than prompting alone. These results suggest that generative reasoning should augment calibrated prediction through evidence-grounded, controlled corrections rather than direct output replacement. Code is available at https://github.com/littlepeachs/CoMPASS.
Sources
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- Strategies for Pre-training Graph Neural Networks
- Fast Inference from Transformers via Speculative Decoding
- Qwen2.5 Technical Report
- Galactica: A Large Language Model for Science
- ChemLLM: A Chemical Large Language Model
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks