Multi-Granular Rationale-Guided Molecular LLM for Property Prediction

arXiv:2608.10480 · cs.AI, cs.LG · Submitted 2026-08-11 · Read on arXiv

Junwoo Park, Minyoung Shin, Cheol Soon Lee, Sujee Lee

Sungkyunkwan University · Korea University

cs.AI, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 16 pages, 5 figures, 18 tables. Code: https://github.com/skku-aihclab/MR-MoL

Code: https://github.com/skku-aihclab/MR-MoL

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: MR-MoL is a multi-granular rationale-guided molecular LLM for property prediction.

Terminology

Summary

MR-MoL is a multi-granular rationale-guided molecular LLM for property prediction. It is, to the authors' knowledge, the first method to feed GNN-derived attributions into an LLM's prompt as evidence for property prediction. The model takes four inputs: a task instruction, a 1D SMILES sequence, a 2D molecular graph, and a multi-granular rationale. The rationale is built from a fine-tuned GNN that scores each substructure's contribution to its prediction through masking. The most influential substructures are serialized as a ranked, direction-tagged list indicating whether each substructure pushes the GNN's prediction higher or lower. The rationale spans three levels of granularity: Murcko scaffolds with their side chains, BRICS fragments, and functional groups.

The model uses two complementary information paths. The graph embedding path turns the 2D molecular graph into molecular tokens through a GNN encoder and a Q-Former projector. The rationale path masks substructures against a fine-tuned GNN and serializes the most influential ones as ranked items. Training happens in two stages. Stage 1 performs molecular graph-language alignment for the graph embedding path, where the target is a molecule description and the rationale plays no role. Stage 2 performs multi-task, rationale-guided instruction tuning for property prediction, where the target is the answer to the property prediction task and the rationale is added as task-conditioned evidence.

On eight MoleculeNet tasks, MR-MoL achieves the best overall results among generalist models and narrows the gap to specialist models tuned for each task. Among the generalist models, MR-MoL outperforms every baseline on six of the eight tasks, leading by more than 11 ROC-AUC points on BACE and by nearly 7 on SIDER. In regression, every other generalist exceeds 3.7 RMSE on ESOL against MR-MoL's 1.210. The two exceptions are HIV and Lipo, where nach0 leads. MR-MoL surpasses both specialist molecular LLMs, 3D-MolT5 and HIGHT, on four of the six classification tasks and on ESOL, and on SIDER it achieves the best result across all baselines. In regression the strongest GNN-based models retain an advantage on both tasks.

Five diagnostics confirm that the model reads the rationale rather than merely benefiting from its presence. Direction sensitivity: flipping only the direction tag in each rationale item causes ROC-AUC to fall below chance on classification datasets and the MCC to flip sign, showing the model follows the inverted tag. Rank sensitivity: removing the rank-1 item produces a prediction shift 1.6 to 6.2 times larger than removing a random lower-rank item, with all p < 10−5. Substructure sensitivity: replacing each item's substructure with a random one lowers ROC-AUC by 2 to 4 points on BACE and BBBP and raises RMSE on ESOL and Lipo. Chemical validity: the attributions reproduce known structure-property relationships, such as polar functional groups raising predicted solubility and lowering predicted lipophilicity, mirroring the textbook inverse relation captured by the general solubility equation. Correction of individual predictions: adding the rationale corrects wrong predictions, such as flipping a blood-brain barrier permeability prediction for nalidixic acid from Yes to the correct No.

The ablation shows that removing the rationale lowers performance on seven of the eight tasks, with the largest classification drops on ClinTox and BACE. Removing the graph also hurts most tasks, though typically by a smaller margin on classification. The rationale matters more on classification and the graph is the larger contributor on the regression tasks. The sole exception is HIV, where neither added channel helps.

Limitations include that the rationale inherits the quality of the source predictor and may carry misleading evidence, the three views are not exhaustive, and the rationale channel covers only classification and regression tasks, leaving tasks such as molecule captioning and reaction prediction beyond its current reach.

Improvements for AI systems

Improvements to AI Systems:

  1. Integrate multi-granular, direction-tagged rationales into any LLM’s prompt for property prediction tasks (classification and regression). The system can now generate and consume ranked lists of substructures (scaffolds, BRICS fragments, functional groups) with explicit “increases” or “decreases” direction tags, enabling the LLM to reason about causal contributions rather than pattern-matching.

  2. Add a dual-path architecture (graph embedding + rationale) with staged training to any molecular LLM. Stage 1 aligns graph embeddings with text descriptions; Stage 2 fine-tunes on property tasks with rationale as evidence. This improves performance on small-molecule datasets, especially where specialist models are unavailable—e.g., achieving 1.210 RMSE on ESOL (vs. >3.7 for other generalists) and leading by 11 ROC-AUC points on BACE.

  3. Implement a rationale-sensitivity diagnostic suite to verify that the model genuinely uses provided evidence. The system can now self-audit by (a) flipping direction tags (expect performance to drop below chance), (b) removing rank-1 vs. random items (expect 1.6–6.2× larger prediction shifts), and (c) replacing substructures with random ones (expect 2–4 point ROC-AUC drops). This ensures trustworthiness in high-stakes applications like drug discovery.

  4. Enable correction of individual predictions via rationale injection. The system can now take a wrong prediction (e.g., blood-brain barrier permeability for nalidixic acid) and flip it to the correct answer by adding the rationale to the prompt—useful for human-in-the-loop review and iterative refinement.

  5. Extend the rationale-generation pipeline to new tasks by fine-tuning a GNN on any property dataset, then serializing its attributions. The system can now produce chemically valid explanations that align with known structure-property relationships (e.g., polar groups increase solubility, decrease lipophilicity), enabling interpretable predictions for novel molecules.

  6. Combine graph and rationale channels adaptively—the system can learn when each channel matters (rationale more for classification, graph more for regression) and optionally disable a channel if it hurts (e.g., HIV task where neither helps), improving robustness across diverse datasets.

What the improved AI system can do:

  • Predict molecular properties (e.g., solubility, toxicity, BBB permeability) with accuracy approaching specialist models, while remaining a single generalist.

  • Provide human-readable, ranked, direction-tagged explanations for each prediction, verified via sensitivity checks.

  • Correct its own mistakes when given additional rationale evidence.

  • Automatically adapt its reliance on graph vs. rationale per task, and flag cases where added evidence is unhelpful.

  • Be applied to new property prediction tasks by simply fine-tuning a GNN and reusing the same LLM prompt template.

Sources

Related papers