Localize, Then Reason: Visual Latent Structural Reasoning for Molecular Properties and Edits

arXiv:2608.13244 · cs.CL, cs.CE, q-bio.BM · Submitted 2026-08-13 · Read on arXiv

Xingqiao Lin, Junmei Wang, Haocheng Tang

Carnegie Mellon University · University of Pittsburgh · Northeastern University

cs.CL, cs.CE, q-bio.BM

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: The paper introduces Visual Latent Structural Reasoning (VLSR), an end-to-end framework for molecular property reasoning from molecular images that follows a "localize-then-reason" strategy.

Terminology

Summary

The paper introduces Visual Latent Structural Reasoning (VLSR), an end-to-end framework for molecular property reasoning from molecular images that follows a localize-then-reason strategy. The authors argue that current LLM-based chemical reasoning methods either provide explicit structural annotations (e.g., functional groups) as inputs, which reduces the model's need to discover chemically relevant regions, or reason directly from molecular images without focusing on chemically meaningful areas first. VLSR addresses this gap by first learning to locate chemically meaningful regions in a molecular image and then reasoning about their property effects in a compact latent workspace before producing the final answer.

The problem is formulated as: given a natural-language query q and either one molecular image IA or a pair (IA, IB), the model predicts an answer y (binary, directional, categorical, or continuous). The model must derive local structure R from the images before estimating pθ(y IA, IB, q, R). RDKit-derived boxes and labels (functional groups, ring systems, hetero-atom regions) are used as training targets only, not as inference inputs.

The VLSR pipeline is: (I, q) or (IA, IB, q) → Z → (R, E) → L → y, where Z is image tokens, R is localized regions, E is edit evidence (for paired inputs), and L is latent reasoning states. The method uses a Qwen-VL vision encoder to map depictions to patch tokens. Trainable region queries retrieve local chemical evidence through cross-attention, producing region tokens that integrate multi-patch motifs and ring context. A Hungarian one-to-one assignment with a lightweight one-to-many hybrid branch supervises localization during training, but boxes and labels are not passed to the reasoning core and are absent at inference.

For paired questions, VLSR uses bidirectional soft alignment rather than index-wise patch subtraction, computing differences, products, and learned pooling to preserve additions and removals without requiring atom maps. Three latent tokens—⟨LAT EDIT⟩, ⟨LAT PROP⟩, and ⟨LAT EFFECT⟩—organize structural evidence, the queried property, and its predicted effect. The reasoning core performs M Transformer updates over question states, image tokens, region tokens, optional edit tokens, and the three latent states, then decodes only the final answer without generating intermediate textual rationales.

Training uses three stages. Stage I is OCSR-style continued pre-training on one million synthetic PubChem molecules and 680K patent examples with a frozen vision encoder. Stage II is supervised fine-tuning on FGBench-Scaffold, a pair-aware scaffold-component split of FGBench with 625,936 examples (507,000 training, 56,348 validation, 62,588 test). The split uses Bemis–Murcko scaffold components, an ECFP4 similarity threshold of 0.7, and prevents component overlap between training and test sets. LoRA jointly optimizes answer prediction and auxiliary region localization. Stage III applies GRPO with answer-level rewards (correctness, regression closeness, sign consistency, output cleanliness) without accessing annotations, attention maps, or latent states.

The supervised objective combines answer loss with auxiliary losses: Ltotal = Lanswer + 0.05Lce attn + 0.05Liou attn + 0.01Ldiv + 0.02Lent + 0.02Lfg + 0.02Lce hyb + 0.02Liou hyb + 0.01Lfg hyb, where attention cross-entropy and soft-IoU supervise coverage, diversity and entropy regularize maps, and Lfg predicts region type.

Main results on FGBench-Scaffold show VLSR achieves the best overall performance among image-input methods and outperforms the strongest symbolic baseline. Relative to Qwen3.5-4B + MSR-SFT, accuracy rises from 0.767 to 0.842, F1 from 0.679 to 0.786, and balanced accuracy from 0.745 to 0.829. Regression also improves. VLSR achieves 39.84 samples/s inference throughput compared with 4.13 samples/s for image-input Qwen3.5-4B-SFT with textual reasoning, a 9.6× throughput improvement.

Controlled experiments test whether learning where to look explains the gain. Localization analysis shows learned queries outperform random regions for functional groups, rings, and hetero-atom regions on held-out molecules, and distinct queries attend to different annotated structures within the same molecule. Occlusion intervention masks either the attention-ranked proposal or a size-matched random region at inference: learned-region masking reduces accuracy to 0.789 and raw-pooled Pearson correlation to 0.576, compared with 0.791 and 0.626 for random masks averaged over 20 deterministic draws. The larger Pearson drop under learned-region masking indicates proposals carry more regression-relevant evidence than equally sized random regions.

Relational reasoning controls replace learned image regions with RDKit-derived region descriptors and their relative spatial relationships. Adding these descriptors to SMILES increases Pearson correlation from 0.639 to 0.674. Removing chemical regions yields 0.600 Pearson correlation, while random region inputs yield 0.604. Combining SMILES with image features achieves 0.793 accuracy and 0.674 Pearson correlation. None of these controls matches VLSR, which achieves 0.842 accuracy and 0.718 Pearson correlation.

Renderer robustness tests show VLSR obtains 0.788 accuracy and 0.632 Pearson correlation on PyMOL2D depictions, and 0.799 accuracy and 0.640 Pearson correlation under 3D-projected RDKit depictions, compared with 0.842 and 0.718 on canonical images, indicating the model can recover useful local structure under substantial depiction shifts.

Zero-shot evaluation on protein-conditioned FEP-derived ligand comparisons combines 317 ligand-pair comparisons from Schrödinger JACS and Merck subsets with 467 comparisons from additional Schrödinger pairwise datasets, covering 28 protein targets and 784 ligand comparisons. No FEP labels, affinity values, or protein contexts are used during training, and no fine-tuning or calibration is performed on this evaluation set. The model receives a ligand pair with the corresponding target protein sequence but no binding pockets, poses, docking structures, or 3D complexes. Both ligands in every retained pair have maximum ECFP4 Tanimoto similarity below 0.7 with training molecules. VLSR achieves 0.691 accuracy, outperforming the strongest SMILES-based baseline, Qwen3.5-4B + MSR-SFT (0.628), by 6.3 percentage points. In a tool-assisted setting, DeepSeek-V4-Flash with MSR improves the Instruct setting from 0.540 to 0.589 but does not consistently improve Think-style prompting. A HIF-2α case study shows VLSR correctly predicts the approximately 20-fold potency loss for the lig 163→lig 165 OH→NH2 edit, with the edited group near His293 and a structural water in the PT2385-bound structure (PDB 5TBM).

Ablation studies show removing learned regions causes a clear drop, replacing them with random regions degrades performance further, removing the workspace reduces classification and regression performance, and removing edit alignment hurts particularly when corresponding structures appear at different image positions. The ablations support the intended order: locate the evidence, establish correspondence when needed, and reason about its property effect.

Limitations acknowledged in the paper include: VLSR uses RDKit-derived region boxes and labels as training targets, so it does not eliminate chemical supervision and the annotation inventory may bias learned regions; CPT contains no property-reasoning supervision; completely excluding molecular overlap with general chemical pretraining corpora is impractical; 2D depictions cannot recover conformational energetics or protein–ligand contacts; region supervision is defined by rectangular boxes, so overlapping motifs and diffuse electronic effects may not correspond to a single target; and matched IoU should be read as spatial agreement rather than causal faithfulness. The FEP-derived results demonstrate directional transfer rather than full recovery of physical effects.

The paper concludes that VLSR establishes a localize-then-reason framework for molecular visual reasoning, learning chemically relevant regions before integrating their property effects within a compact latent workspace. Across scaffold-controlled evaluations, this design improves both classification and regression performance, while ablations show region localization, edit alignment, relative spatial relationships, and latent reasoning make complementary contributions. By avoiding lengthy textual rationales and decoding only the final answer, VLSR achieves 9.6× higher inference throughput than the corresponding image-input reasoning baseline. Its robustness across molecular depiction styles and transfer to protein-conditioned comparisons without additional training indicate the learned regional representations capture reusable chemical evidence.

Improvements for AI systems

Improvements to AI systems:

  1. Add a localize-then-reason stage to multimodal LLMs for scientific image tasks. Instead of letting the model attend freely over the entire image, insert a trainable region-query module that first identifies chemically (or domain-relevant) subregions via cross-attention, then feeds those region tokens into the reasoning core. This improves accuracy (0.767→0.842) and F1 (0.679→0.786) on molecular property tasks and can generalize to other structured images (e.g., anatomical scans, circuit diagrams, geological maps).

  2. Replace textual chain-of-thought with a compact latent reasoning workspace. Use a small set of latent tokens (e.g., ⟨EDIT⟩, ⟨PROP⟩, ⟨EFFECT⟩) that are updated over multiple transformer layers and decoded only at the final answer. This yields a 9.6× throughput increase (39.84 vs. 4.13 samples/s) while maintaining or improving accuracy, making it suitable for real-time or high-volume inference in drug discovery, materials screening, and automated lab analysis.

  3. Implement bidirectional soft alignment for paired-image comparisons. Instead of patch-wise subtraction or index-based matching, compute learned differences, products, and pooled interactions between region tokens from two images. This preserves additions/removals without atom maps and improves relational reasoning when corresponding structures appear at different positions—useful for before/after images, medical scans, or reaction monitoring.

  4. Use auxiliary localization supervision during training only, not at inference. Train with box/label targets (e.g., functional groups, rings) via Hungarian assignment and auxiliary losses (attention cross-entropy, soft-IoU, diversity, entropy), then discard these annotations at inference. This forces the model to discover relevant regions autonomously, improving robustness to depiction shifts (e.g., 3D-projected or PyMOL2D renderings) and enabling zero-shot transfer to new tasks without retraining.

  5. Apply reinforcement learning with answer-level rewards (GRPO) to refine reasoning without intermediate rationales. Reward correctness, regression closeness, sign consistency, and output cleanliness—no need for attention maps or latent state access. This improves directional accuracy (e.g., predicting potency loss from OH→NH2 edits) and transfers to protein-conditioned ligand comparisons (0.691 vs. 0.628 baseline) without fine-tuning.

  6. Design scaffold-aware data splits for training and evaluation. Use Bemis–Murcko scaffold components with ECFP4 similarity thresholds to prevent train/test overlap. This yields more reliable generalization estimates and reduces overfitting to molecular families, applicable to any domain with compositional structures (e.g., proteins, polymers, crystal structures).

What the improved AI system can do:

  • Accurately predict molecular properties (e.g., solubility, toxicity, binding affinity) directly from 2D structure images, with higher accuracy and F1 than symbolic baselines, while running 10× faster than text-reasoning LLMs.

  • Compare two molecular images and infer the effect of structural edits (e.g., potency change) without needing atom-mapped SMILES or 3D conformers, even when the edited region appears at different positions.

  • Generalize to unseen depiction styles (e.g., different renderers, 3D projections) and to new protein targets without additional training, enabling zero-shot screening of ligand pairs against protein sequences.

  • Operate in high-throughput settings (e.g., virtual screening of millions of compounds) due to fast inference and no textual rationale generation.

  • Provide interpretable localization (via attention maps) that highlights chemically relevant regions (functional groups, rings, heteroatoms) during training, aiding model debugging and scientific insight, while keeping inference annotation-free.

  • Handle continuous regression tasks (e.g., pIC50, ΔΔG) with improved Pearson correlation (0.718 vs. 0.674 for best baseline), useful for quantitative structure-activity relationship modeling.

Sources

Related papers