Unlocking Multimodal Protein Language Models at Inference Time
cs.CE, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: Accepted to EMNLP 2026 Main Conference
Code: https://github.com/EchoChou990919/mplm_inference
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies.
Terminology
Abstract
Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies. Yet prior work has focused more on model training than on how inference-time strategies behave. In this paper, we establish a three-stage investigation framework to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks. We evaluate vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search on multimodal pLMs, corresponding to controls over sampling distributions, per-step logits, and parallel trajectories. Throughout the complementary advancements centered on exploration-exploitation trade-off, we (1) reveal the suboptimality of default inference protocols and identify task-oriented sampling preferences; (2) observe substantial quantitative gains across tasks, consistently boosting the upper bound performance of multimodal pLMs without updating model parameters; (3) derive conclusions about base models that differ from prior consensus.
Sources
- RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance
- Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute
- Computational Protein Science in the Era of Large Language Models (LLMs)
- Classifier-Free Diffusion Guidance
- ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search
- Geometric Flow Matching for Molecular Conformation Generation via Manifold Decomposition
- Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
- Improved motif-scaffolding with SE(3) flow matching
- Hierarchical protein backbone generation with latent and structure diffusion
- Controlling Repetition in Protein Language Models
- MotifBench: A standardized protein design benchmark for motif-scaffolding problems
- HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens
Related papers
- Constrained Sensing and Reliable State Estimation with Shallow Recurrent Decoders on a TRIGA Mark II Reactor
- Evidence-Unit Fairness and the Limits of Query-Adaptive Sparse-Dense Fusion in Financial Document Retrieval
- Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad
- Lightweight Adaptation of EEG Foundation Models for Stroke Motor Imagery Decoding: Domain Shift and Subject-Level Robustness
- RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning
- Wildfire Suppression: Complexity, Models, and Instances