Speak to a Protein: An Interactive Multimodal Co-Scientist

arXiv:2510.17826 · q-bio.BM, cs.AI · Submitted 2025-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "Speak to a Protein: An Interactive Multimodal Co-Scientist".

Marcus: Building a working mental model of a protein typically requires weeks of reading, crossreferencing crystal and predicted structures, and inspecting ligand complexes, an effort that is slow, unevenly accessible,

Ines: First, who's behind it and why it matters.

Title and authors: Ines: So, we've been looking at this paper titled "Speak to a Protein: An Interactive Multimodal Co-Scientist," and it sounds like it’s tackling that huge time sink involved in understanding protein structures. It suggests moving away from weeks of manual reading and cross-referencing crystal structures to something much more immediate with an interactive dialogue.

Marcus: Right, I was thinking the authors are really aiming to reduce the friction for people who don't have deep structural biology backgrounds. They're proposing a system that can pull in literature, structures, and ligand data all at once and then ground the answers in a live three dee scene <ref:2510.17826#pg0,answers in a live 3D scene>. That sounds like it would dramatically speed up preliminary discovery phases.

Yuki: From a population genetics standpoint, I see this as a way to rapidly test hypotheses about protein function across different structural states or binding partners without needing massive, dedicated lab resources for every single test. It could accelerate our understanding of how these molecular interactions play out in biological systems over time.

Ines: Exactly, Yuki; the core idea is turning protein analysis into a dialogue rather than a solitary deep dive into static files. The paper shows how the AI agent orchestrates tools like literature searches and structure lookups to build this context dynamically for each question asked.

Marcus: And I’m interested in how that data retrieval works, because if it can pull in information from UniProt, PDB, and ChEMBL simultaneously, we’re talking about a lot of organized data ingestion. My concern is always whether that automated pipeline will handle the complexity of real-world biological variability without introducing significant batch effects into the results.

Yuki: That variability is where I think it gets interesting; if the system can correlate structural features with known functional changes derived from literature, it could provide a much richer context for evolutionary analysis than just looking at a single static structure.

Ines: The paper also highlights how this isn't just about retrieval; the AI is able to generate code when needed and explain its results through both text and graphics, which allows for real-time verification. That tight coupling between the language, the code execution in that sandbox, and the visualization is a big part of what they’re proposing here.

Marcus: I see that capability translating directly into reproducibility; having the Python code shown in real time makes it much easier to debug why a particular analysis yielded a specific result compared to another. That transparency is something we need when dealing with complex genomic datasets.

Yuki: It gives us a way to visualize the molecular logic, which bridges the gap between sequence information and actual functional outcomes in a more tangible way for anyone involved in the study.

Title and authors: Ines: Moving into their proposed improvements, they suggest enhancing this tight loop by making it even more direct, moving beyond simple retrieval to direct structural manipulation via natural language commands. This means if you want to measure something specific in the three dee scene, you don't need to write a script; you just ask the AI directly <ref:2510.17826#pg0>.

Marcus: That would simplify things immensely for users who aren't specialized coders; it lowers that barrier I mentioned earlier by abstracting away the need for deep Python knowledge. It sounds like they are trying to make the interface feel more intuitive for a broader group of scientists.

Yuki: From my perspective, this direct manipulation capability could allow us to quickly probe subtle conformational differences in proteins that might not be obvious from a standard static view, which is crucial when studying how function emerges from sequence variation.

Ines: And they also emphasize the need for automated, reproducible data pipeline generation. They want the AI to handle the whole sequence—from finding the relevant PDB entry to cleaning and normalizing assay data in ChEMBL into a usable format automatically.

Marcus: That's a huge step because manual normalization of heterogeneous biological datasets is often where errors creep in; having an AI do that systematically, based on defined rules, should improve the statistical rigor significantly.

Yuki: If the pipeline is robust enough to handle those complexities, it could mean we can focus less on data wrangling and more on interpreting the actual biological signals we are looking for in these large cohorts.

Ines: The paper also touches on context-aware deduplication and filtering for large datasets like those PDB structures. It suggests the AI should intelligently apply rules, such as keeping only entries with the most relevant metrics, based on what the user is asking.

Marcus: That addresses a real practical issue we face with massive structural databases; wading through thousands of potentially redundant entries manually is inefficient and prone to human error in selection.

Yuki: It connects back to how we analyze large genomic cohorts; if the AI can intelligently filter based on statistical relevance, it helps narrow down the search space for meaningful biological connections.

Ines: The paper also points toward automated sequence analysis directly on the structures, such as comparing binding pocket amino acid sequences and summarizing conservation patterns within a report. That moves beyond just looking at where a ligand binds to understanding what's structurally conserved there.

Marcus: Sequence conservation is vital for understanding functional constraints; if the AI can automate that comparison across many structures, it provides statistical evidence about which parts of the protein are most critical for its function.

Yuki: That kind of structural insight, when linked to evolutionary history, could help us map out how functional domains have been maintained or lost across different species.

Ines: Finally, they discuss generating publication-ready reports in formats like LaTeX or markdown that summarize all the gathered evidence for other researchers to interrogate. This turns the interactive session into a reusable knowledge base.

Title and authors: Marcus: I think that ability to synthesize all those disparate data sources—the literature, the statistics from ChEMBL, and the visual structural data—into a clean report is where we see real utility for accelerating research workflows.

Yuki: It’s about creating a shared language where complex findings from molecular biology and population genetics can be presented in a format that is accessible to both disciplines.

Ines: So, to wrap up on "Speak to a Protein: An Interactive Multimodal Co-Scientist," the paper shows an architecture that tightly couples language, code execution, and three dee visualization into one interactive loop for protein analysis <ref:2510.17826#pg0,Speak to a Protein: An Interactive Multimodal Co-Scientist>. It demonstrates how this approach can lower the barrier for complex data analysis by allowing users to ask questions and get immediate visual updates grounded in retrieved literature and structures.

Marcus: And they show concrete examples using proteins like the dopamine D3 receptor, where they managed to retrieve four hundred seventy-nine unique ligand/structure pairs and query ChEMBL for bioactivity data, generating an interactive table of inhibitors with SMILES strings. That shows the practical scale of what's possible with this architecture.

Yuki: The implication for us is that we can start asking more complex, integrative questions about molecular function much faster than we currently can through traditional methods alone. We gain speed in hypothesis generation by having a system that handles the heavy lifting of data synthesis and visualization.

Ines: That speed in generating hypotheses is what really matters for moving forward with our research programs; it allows us to test more ideas within a compressed timeline. Overall, the "Speak to a Protein" paper lays out a framework where the AI acts as an expert co-scientist, streamlining the entire process from initial question to evidence synthesis.

Marcus: It sounds like they've built something that helps manage the complexity of large structural datasets while maintaining enough transparency through code execution to ensure those results are sound. That level of control over the data pipeline is important when you’re dealing with population statistics derived from these structures.

Yuki: I think the real impact lies in democratizing access to high-level structural biology interpretation; if this becomes a standard way to explore protein-ligand interactions, it could open up new avenues for us to link molecular details directly to broader biological or evolutionary narratives.

Ines: It’s a powerful demonstration of how integrating these different modalities—literature, structure, and code—can create a much more efficient path toward understanding complex protein behavior than any single modality can offer on its own.

Marcus: We're ready to see how this kind of interactive tool can start streamlining our own data processing pipelines for cohort analysis next.

Yuki: And I look forward to seeing how this technology helps us connect the dots between molecular structure and the history of life in future studies.

The paper's summary: Ines: So, to recap, we're looking at how this new system essentially takes weeks of complex manual data gathering for protein structure analysis and condenses it into a single interactive session where you talk to an AI that pulls from literature, structures, and activity tables simultaneously.

Marcus: Exactly; it’s about compressing a huge workflow—from searching PDB files to filtering ChEMBL assay data—into something you can do in less than an hour, which is huge for managing big genomic cohorts.

Yuki: From a population genetics angle, it suggests we can test hypotheses about protein function and structure across many different biological samples much faster than before, giving us more context on how these molecular interactions evolve within species.

Ines: But what does this actually recover biologically? I mean, the paper shows the AI isn't just pulling text; it’s grounding its answers in a live three dee scene that highlights specific residues or binding pockets based on your natural language prompts.

Marcus: That’s where the engineering gets interesting; it means you don't need to learn complicated software interfaces to manipulate a structure, you just describe what you want—like asking the AI to measure a distance between two atoms—and it executes the underlying code in real time.

Yuki: That ability to visualize molecular logic directly from natural language gives us a tangible way to connect sequence variations we find in structures with observed functional differences in living organisms.

Ines: And the paper also points out that this system can generate reports, like LaTeX documents, summarizing all that gathered evidence, so it becomes a usable knowledge base for other researchers.

Marcus: That’s important because it addresses the reproducibility issue; if you can get a report generated from an AI session, you have a traceable record of how those specific structural and statistical data points were combined.

Yuki: It really shifts the focus from being stuck in the data collection phase to actually interpreting the biological story that those structures tell us about life history.

Ines: Looking ahead, they mention that future work involves giving this system tools to access proprietary datasets, which suggests it could move beyond public information and become a more comprehensive research partner.

Marcus: That would mean we could potentially apply this same framework to analyzing much larger internal company datasets or sensitive genomic archives without having to build entirely new data pipelines from scratch.

Yuki: If we can integrate this level of interactive analysis into our studies, it could dramatically speed up the pace at which we can map out the evolutionary pressures acting on key protein families.

Ines: So, the core takeaway is that this paper demonstrates a powerful way to couple language understanding with structural biology and code execution to make complex scientific inquiry immediate and visual.

Marcus: It shows that automating the tedious data wrangling step by step, while keeping a live three dee view active, really makes a massive difference in how quickly we can move from an idea to concrete evidence.

Yuki: The real implication is democratizing access to this kind of deep structural analysis, allowing researchers who might not have years of specialized training to make high-level structural comparisons and draw biological conclusions.

Ines: So, the next thing we'll be talking about is how this system handles the specific challenges of data deduplication when dealing with thousands of overlapping structures in a single database.

The paper's improvements: Ines: So, we're shifting gears now to how this paper suggests making the AI even better than what they have already built by focusing on several key improvements.

Marcus: Right, I was looking at their suggestions to refine that tight loop between language, code, and visualization. They are pushing for a more direct structural manipulation feature so users can just issue a command like "measure this distance" without writing any specialized Python script.

Yuki: That direct manipulation capability is key because it lowers the barrier for bench scientists; they don't need to be deep coders to probe subtle conformational differences in proteins when they are looking at how species have diverged over time.

Ines: I agree, and then there's the idea for automated, reproducible data pipeline generation. The authors want the AI to autonomously handle the whole process, from finding structures to cleaning up ChEMBL assay data into a standardized format without needing step-by-step prompting from us.

Marcus: That sounds like a huge win for my work on cohorts; automating that normalization process means we can focus less on tedious data wrangling and more on analyzing the actual statistical trends in the protein activity.

Yuki: If the AI can manage that complexity, it allows us to rapidly test hypotheses about how different structural features correlate with population-level functional changes across species.

Ines: They also emphasize context-aware deduplication and filtering for massive datasets, suggesting the AI should intelligently select only the most relevant entries based on what we’re asking of it at any given moment.

Marcus: That addresses a real practical issue with PDB structures; wading through thousands of potentially redundant entries manually is inefficient, so having the AI apply those complex filters based on user intent is much smarter for cohort analysis.

Yuki: It connects back to how we analyze large biological samples; if the AI can intelligently filter based on statistical relevance, it helps us narrow down the search space for meaningful evolutionary connections.

Ines: Plus, they propose automated sequence and pocket analysis directly on the visualized structures, meaning we could see conservation patterns in binding pockets immediately without doing manual sequence alignment.

Marcus: That structural insight is valuable; seeing those conservation maps automatically generated gives us concrete evidence about which parts of the protein are functionally constrained across different species.

Yuki: That kind of structural evidence, when linked to evolutionary history, really helps us map out how functional domains have been maintained or lost across different lineages.

Ines: And finally, they stress the need for automated synthesis into publication-ready reports in formats like LaTeX or markdown, turning the interactive session into a reusable knowledge base.

Marcus: That's where we see real utility; being able to export a structured report summarizing all that gathered evidence means we have a traceable record of how those specific structural and statistical data points were combined.

Yuki: It really shifts the focus from just gathering raw numbers to creating a shared language where complex findings from molecular biology and population genetics can be presented in an accessible format.

Ines: So, these improvements focus on making the AI not just a retriever, but an active analyst that proactively manages complexity and generates structured outputs for us.

Marcus: It seems like they're building a much more robust data pipeline that handles the inherent messiness of real biological data much more efficiently than current methods allow.

Yuki: I think this entire direction points toward a future where we can connect molecular structure directly to evolutionary theory with unprecedented speed and rigor.

Ines: We’ll be looking at the next segment to see how these suggested improvements translate into actual performance gains in terms of analysis speed.

Conclusion: Ines: So, to wrap up our discussion on "Speak to a Protein: An Interactive Multimodal Co-Scientist," we've covered how this paper creates an interactive loop that merges literature retrieval, structural biology visualization, and code execution into one dialogue for analyzing protein structures.

Marcus: It really shows how compressing the traditional weeks of manual data gathering into a single hour can dramatically accelerate the process for handling large cohorts and complex biological systems.

Yuki: From a population genetic perspective, this technology suggests we can test hypotheses about protein function across many different biological samples much faster than before, giving us more context on how these molecular interactions evolve within species.

Ines: The implication is that we’re moving toward a system where hypothesis generation isn't just possible but immediate and visually verifiable right in the moment.

Marcus: That speed in generating evidence means we can start testing more complex ideas much sooner, which is essential when managing the statistical noise inherent in large genomic datasets.

Yuki: It opens up new avenues for linking molecular details directly to broader biological or evolutionary narratives without getting bogged down by massive amounts of static data.

Ines: Overall, this paper demonstrates a powerful framework for acting as an AI co-scientist that streamlines the entire research workflow from initial question to evidence synthesis.

Marcus: It’s a solid demonstration of how coupling language with structural visualization and code execution can make complex structural analysis feel much more intuitive and manageable for a wider range of researchers.

Yuki: I think this is about democratizing access to deep structural biology interpretation, allowing people who might not have years of specialized training to draw high-level conclusions about protein behavior.

Ines: It’s a really exciting development because it fundamentally changes the pace at which we can explore new biological questions using these tools.

Marcus: And for us in genomics, seeing this kind of structured synthesis makes the idea of managing batch effects and cohort variability feel much more tractable through an automated pipeline.

Yuki: We're looking forward to seeing how this technology helps us connect the dots between molecular structure and the history of life in future studies.

Ines: That’s right, we'll be moving on now to discuss another exciting paper that tackles a different kind of data challenge in the field.

Carles Navarro, Mariona Torrens, Philipp Tholke, Stefan Doerr, Gianni De Fabritiis

Acellera Labs

q-bio.BM, cs.AI

Submitted: 2025-10-01

Updated: 2026-10-06

Code: https://github.com/jerryjliu/llama_index

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Building a working mental model of a protein typically requires weeks of reading, crossreferencing crystal and predicted structures, and inspecting ligand complexes, an effort that is slow, unevenly

Key concepts

Agentic Co-Scientist
This is the core AI orchestrator that plans a sequence of actions based on natural language requests. It uses Model Context Protocol (MCP) to coordinate various specialized tools, such as searching literature or querying protein databases, to achieve complex scientific goals.
Multimodal Tooling Layer
This layer allows the AI model to interact with different types of data simultaneously. It enables the system to retrieve text from literature, structured data from UniProt, and visual information from 3D structures during a single interaction.
Interactive Visualization and Code Execution
This feature lets users visualize protein data in real-time within a 3D viewer. Users can request actions like highlighting binding sites, which triggers Python code execution in the browser to produce immediate visual changes or perform custom measurements.
Model Context Protocol (MCP)
MCP is the framework that governs how the LLM agent communicates with all available tools. It allows the agent to intelligently decide which tool to use next—whether it's a literature search, a database query, or running code—to build a comprehensive answer.

Terminology

Summary

Building a working mental model of a protein typically requires weeks of reading, crossreferencing crystal and predicted structures, and inspecting ligand complexes, an effort that is slow, unevenly accessible, and often requires specialized computational skills.

The gist

Speak to a Protein introduces an interactive multimodal dialogue with an expert co-scientist that retrieves literature, structures, and ligand data; grounds answers in a live 3D scene; generates code when needed, and explains results in both text and graphics.

System Architecture: The Agentic Co-Scientist

The system is organized around a frontend for user interaction and visualization (chat panel with 3D viewer) and a back-end for language understanding, tool coordination, and data retrieval. At its core is an LLM Agent that orchestrates all available tools through Model Context Protocol (MCP). This agent plans a sequence of actions based on natural language requests, invoking domain-specific tools such as Literature Search, UniProt Search, ChEMBL Search, PDB Search (PDBe), MoleculeKit Search, and a Python sandbox. The system’s multimodal tooling layer is implemented as MCP servers that the model can invoke and compose during an interaction.

Data Retrieval and Grounding Modalities

The system utilizes several specialized tools to ground its responses across multiple modalities:

  1. Literature Search: This tool constructs a protein-specific corpus by searching PubMed Central, retrieving relevant articles in XML format, cleaning the text, and embedding it into a vector space for retrieval-augmented generation (RAG). The index is cached on disk for future queries.

  2. UniProt Search: This MCP server allows for both text search to resolve canonical entries from colloquial names and data lookup to retrieve detailed entry information, including sequence data, functional annotations, and cross-references to other databases like PDB and DrugBank.

  3. ChEMBL Search: This tool retrieves bioactivity and assay data from the ChEMBL database based on a target identifier and assay type. It normalizes heterogeneous datasets into a streamlined representation (CSV format) stored locally for downstream analysis using Python libraries like pandas.

  4. PDB Search: This tool provides structured access to the Protein Data Bank API, retrieving detailed entry information such as metadata, experimental method, resolution, and crucially, lists of co-crystallized small molecules (ligands).

Interactive Visualization and Code Execution

The core innovation lies in grounding responses across multiple synchronized modalities. When a user asks a question, the AI interacts with a live 3D structural viewer to highlight residues or annotate binding pockets. The front-end includes a Python sandbox powered by Pyodide, enabling the execution of Python code directly in the browser to perform custom analyses or visual manipulations. Users can request actions such as:

: highlight the ATP-binding site is translated into tool calls that execute Python code within the viewer, producing visual changes in real time. ˚

The visualization toolkit allows users to control representations by switching between styles (cartoon, ball-and-stick, surface) and coloring them according to properties like chain or residue type. Users can perform measurements of distances, angles, and dihedrals between atoms or residues.

Workflow Efficiency and Case Studies

Speak to a Protein reduces the time from question to evidence by tightly coupling language, code, and 3D structures. Case studies demonstrate this capability:

  1. For the dopamine D3 receptor (D3R), the system retrieved all PDB structures, identified relevant ligand/structure pairs (479 unique pairs), queried ChEMBL for bioactivity data (finding 132 annotated ligands), and generated an interactive table of inhibitors, including SMILES strings and ChEMBL IDs.

  2. For Cyclin-dependent kinase 2 (CDK2), the system systematically explored all available PDB structures, identified 479 unique ligand/structure pairs, filtered for relevant drug discovery compounds, and performed iterative cleaning to focus on IC50 values. It then loaded the top 20 most potent complexes into the viewer and performed structural alignment using robust structure-based methods.

  3. The system can generate publication-ready reports in LaTeX or markdown format summarizing all gathered information, creating a knowledge base that other users can interrogate.

Conclusion and Limitations

The study demonstrates how Speak to a Protein compresses traditional workflows involving several hours of manual data gathering into an interactive session taking less than an hour, enabling fluid hypothesis generation. Current limitations include reliance on public information (literature, structure-activity relationships) and potential performance issues when rendering a large number of complex structures simultaneously. Future work plans involve extending the system with tools to access internal or proprietary datasets and improving robustness for handling large data payloads.

Speak to a Protein is freely accessible at https://open.playmolecule.org.

References

(The paper cites numerous references, including works by Achiam et al., Boiko et al., Bran et al., Burley et al., Chien et al.

Improvements for AI systems

Here are specific improvements for AI systems based on the Speak to a Protein research, focusing on enhancing their capability as an agentic co-scientist in structural biology:


) 1. Unified Multimodal Grounding (Language-Code-Structure Loop):

The system should move beyond simple retrieval and integrate language reasoning, code generation, and 3D visualization into a single tightly coupled loop.

  1. Direct Structural Manipulation via Natural Language: The AI can translate high-level natural language commands directly into executable Python code for the molecular viewer (e.g., highlight the ATP-binding site or measure the distance between residue X and ligand Y). This eliminates the friction of needing to learn specialized software interfaces.

  2. Automated, Reproducible Data Pipeline Generation: The AI should be able to autonomously generate entire analytical workflows—from PDB retrieval and ligand extraction to data normalization (e.g., converting heterogeneous ChEMBL assay data into a standardized CSV)—without explicit user prompting for each step.

  3. Context-Aware Data Deduplication and Filtering: The system must intelligently manage large, redundant datasets (like the 460+ CDK2 structures). It should automatically apply complex filtering rules (e.g., keep only entries with the lowest IC50 value per PDB structure) based on user intent, rather than requiring sequential manual commands for cleaning.

  4. Automated Sequence/Pocket Analysis: The AI can perform sequence analysis directly on the visualized structures, such as automatically extracting and comparing binding pocket amino acid sequences (as demonstrated in the paper), generating pairwise alignments, and summarizing conservation patterns in a structured report.

  5. Cross-Modal Hypothesis Generation: The system should be able to correlate findings across modalities—for instance, linking a specific sequence variation found in the pocket (from sequence analysis) directly to an observed change in ligand potency (from SAR data) and visualizing the structural difference simultaneously.

  6. Publication-Ready Report Synthesis: The AI must transition from answering questions to generating deliverables. It should be able to synthesize all gathered evidence (literature, structural data, activity tables) into structured formats like LaTeX reports or executive summaries optimized for specific audiences (e.g., a report for a chemist vs. one for a molecular biologist).

  7. Proactive Tool Orchestration: Instead of waiting for the user to ask the next question, the AI should proactively suggest logical next steps based on its accumulated knowledge and data (e.g., Given these potent inhibitors, would you like me to run a docking simulation against this pocket or perform a sensitivity analysis on these residues?).

) Improved AI System Capabilities:

The improved system will function as a true, interactive, and autonomous AI Co-Scientist capable of:

  1. Reducing the time from question to evidence from hours/days to minutes by automating the entire data gathering, cleaning, and preliminary analysis pipeline.

  2. Lowering the barrier to entry for advanced structural analysis by allowing non-specialists (bench scientists) to conduct complex structure-activity relationship (SAR) studies using only natural language dialogue.

  3. Enabling rapid hypothesis generation by enabling real-time testing of ideas directly within a synchronized 3D visualization environment, allowing users to visually verify predictions immediately after the AI generates them.

  4. Serving as an automated research assistant capable of producing publication-quality reports and sequence analyses (like pocket conservation maps) with high fidelity, ensuring scientific rigor in every output.

Abstract

Building a working mental model of a protein typically requires weeks of reading, cross-referencing crystal and predicted structures, and inspecting ligand complexes, an effort that is slow, unevenly accessible, and often requires specialized computational skills. We introduce Speak to a Protein, a new capability that turns protein analysis into an interactive, multimodal dialogue with an expert co-scientist. The AI system retrieves and synthesizes relevant literature, structures, and ligand data; grounds answers in a live 3D scene; and can highlight, annotate, manipulate and see the visualization. It also generates and runs code when needed, explaining results in both text and graphics. We demonstrate these capabilities on relevant proteins, posing questions about binding pockets, conformational changes, or structure-activity relationships to test ideas in real time. Speak to a Protein reduces the time from question to evidence, lowers the barrier to advanced structural analysis, and enables hypothesis generation by tightly coupling language, code, and 3D structures. Speak to a Protein is freely accessible at https://open.playmolecule.org.

Sources

Related papers