Automating MD simulations for Proteins using Large language Models: NAMD-Agent
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Automating MD simulations for Proteins using Large language Models".
Tom: Molecular dynamics simulations are an essential tool for understanding protein structure, dynamics, and function at the atomic level;
Jane: First, who's behind it and why it matters.
Paper summary: Jane: We’ve covered how the paper introduces the NAMD-Agent pipeline, which uses an LLM to automate the creation of simulation input files and runs production simulations on NAMD3. We also discussed their findings regarding accuracy and stability across seven setups.
Lu: I think what stands out from "Automating MD simulations for Proteins using Large language models: NAMD-Agent" is the agentic framework they employed, which allows the AI to reason through the multi-step process of structural refinement and system assembly before execution.
Meng: From my perspective as an engineer, the most important thing is that they’ve managed to automate interactions with CHARMM-GUI via web automation and Selenium; that’s a concrete technical achievement for handling those specific software interfaces reliably.
Lalam: I see the most important vision here being how this advancement can improve our culture by making complex molecular modeling workflows significantly more accessible, moving the focus from tedious setup to meaningful scientific investigation.
Tom: So, in simple terms, this paper shows that we can use large language models to create a system that handles the entire preparation and execution phase of molecular dynamics simulations based on a simple text request. It’s about taking the heavy lifting out of getting simulation-ready files ready.
Jane: And its implication is that researchers with less direct experience with complex software can start running more detailed, high-quality MD simulations much faster than before <ref:2507.07887#pg1>.
Lu: The impact on the field is that it opens up a new avenue for exploring system dynamics that were previously inaccessible due to the time constraints of manual setup procedures.
Meng: Practically, it means quicker iterations in research cycles because generating those input decks takes minutes instead of hours, which directly impacts our ability to test hypotheses more frequently.
Lalam: The cultural shift is toward a future where the barrier to entry for running sophisticated computational simulations drops substantially, allowing more people to contribute meaningful work in this area.
Tom: So that's the high-level summary of "Automating MD simulations for Proteins using Large language models: NAMD-Agent." It’s a system that automates the preparation and execution stages using an LLM agent.
Conclusion: Tom: So, we've been talking about how this NAMD-Agent pipeline uses an AI to handle everything from setting up the initial protein structure to running the actual molecular dynamics simulations on NAMD3 without needing a ton of manual input.
Jane: Exactly, Tom; it’s really about using language to turn a simple idea into a complete simulation workflow, which simplifies things immensely for those trying to get complex models running.
Lu: I think what's fascinating is how they structured the agentic framework; it doesn't just guess the next step, it actually reasons through structural fixes like pre-processing PDB files and then uses retrieval augmented generation to pull in specific simulation code templates.
Meng: From my side, I’m focused on the automation aspect; they managed to get CHARMM-GUI and NAMD interacting through that browser automation layer using Selenium, which shows a really practical approach to interfacing with existing software.
Lalam: And from where I sit, the most profound vision here is how this moves us toward a future where the barrier to entry for running detailed molecular simulations drops substantially, allowing more people to contribute meaningful work in this area.
Tom: It really is about taking that heavy lifting out of getting simulation-ready files ready so researchers can actually focus on asking better scientific questions instead of wrestling with setup scripts.
Jane: That makes total sense; it’s about making the process less tedious and more focused on the actual science being done with the protein data.
Lu: The real power is in that code-aware RAG mechanism they used, which grounds the AI's generated workflows in tested examples instead of just pulling random text from a massive general knowledge base.
Meng: It's solid, but I wonder how robust this setup is when you push it toward much larger systems or different simulation engines; we have to check if those web interface dependencies will hold up under real stress.
Lalam: That’s a crucial point Meng brings up; ensuring reproducibility across different hardware and software platforms is the next big challenge for making this tool truly universal.
Tom: So, looking at the title, "Automating MD simulations for Proteins using Large language models: NAMD-Agent," it really captures that essence of using an AI agent to manage the entire MD pipeline.
Jane: And the authors have clearly done a good job demonstrating that this isn't just theoretical; they show us concrete results with successful runs on various protein and membrane systems.
Lu: Indeed, they achieved a seventy-one point four percent overall accuracy across seven different simulation setups, which gives us some solid empirical data to look at when we consider the practical application of such tools.
Meng: Those success rates are encouraging, but we still need to see how consistently that performance holds up when the input protein structures become significantly more complex or when we move beyond NAMD.
Lalam: This work points toward a future where complex computational modeling becomes accessible not just to specialists, but to a much broader community of researchers exploring biological systems at an atomic level.
Tom: It’s exciting stuff; this could seriously speed up the pace of discovery in structural biology if we can deploy this technology widely.
Carnegie Mellon University
cs.CL, cs.CE, q-bio.BM
Submitted: 2025-07-10
Updated: 2026-10-02
Comments: 25 pages
Code: https://github.com/SeleniumHQ/selenium
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 73/100
The gist: Molecular dynamics simulations are an essential tool for understanding protein structure, dynamics, and function at the atomic level; this work introduces an automated pipeline that leverages Large
Key concepts
- Agentic Framework
- This is an AI structure where the LLM acts as a central agent that can reason and take actions. It uses a 'think-act-observe' loop to solve complex problems by breaking them down into steps, allowing it to adapt its strategy when encountering errors or needing more information.
- Retrieval-Augmented Generation (RAG)
- RAG is a technique where the LLM doesn't just rely on its training data. Instead, it searches a specific collection of pre-written code and scripts related to MD simulations to find relevant instructions. This grounds the LLM's output in real, tested examples, significantly reducing errors or 'hallucinations'.
- CHARMM-GUI Automation
- CHARMM-GUI is a tool used to build complete molecular systems (like proteins and membranes) for MD simulations. The pipeline automates this process by using the LLM agent to interact with CHARMM-GUI's modules, effectively building the entire simulation setup from a simple text description.
Terminology
Summary
Molecular dynamics simulations are an essential tool for understanding protein structure, dynamics, and function at the atomic level; this work introduces an automated pipeline that leverages Large Language Models (LLMs) to streamline the generation of Molecular Dynamics (MD) input files by automating interactions with CHARMM-GUI and NAMD.
The gist
This work introduces an end-to-end, retrieval-augmented, large language model (LLM) pipeline that turns a short natural-language prompt into fully prepared NAMD input decks, launches production simulations, and returns standard structural analyses, all with negligible human intervention.
How it works: The Agentic Framework
The system operates using an agentic AI framework where the LLM backbone is Gemini-2.0-Flash, integrated via LlamaIndex to enable code generation and execution capabilities. This approach utilizes ReAct (Reasoning and Acting) agents, which interleave explicit reasoning steps with tool-using actions to solve complex tasks efficiently by thinking, acting, observing outcomes, and adapting its strategy accordingly. The workflow begins with user input—initial queries, system specifications (such as protein structure and membrane composition), and simulation parameters—which Gemini interprets to ensure clarity, completeness, and consistency before proceeding.
How it works: Automated Workflow Stages
The operational workflow is structured into several distinct stages managed by the LLM supervision. The agent first utilizes PDBFixer to preprocess downloaded protein structures by resolving common issues such as missing atoms, incomplete residues, and nonstandard residues,
ensuring compatibility. Next, the agent employs CHARMM-GUI in its Solution Builder and Membrane Builder modules
to automate system construction. This automation is accomplished using the Gemini model’s agentic capabilities with reference to the Auto CGUI github repository as a structured codebase. Following system assembly, files generated by CHARMM-GUI are organized systematically into structured directories for reproducibility and scalability.
How it works: RAG for Code-Aware Simulation Pipelines
A critical component of the pipeline is Retrieval-Augmented Generation (RAG), which is adapted to MD simulations by restricting the retrieval corpus to a collection of specific automation scripts in a code repository. This code-aware RAG mechanism retrieves function definitions, API usage patterns, shell scripts, and simulation templates.
This allows the Gemini agent to compose, modify, and execute simulation workflows based on natural language prompts
while ensuring that the generated code adheres to syntactic and semantic conventions consistent with prior workflows. This retrieval significantly reduces hallucinations by grounding the output in real, tested examples.
How it works: Toolset for Simulation Execution and Analysis
NAMD-Agent employs a comprehensive set of tools across three main areas:
-
Preprocessing Tools: Utilizing PDBFixer to address structural issues.
-
Simulation Setup Tools: Automating interactions with CHARMMGUI via a Selenium-driven browser automation layer to handle
form submissions, handling downloads, and selecting parameters.
-
Execution and Analysis Tools: Production simulations are run using NAMD3, while post-processing involves domain-specific Python scripts drawing from MDTraj and OpenMM packages for analyses such as RMSD, RMSF, SASA, Radius of Gyration (Rg), and hydrogen-bond profiling.
Results Summary
The framework was evaluated across seven distinct molecular dynamics simulation setups. Five simulations were successfully completed, yielding an overall accuracy of 71.4% for the automated pipeline.
Successful runs included two solution systems (1UBQ and 1L2Y) and three membrane systems (1AFO in two configurations, and 1CRN). The agent generated valid topologies, coordinates, and parameter files in minutes
and produced stable trajectories whose analyses—RMSD traces plateaued after initial fluctuations, suggesting stable system configurations
—matched literature expectations. Failures were identified in edge cases such as membrane packing constraints or structural instability.
Limitations and Future Work
Limitations include dependency on web interfaces (CHARMM-GUI HTML elements), engine specificity (currently supporting only NAMD/CHARMM-GUI), potential LLM hallucination, and scalability challenges for large systems. Future work aims to enhance robustness by expanding multi-engine generalization to GROMACS, AMBER, and OpenMM; incorporating adaptive simulation steering via reinforcement learning; ensuring reproducibility through end-to-end provenance tracking (e.g., CWL or Nextflow); and improving reliability by integrating multiple specialized language models under a supervisory framework for error checking.
Acknowledgments
The code is available at the following link: https://github.com/BaratiLab/NAMD AGENT. All simulations were performed using a GeForce GTX 1080 Ti GPU with 11GBs of memory. The paper cites numerous foundational works in LLMs, MD automation tools (like Chaperong, easyAmber), and specific MD analysis metrics (RMSD, Rg).
Improvements for AI systems
Here are specific improvements for the NAMD-Agent system, focusing on enhancing its reliability, generality, and scientific rigor:
-
Enhance Error Handling through Contextual Validation (Mitigating Hallucinations)
-
Implement Dynamic Parameter Steering via Bayesian Optimization (Adaptive Simulation)
-
Achieve Multi-Engine Generality via Modular API Integration (Versatility)
-
Integrate a Hierarchical LLM Verification Layer (Robustness and Trust)
-
A system that can autonomously generate, execute, and analyze complex molecular dynamics simulations across multiple simulation engines (NAMD, GROMACS, AMBER) by dynamically selecting the most appropriate configuration scripts based on the input protein/membrane type. It would handle
edge cases
(like membrane size constraints or force field incompatibilities) by querying a verified database of known system limitations before execution. -
An agent capable of receiving a high-level scientific goal (e.g.,
Simulate protein X in a solvated POPC bilayer for 50 ns and determine the RMSD convergence
) and iteratively refining simulation parameters (temperature, pressure, ion concentration) in real-time using reinforcement learning or Bayesian optimization based on initial trajectory data. This allows the AI to actively steer the simulation toward desired conformational states or rare events rather than simply following a static protocol. -
A framework that can generate complete, executable input decks for GROMACS and AMBER, not just NAMD/CHARMM-GUI configurations, by retrieving and adapting engine-specific syntax from a structured code repository (CodeRAG). This allows the system to support a broader range of established MD software platforms without requiring bespoke training for each engine.
-
A supervisory architecture where multiple specialized LLMs (e.g., one focused on force field knowledge, another on syntax generation, and a third acting as a final verifier) vote on the generated configuration before execution. This hierarchical verification layer drastically reduces the risk of hallucinated parameters (like unsupported keywords or incorrect topology formats) by requiring consensus among domain experts, ensuring higher fidelity in the final simulation inputs.
Abstract
Molecular dynamics (MD) simulations are essential for understanding protein structure, dynamics, and function, but preparing, running, and analyzing simulations remains time-consuming and error-prone. We present an automated pipeline that combines large language model (LLM) agents with Python scripting and HTMD MCP tools to generate simulation-ready inputs for NAMD3/CHARMM, execute simulations, analyze outputs, and recover from build or runtime failures. The framework was evaluated across five biomolecular system classes: a protein-DNA complex (p53 DNA-binding domain bound to its response element), a protein-membrane system (M2 muscarinic receptor with iperoxo in a POPC/cholesterol bilayer), a protein-ligand series (five congeneric TYK2 inhibitors), a protein-water reference (ubiquitin), and a protein-protein complex (barnase-barstar). For the protein-DNA system, the automated workflow reproduced key metrics from an independent published benchmark. To distinguish framework performance from model-specific behavior, we repeated the complete protein-DNA study using three LLM orchestrators: Claude Opus-4.8, GPT-5.6 Sol, and Nemotron 3 Ultra. All three completed the workflow, but they differed in benchmark-ranking fidelity and by up to two orders of magnitude in token consumption and cost. Across all system classes, the agent recovered experimentally and computationally established behavior. Additional post-processing software was used to refine simulation outputs, enabling a complete and largely hands-free workflow. This approach reduces setup effort, limits manual errors, supports parallel handling of diverse biomolecular systems, and provides a robust, adaptable foundation for LLM-driven automation in computational structural biology.
Sources
- Agent AI: Surveying the Horizons of Multimodal Interaction
- LLM-3D Print: Large Language Models To Monitor and Control 3D Printing
- Adsorb-Agent: Autonomous Identification of Stable Adsorption Configurations via Large Language Model Agent
- LLM-Drone: Aerial Additive Manufacturing with Drones Planned Using Large Language Models
- LLM-guided Chemical Process Optimization with a Multi-Agent Approach
- MDCrow: Automating Molecular Dynamics Workflows with Large Language Models
- Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions
- ReAct: Synergizing Reasoning and Acting in Language Models
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Augmented Language Models: a Survey
- NANOGPT: A Query-Driven Large Language Model Retrieval-Augmented Generation System for Nanotechnology Research
- Code Llama: Open Foundation Models for Code
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering