Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation".
Jane: The paper was written by Jiatong Li, Junxian Li, Weida Wang, Yunqing Liu, Changmeng Zheng et al. from Hong Kong Polytechnic University and Shanghai Jiao Tong University and Shanghai AI Lab and National University of Singapore.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Introduction and Core Idea: Tom: We’ve just heard about the paper, “Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation,” and it's clear that the researchers identified a massive gap in how we test AI for chemistry.
Jane: It turns out most current benchmarks only allow for one correct answer, like if you ask for a specific molecule, there’s usually only one defined target.
Lu: But this paper shows that in the real world of drug discovery, there is almost never just one perfect chemical solution that works.
Meng: That’s the core problem; we're currently testing models based on rote memorization rather than their ability to actually solve a problem creatively.
Tom: And Jane, you mentioned this "one-to-many" concept—how does that translate into practical use for finding new medicines?
Jane: It means that when a chemist has a target property, like higher solubility, there are often multiple different molecular shapes they can take to achieve it.
Lu: This approach recognizes the inherent flexibility in chemistry, which is something prior models completely missed because their training data was too rigid.
Meng: I see this as a huge win for usability; we aren't forcing users to narrow down their requirements into one single answer anymore.
Lalam: This allows AI to operate within the complexity of nature itself, serving the full breadth of human inquiry instead of just trying to fit into a pre-defined box.
Tom: So, by acknowledging that this is a dynamic relationship, we set the stage for testing something truly novel with all those ideas.
Methodology and Structure: Jane: To move beyond the single-answer limitation, “Speak-to-Structure” proposes three specific tasks to test AI capabilities: MolEdit, MolOpt, and MolCustom.
Tom: They’ve mapped these tasks directly onto real stages of drug discovery, which is incredibly smart because it makes the evaluation relevant.
Lu: For example, the design of "MolEdit" subtasks really tests if an AI understands chemical grammar and can make a localized change without breaking the whole structure.
Meng: When I look at "MolOpt," it feels like a much tougher test because the the model has to satisfy both making a specific structural change and achieving an improvement in a chemical property simultaneously.
Tom: That’s right, and then we have "MolCustom" which is pure creativity—designing brand new molecules from constraints like atom count or functional groups.
Jane: All three are designed to push the AI beyond simple pattern recognition into genuine structural reasoning and design capability.
Lu: The authors are suggesting that if an AI can pass these, it has mastered the fundamentals of chemical design itself, not just repeating examples.
Meng: This provides us with much clearer data on which models truly possess this deep understanding versus those who aren't capable of it.
Lalam: This framework allows us to define a new standard for intelligent assistance in science, ensuring that AI can handle the nuanced demands of complex discovery processes.
Tom: These structured tasks give us a clear path toward understanding how the AI actually works, which is a great lead-in to discussing how they built this data.
Data Construction and Evaluation: Tom: The researchers identified that human labeling is incredibly expensive for building large datasets, so they tackled that bottleneck head-on by using automated chemical toolkits.
Jane: They bypassed the need for manual annotation by programmatically generating millions of instruction-molecule pairs from databases like PubChem.
Lu: The idea of generating such a vast number of instruction sets shows an incredible level of systematic thinking about data generation.
Meng: This approach addresses the "data hunger" problem, meaning we no longer have to rely on slow, expensive human annotation for future training runs.
Tom: And Jane, when the AI generates a molecule using these tools, how do we know if it’s actually good? They put a lot of effort into automated evaluation.
Jane: They use metrics like Tanimoto Similarity to check if the molecule is a *rational* edit, which is much better than just checking for an exact match on its own.
Lu: That distinction between being 'correct' and being 'rationally related' is key, especially in MolOpt where you need that structural connection.
Meng: It’s about making sure the generated output isn't just random noise; it has to be chemically valid and relevant to the input, which is a huge operational improvement.
Lalam: This rigorous evaluation framework ensures that our tools are not just hitting targets, but are actually achieving meaningful progress in human endeavors.
Tom: These methods of building data and scoring results provide a solid foundation for understanding the next steps in AI training.
Conclusion and Future Outlook: Jane: We've seen how “Speak-to-Structure” moves beyond simple pattern matching, setting a new standard for evaluating LLMs in molecular design.
Tom: The findings clearly demonstrate that if we want truly capable AI, we need massive, high-quality instruction tuning on tasks that actually move past memorization.
Lu: I'm just thrilled about the potential of scaling this work; the researchers are only scratching the surface of what these models could achieve.
Meng: I think the practical impact is enormous; as a tool, it will help us accelerate drug discovery by allowing complex chemical tasks to be handled more reliably.
Lalam: We are seeing how AI can transition from merely reflecting existing data to actively synthesizing new capabilities for humanity.
Tom: Before we wrap up, let's get our final thoughts from the whole team on this groundbreaking research.
Lu: I'm just thrilled about the potential of scaling this work; we're only scratching the surface of what this could achieve.
Meng: The engineering path forward looks much clearer when knowing which models are capable and how they need to be trained.
Lalam: I feel that this allows us to see AI as a true partner in science, not just a passive data processor.
Tom: Thanks for joining us today! We're looking forward to the next paper on arXiv, but it was a fascinating discussion of "Speak-to-Structure" with all of you.
Jane: It truly is the kind of work that shows what AI can be, moving beyond pattern matching to genuine chemical reasoning and discovery.
Jiatong Li, Junxian Li, Weida Wang, Yunqing Liu, Changmeng Zheng, Yatao Bian, Dongzhan Zhou, Xiao-Yong Wei, Qing Li
Hong Kong Polytechnic University · Shanghai Jiao Tong University · Shanghai AI Lab · National University of Singapore
cs.CL
Submitted: 2026-05-22
Updated: 2026-08-25
Code: https://github.com/phenixace/S2TOMG-Bench
Importance score: 88/100
The gist: This paper introduces Speak-to-Structure (S2-Bench), the first benchmark designed to evaluate Large Language Models (LLMs) in "open-domain natural language-driven molecule generation." It addresses a
Key concepts
- One-to-many concept
- In chemistry and drug discovery, a target property (like higher solubility) can often be achieved by multiple different molecular shapes. This concept recognizes the inherent flexibility in chemistry, moving beyond the assumption that only one perfect chemical solution exists.
- Speak-to-Structure tasks
- To test AI's capabilities beyond simple memorization, the paper proposes three specific tasks: MolEdit (localized structural change), MolOpt (structural change plus property improvement), and MolCustom (designing new molecules from constraints).
- Tanimoto Similarity
- This is an automated evaluation metric used to check if a generated molecule is a *rational* edit. It is crucial because it determines if the output is chemically valid and structurally related to the input, rather than just checking for an exact match.
Terminology
Summary
This paper introduces Speak-to-Structure (S2-Bench), the first benchmark designed to evaluate Large Language Models (LLMs) in open-domain natural language-driven molecule generation.
It addresses a critical gap in existing research where molecule-text alignment datasets are predominantly built on one-to-one mappings,
which measure a model's ability to retrieve a single predefined answer rather than its creative potential to generate diverse, yet equally valid, molecular candidates.
This matters because real-world molecule discovery involves one-to-many relationships
where multiple distinct structures can satisfy the same desired properties.
The S2-Bench Framework
The benchmark is structured around three primary tasks that mirror critical phases of drug and materials science:
-
Molecule Editing (MolEdit): Tests the ability to perform
precise, localized structural modifications while preserving main structures,
analogous to lead optimization. -
Molecule Optimization (MolOpt): Evaluates the ability to refine a lead compound under
specified property constraints,
such as increasing solubility or reducing toxicity. -
Customized Molecule Generation (MolCustom): Challenges models to synthesize a novel molecule from scratch based on
quantitative and qualitative constraints,
such as a specific number of atoms, bonds, or functional groups.
To ensure objective assessment, the authors developed an automated evaluation system using cheminformatics toolkits to verify validity, check constraint satisfaction, and quantify quality through task-specific metrics like Tanimoto Similarity and novelty.
OpenMolIns Instruction Tuning
To facilitate model training, the researchers introduced OpenMolIns, a large-scale instruction tuning dataset
comprising up to 1.2 million instruction-molecule pairs programmatically constructed from the PubChem database. Unlike previous datasets that rely on scarce and costly human annotations,
OpenMolIns uses automated chemical toolkits to create diverse, one-to-many training examples. The dataset is provided in five distinct scales—light, small, medium, large, and xlarge—allowing researchers to systematically analyze the data scaling law
for molecular generation. This approach ensures the model learns genuine chemical reasoning
rather than simply memorizing specific input-output pairs.
Key Research Findings
Through an extensive evaluation of 31 LLMs, the study reveals several critical insights:
** Current LLMs lack the structural understanding necessary for precise molecule generation,
particularly in the MolCustom task, where they struggle to satisfy fine-grained structural constraints.
**
** Instruction tuning is indispensable for guiding and optimizing
capabilities; for instance, Llama3.1-8B fine-tuned on the xlarge OpenMolIns dataset surpassed powerful proprietary models like GPT-4o and Claude-3.5. **
** Existing one-to-one mapping datasets cause LLMs to rely on pattern recognition and recall
rather than true comprehension, as evidenced by models fine-tuned on ChEBI-20 performing worse on S2-Bench than general LLMs. **
** The data scaling law is task-dependent; while complex de novo synthesis (MolCustom) benefits immensely from larger datasets, simpler tasks like MolEdit may be constrained more by model capacity than by dataset size.
**
Conclusion and Impact
The paper concludes that by shifting the focus from simple pattern recall to realistic molecular design,
S2-Bench provides a more accurate measure of an LLM's potential in molecule discovery. The work aims to pave the way for more capable LLMs in natural language-driven molecule discovery,
ultimately assisting chemists in streamlining the discovery of new pharmaceuticals and materials.
Improvements for AI systems
To improve AI systems based on the findings and methodologies in this paper, I would implement the following specific architectural and training improvements:
-
Implement a
One-to-Many
Instruction Tuning Paradigm using Programmatic Data Generation. -
Integrate a Multi-Objective Weighted Success Rate (WSR) Reward Function for Reinforcement Learning from Human Feedback (RLHF).
-
Develop Discrete Constraint-Aware Reasoning Modules for Molecular Synthesis.
Sources
- Microsoft COCO Captions: Data Collection and Evaluation Server
- NovoMolGen: Rethinking Molecular Language Model Pretraining
- The Llama 3 Herd of Models
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Mistral 7B
- MolViBench: Evaluating LLMs on Molecular Vibe Coding
- Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
- A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language
- Galactica: A Large Language Model for Science
- Gemma 3 Technical Report
- Chem-R: Learning to Reason as a Chemist
- Qwen2 Technical Report
- Yi: Open Foundation Models by 01.AI
- ChemLLM: A Chemical Large Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering