Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

arXiv:2412.14642 · cs.CL · Submitted 2026-05-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation".

Jane: The paper was written by Jiatong Li, Junxian Li, Weida Wang, Yunqing Liu, Changmeng Zheng et al. from Hong Kong Polytechnic University and Shanghai Jiao Tong University and Shanghai AI Lab and National University of Singapore.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Introduction and Core Idea: Tom: We’ve just heard about the paper, “Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation,” and it's clear that the researchers identified a massive gap in how we test AI for chemistry.

Jane: It turns out most current benchmarks only allow for one correct answer, like if you ask for a specific molecule, there’s usually only one defined target.

Lu: But this paper shows that in the real world of drug discovery, there is almost never just one perfect chemical solution that works.

Meng: That’s the core problem; we're currently testing models based on rote memorization rather than their ability to actually solve a problem creatively.

Tom: And Jane, you mentioned this "one-to-many" concept—how does that translate into practical use for finding new medicines?

Jane: It means that when a chemist has a target property, like higher solubility, there are often multiple different molecular shapes they can take to achieve it.

Lu: This approach recognizes the inherent flexibility in chemistry, which is something prior models completely missed because their training data was too rigid.

Meng: I see this as a huge win for usability; we aren't forcing users to narrow down their requirements into one single answer anymore.

Lalam: This allows AI to operate within the complexity of nature itself, serving the full breadth of human inquiry instead of just trying to fit into a pre-defined box.

Tom: So, by acknowledging that this is a dynamic relationship, we set the stage for testing something truly novel with all those ideas.

Methodology and Structure: Jane: To move beyond the single-answer limitation, “Speak-to-Structure” proposes three specific tasks to test AI capabilities: MolEdit, MolOpt, and MolCustom.

Tom: They’ve mapped these tasks directly onto real stages of drug discovery, which is incredibly smart because it makes the evaluation relevant.

Lu: For example, the design of "MolEdit" subtasks really tests if an AI understands chemical grammar and can make a localized change without breaking the whole structure.

Meng: When I look at "MolOpt," it feels like a much tougher test because the the model has to satisfy both making a specific structural change and achieving an improvement in a chemical property simultaneously.

Tom: That’s right, and then we have "MolCustom" which is pure creativity—designing brand new molecules from constraints like atom count or functional groups.

Jane: All three are designed to push the AI beyond simple pattern recognition into genuine structural reasoning and design capability.

Lu: The authors are suggesting that if an AI can pass these, it has mastered the fundamentals of chemical design itself, not just repeating examples.

Meng: This provides us with much clearer data on which models truly possess this deep understanding versus those who aren't capable of it.

Lalam: This framework allows us to define a new standard for intelligent assistance in science, ensuring that AI can handle the nuanced demands of complex discovery processes.

Tom: These structured tasks give us a clear path toward understanding how the AI actually works, which is a great lead-in to discussing how they built this data.

Data Construction and Evaluation: Tom: The researchers identified that human labeling is incredibly expensive for building large datasets, so they tackled that bottleneck head-on by using automated chemical toolkits.

Jane: They bypassed the need for manual annotation by programmatically generating millions of instruction-molecule pairs from databases like PubChem.

Lu: The idea of generating such a vast number of instruction sets shows an incredible level of systematic thinking about data generation.

Meng: This approach addresses the "data hunger" problem, meaning we no longer have to rely on slow, expensive human annotation for future training runs.

Tom: And Jane, when the AI generates a molecule using these tools, how do we know if it’s actually good? They put a lot of effort into automated evaluation.

Jane: They use metrics like Tanimoto Similarity to check if the molecule is a *rational* edit, which is much better than just checking for an exact match on its own.

Lu: That distinction between being 'correct' and being 'rationally related' is key, especially in MolOpt where you need that structural connection.

Meng: It’s about making sure the generated output isn't just random noise; it has to be chemically valid and relevant to the input, which is a huge operational improvement.

Lalam: This rigorous evaluation framework ensures that our tools are not just hitting targets, but are actually achieving meaningful progress in human endeavors.

Tom: These methods of building data and scoring results provide a solid foundation for understanding the next steps in AI training.

Conclusion and Future Outlook: Jane: We've seen how “Speak-to-Structure” moves beyond simple pattern matching, setting a new standard for evaluating LLMs in molecular design.

Tom: The findings clearly demonstrate that if we want truly capable AI, we need massive, high-quality instruction tuning on tasks that actually move past memorization.

Lu: I'm just thrilled about the potential of scaling this work; the researchers are only scratching the surface of what these models could achieve.

Meng: I think the practical impact is enormous; as a tool, it will help us accelerate drug discovery by allowing complex chemical tasks to be handled more reliably.

Lalam: We are seeing how AI can transition from merely reflecting existing data to actively synthesizing new capabilities for humanity.

Tom: Before we wrap up, let's get our final thoughts from the whole team on this groundbreaking research.

Lu: I'm just thrilled about the potential of scaling this work; we're only scratching the surface of what this could achieve.

Meng: The engineering path forward looks much clearer when knowing which models are capable and how they need to be trained.

Lalam: I feel that this allows us to see AI as a true partner in science, not just a passive data processor.

Tom: Thanks for joining us today! We're looking forward to the next paper on arXiv, but it was a fascinating discussion of "Speak-to-Structure" with all of you.

Jane: It truly is the kind of work that shows what AI can be, moving beyond pattern matching to genuine chemical reasoning and discovery.

Jiatong Li, Junxian Li, Weida Wang, Yunqing Liu, Changmeng Zheng, Yatao Bian, Dongzhan Zhou, Xiao-Yong Wei, Qing Li

Hong Kong Polytechnic University · Shanghai Jiao Tong University · Shanghai AI Lab · National University of Singapore

cs.CL

Submitted: 2026-05-22

Updated: 2026-08-25

Code: https://github.com/phenixace/S2TOMG-Bench

Importance score: 88/100

The gist: This paper introduces Speak-to-Structure (S2-Bench), the first benchmark designed to evaluate Large Language Models (LLMs) in "open-domain natural language-driven molecule generation." It addresses a

Key concepts

One-to-many concept
In chemistry and drug discovery, a target property (like higher solubility) can often be achieved by multiple different molecular shapes. This concept recognizes the inherent flexibility in chemistry, moving beyond the assumption that only one perfect chemical solution exists.
Speak-to-Structure tasks
To test AI's capabilities beyond simple memorization, the paper proposes three specific tasks: MolEdit (localized structural change), MolOpt (structural change plus property improvement), and MolCustom (designing new molecules from constraints).
Tanimoto Similarity
This is an automated evaluation metric used to check if a generated molecule is a *rational* edit. It is crucial because it determines if the output is chemically valid and structurally related to the input, rather than just checking for an exact match.

Terminology

Summary

This paper introduces Speak-to-Structure (S2-Bench), the first benchmark designed to evaluate Large Language Models (LLMs) in open-domain natural language-driven molecule generation. It addresses a critical gap in existing research where molecule-text alignment datasets are predominantly built on one-to-one mappings, which measure a model's ability to retrieve a single predefined answer rather than its creative potential to generate diverse, yet equally valid, molecular candidates. This matters because real-world molecule discovery involves one-to-many relationships where multiple distinct structures can satisfy the same desired properties.

The S2-Bench Framework

The benchmark is structured around three primary tasks that mirror critical phases of drug and materials science:

  1. Molecule Editing (MolEdit): Tests the ability to perform precise, localized structural modifications while preserving main structures, analogous to lead optimization.

  2. Molecule Optimization (MolOpt): Evaluates the ability to refine a lead compound under specified property constraints, such as increasing solubility or reducing toxicity.

  3. Customized Molecule Generation (MolCustom): Challenges models to synthesize a novel molecule from scratch based on quantitative and qualitative constraints, such as a specific number of atoms, bonds, or functional groups.

To ensure objective assessment, the authors developed an automated evaluation system using cheminformatics toolkits to verify validity, check constraint satisfaction, and quantify quality through task-specific metrics like Tanimoto Similarity and novelty.

OpenMolIns Instruction Tuning

To facilitate model training, the researchers introduced OpenMolIns, a large-scale instruction tuning dataset comprising up to 1.2 million instruction-molecule pairs programmatically constructed from the PubChem database. Unlike previous datasets that rely on scarce and costly human annotations, OpenMolIns uses automated chemical toolkits to create diverse, one-to-many training examples. The dataset is provided in five distinct scales—light, small, medium, large, and xlarge—allowing researchers to systematically analyze the data scaling law for molecular generation. This approach ensures the model learns genuine chemical reasoning rather than simply memorizing specific input-output pairs.

Key Research Findings

Through an extensive evaluation of 31 LLMs, the study reveals several critical insights:

** Current LLMs lack the structural understanding necessary for precise molecule generation, particularly in the MolCustom task, where they struggle to satisfy fine-grained structural constraints. **

** Instruction tuning is indispensable for guiding and optimizing capabilities; for instance, Llama3.1-8B fine-tuned on the xlarge OpenMolIns dataset surpassed powerful proprietary models like GPT-4o and Claude-3.5. **

** Existing one-to-one mapping datasets cause LLMs to rely on pattern recognition and recall rather than true comprehension, as evidenced by models fine-tuned on ChEBI-20 performing worse on S2-Bench than general LLMs. **

** The data scaling law is task-dependent; while complex de novo synthesis (MolCustom) benefits immensely from larger datasets, simpler tasks like MolEdit may be constrained more by model capacity than by dataset size. **

Conclusion and Impact

The paper concludes that by shifting the focus from simple pattern recall to realistic molecular design, S2-Bench provides a more accurate measure of an LLM's potential in molecule discovery. The work aims to pave the way for more capable LLMs in natural language-driven molecule discovery, ultimately assisting chemists in streamlining the discovery of new pharmaceuticals and materials.

Improvements for AI systems

To improve AI systems based on the findings and methodologies in this paper, I would implement the following specific architectural and training improvements:

  1. Implement a One-to-Many Instruction Tuning Paradigm using Programmatic Data Generation.

  2. Integrate a Multi-Objective Weighted Success Rate (WSR) Reward Function for Reinforcement Learning from Human Feedback (RLHF).

  3. Develop Discrete Constraint-Aware Reasoning Modules for Molecular Synthesis.


Sources

Related papers