LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction".
Jane: The paper was written by Andre Niyongabo Rubungo, Kangming Li, Jason Hattrick-Simpers and Adji Bousso Dieng from Princeton University and Vertaix and University of Toronto and Acceleration Consortium and Vector Institute for Artificial Intelligence and Schwartz Reisman Institute for Technology and Society.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re cracking open a paper that’s got a title that’s a mouthful: “LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction.” Jane, I’ve got to say, just reading that title gets me excited.
Jane: It really does, Tom. And for our listeners who might not be deep in the materials science world, let’s break that down. We’re talking about using those big AI models, the ones that write poems and code, to actually predict things like how much energy it takes to make a new battery material or whether a crystal will be a good conductor.
Tom: Right, and the paper is essentially saying, “Hey, everyone is using these models, but nobody has a standard way to test how good they actually are.” It’s like if every car manufacturer used a different crash test. You couldn’t compare them.
Jane: Exactly. So they built this massive test. I mean, we’re talking about nearly two million crystal structures from ten different public databases. That’s not a small sample size.
Tom: And it’s not just about the size, it’s about the variety. They’ve got everything from simple salts to complex metal-organic frameworks. They’re testing the models on forty-five different properties. We’re talking about band gaps, formation energies, even how much gas a material can soak up.
Lu: And that’s the crucial part, Tom. The title says “benchmarking,” but the real story is in the results. They didn’t just build a bigger test; they used it to ask a very pointed question: are these general-purpose AI models actually any good at this specific, highly technical job?
Jane: And the answer, Lu, seems to be a bit of a reality check. It’s not that they’re useless, but the paper shows that a smaller, specialized model trained just for this task beats the big general-purpose ones by a mile.
Tom: So the title is almost a challenge. It’s saying, “Here’s the arena, here are the rules. Let’s see who really wins.” And the winners are not who you’d expect from the hype.
Jane: It’s a fantastic setup. We have a standard, we have a clear winner, and we have a clear path forward for anyone who wants to actually use these tools for real discovery. That’s what I’m most excited to dig into next.
Summary: Tom: So, we’ve set the stage with the title. Now, let’s get into the meat of the paper, “LLM4Mat-Bench.” Jane, what did they actually do?
Jane: They did a huge experiment. They took these models and fed them the material’s information in three different ways. First, just the chemical formula, like NaCl for salt. Second, the full crystal structure file, which is a bunch of numbers and coordinates. And third, a written description of the structure, like a paragraph a scientist might write.
Tom: And that’s a key insight, right? The way you present the data to the model changes everything.
Jane: Absolutely. And what they found was that the models, the big chatty ones like Llama and Mistral, they really struggled with the raw data files. They’d often just make stuff up or refuse to give a number at all.
Meng: I saw that in the paper. They called it “hallucination.” For a lot of the tasks, the model would just output something that wasn’t even a valid property value. From an engineer’s standpoint, that’s a non-starter. You can’t build a pipeline on a model that sometimes just doesn’t answer.
Tom: But when they gave the model a nice, human-readable description of the crystal, it did a little better. Still not great, but better.
Jane: Right. But here’s the kicker. They also tested a much smaller model, called LLM-Prop, that was specifically trained on this type of data. And that little guy blew the big models out of the water.
Lu: It’s a classic lesson in specialization. The big models have a lot of general knowledge, but they haven’t been fine-tuned for this specific task. The smaller model has one job, and it does it exceptionally well. The paper shows that on many tasks, it was even better than a traditional graph neural network, which is the old standard for this kind of prediction.
Meng: And that’s a huge deal for practicality. The smaller model is faster, cheaper to run, and more reliable. The paper mentions it’s about two hundred times smaller than some of the big ones. For a startup or a lab with limited compute, that’s the difference between being able to do the work and not.
Tom: So the summary is: bigger isn’t better, and the way you talk to the model matters. It’s a really clean, clear result that challenges a lot of the current hype.
Jane: It really does. And it makes you wonder, if we can get these results just by fine-tuning a small model, what could we achieve with a model that’s actually designed for this from the ground up? That’s the exciting question we should look at next.
Improvements: Tom: We’ve seen the results, and they’re pretty clear. But what does the paper suggest we do about it? What’s the path forward? Jane, you’ve been looking at this.
Jane: The paper doesn’t just say, “Our model is best.” It points to a whole new direction. The main takeaway is that we need task-specific models. We can’t just grab a general chatbot and expect it to be a materials scientist.
Lu: And that’s where the “improvements” come in. They’re not just suggesting we train bigger models. They’re suggesting we need to think about how we represent the data. The paper shows that the textual descriptions are much better for LLMs than raw CIF files. So, a big improvement would be developing better, more standardized ways to describe crystal structures in text.
Meng: That’s a practical problem I can get behind. The current descriptions are generated by a tool, and they can be thousands of words long. That’s a lot of tokens to process. The paper even notes that the models perform better on datasets with shorter descriptions. So, there’s a real engineering challenge in making these descriptions more concise and information-dense.
Tom: So it’s not just about the model, it’s about the data we feed it. We need to speak the model’s language.
Jane: Exactly. And the paper also highlights the need for instruction-tuning. The big chat models failed because they didn’t follow the simple instruction to just output a number. They’d write a paragraph or get confused. So, a key improvement is fine-tuning these models to follow specific output formats, to be more reliable.
Lu: And I think the most exciting improvement they hint at is using this benchmark to develop models that can do more than just predict. If a model can understand the structure and properties, it could potentially help us design new materials. You could ask it, “What’s a stable structure with a high band gap?” and it could help you search for it.
Meng: But to get there, we need the reliability first. We need models that don’t hallucinate a value for a material that doesn’t exist. This benchmark is the first step to measuring that reliability and forcing the field to improve.
Tom: So the improvements aren’t just about tweaking a model. It’s about a whole ecosystem: better data, better fine-tuning, and better evaluation.
Jane: It’s a roadmap. And it’s a roadmap that’s grounded in a massive, shared resource that everyone can use. That’s what makes this paper so important. It’s not just a result; it’s a foundation for the next wave of research.
Conclusion: Tom: Well, we’ve had a great time unpacking “LLM4Mat-Bench: Benchmarking Large Language Models for Materials Property Prediction.” It’s one of those papers that feels like a turning point.
Jane: It really does. We started by looking at the title and seeing a new benchmark. Then we dug into the results and found that smaller, specialized models are the real workhorses. And finally, we saw the roadmap for the future, which is all about better data and more reliable models.
Tom: And the implications are huge. This isn’t just an academic exercise. This is about accelerating the discovery of new batteries, better solar panels, and stronger materials. Every time we can predict a property faster and more accurately, we save years of lab work.
Jane: For me, the most important takeaway is that this benchmark gives the whole community a common language. Now, when someone says their model is good, we can say, “Okay, let’s see how it does on LLM4Mat-Bench.” That’s how a field matures.
Lu: It also democratizes the research. Because the best models are small, more labs can actually run them. They don’t need a supercomputer to make a meaningful contribution.
Meng: And from a practical standpoint, having a reliable, fast, and cheap way to predict properties is a game-changer for anyone trying to build real-world products.
Tom: So, with that, we’re going to say goodbye to this paper. It’s been a fantastic discussion. To all our listeners, if you’re working on materials or AI, this is a benchmark you need to know about.
Jane: Absolutely. Thanks for joining us, and we’ll see you next time with another exciting paper. Take care, everyone.
Andre Niyongabo Rubungo, Kangming Li, Jason Hattrick-Simpers, Adji Bousso Dieng
Princeton University · Vertaix · University of Toronto · Acceleration Consortium · Vector Institute for Artificial Intelligence · Schwartz Reisman Institute for Technology and Society
cond-mat.mtrl-sci, cs.CL
Submitted: 2024-11-30
Updated: 2026-08-17
Comments: Accepted at NeurIPS 2024-AI4Mat Workshop. The Benchmark and code can be found at https://github.com/vertaix/LLM4Mat-Bench
Code: https://github.com/vertaix/LLM4Mat-Bench
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 72/100
The gist: LLM4Mat-Bench is the largest benchmark to date for evaluating the performance of large language models (LLMs) in predicting the properties of crystalline materials.
Terminology
Summary
LLM4Mat-Bench is the largest benchmark to date for evaluating the performance of large language models (LLMs) in predicting the properties of crystalline materials. It contains about 1.9M crystal structures in total, collected from 10 publicly available materials data sources, and 45 distinct properties. LLM4Mat-Bench features different input modalities: crystal composition, CIF, and crystal text description, with 4.7M, 615.5M, and 3.1B tokens in total for each modality, respectively. The benchmark is used to fine-tune models with different sizes, including LLM-Prop and MatBERT, and to provide zero-shot and few-shot prompts to evaluate the property prediction capabilities of LLM-chat-like models, including Llama, Gemma, and Mistral. The results highlight the challenges of general-purpose LLMs in materials science and the need for task-specific predictive models and task-specific instruction-tuned LLMs in materials property prediction.
The dataset comprises approximately two million samples, sourced from ten publicly available materials sources, each containing between 10K and 1M structure samples. LLM4Mat-Bench encompasses several tasks, including the prediction of electronic, elastic, and thermodynamic properties based on a material’s composition, crystal information file (CIF), or textual description of its structure. The data sources include hMOF, Materials Project, OQMD, OMDB, JARVIS-DFT, QMOF, JARVIS-QETB, GNoME, Cantor HEA, and SNUMAT. Crystal structure files (CIFs), material compositions, and material properties were collected from these publicly accessible sources, facilitated by APIs and direct download links. For databases such as Materials Project, OMDB, SNUMAT, JARVIS-DFT, and JARVIS-QETB, user registration is required for access, while databases like hMOF, QMOF, OQMD, and GNoME allow direct data access without registration.
Crystal structures are typically described in file formats such as Crystallographic Information File (CIF) which include predominantly numbers describing lattice vectors and atomic coordinates and are less amenable to LLMs. Instead of directly using these as inputs, the authors use Robocrystallographer to deterministically generate texts that are more descriptive of crystal structures from CIF files. Robocrystallographer leverages predefined rules and existing libraries to extract chemical and structural information, including oxidation states, global structural descriptions (symmetry information, prototype matching, structural fingerprint calculations etc.), and local structural descriptions (e.g. bonding and neighbor analysis, connectivity). This method not only generates deterministic and human-readable texts, but also ensures no data contamination in the fine-tuned LLMs, as the data sources mentioned do not include these crystal text descriptions.
The total samples for each dataset in LLM4Mat-Bench are randomly split into 80%, 10%, and 10% for training, validation, and testing, respectively. OQMD has the highest number of samples at 964,403, while QMOF has the fewest with 7,656 samples. On average, each dataset in LLM4Mat-Bench contains approximately 200,000 samples. In LLM4Mat-Bench, when combined, textual descriptions contain 3.1 billion tokens, crystal structures 615 million, and compositions 4.7 million. OQMD leads in composition tokens (964K), while hMOF has the most description tokens (581M). For CIFs, both OQMD and hMOF have around 96M tokens. On average, compositions have 8 subword tokens per sample, CIFs 1600, and descriptions 1700. hMOF averages the longest inputs for compositions (14.9) and descriptions (5629), while QMOF leads in structures (5876.4). JARVIS-DFT has the most tasks with 20 properties, followed by Materials Project with 10, and OMDB with one.
The authors conducted about 1,235 experiments, evaluating the performance of five models and three material representations on each property for each data source. Consistent with standard practices in materials science, they evaluated performance separately for each data source rather than combining samples from different sources for the same property. This approach accounts for variations in techniques and settings used by different data sources, which can result in discrepancies, such as differing band gaps for the same material. The models benchmarked include CGCNN (a GNN baseline), MatBERT (a BERT-base model with 109 million parameters, pretrained on two million materials science articles), LLM-Prop (a model based on the encoder part of T5-small model with 35 million parameters), and Llama, Gemma, and Mistral variants (3B, 7B, 8B, and 9B parameters).
The main observations from the results are as follows: Small, task-specific predictive LLMs exhibit significantly better performance than larger, generative general-purpose LLMs. This performance disparity is evident across both regression and classification tasks on all 10 datasets. Specifically, LLM-Prop and MatBERT outperform conversational LLMs by a substantial margin, despite being approximately 200 and 64 times smaller in size, respectively. In regression tasks, LLM-Prop achieves the highest accuracy on 8 out of 10 datasets, with MatBERT leading on the remaining 2 datasets. For classification tasks, both LLM-Prop and MatBERT deliver the best performance on 1 out of 2 datasets. LLM-Prop surpasses MatBERT by 1.8% on the SNUMAT dataset, whereas MatBERT outperforms LLM-Prop by 0.8% on the other dataset.
General-purpose generative LLMs hallucinate and often fail to generate valid property values. As shown in the results, Llama, Gemma, and Mistral models produce invalid outputs on multiple tasks, where the expected property value is missing. This issue occurs less frequently when the input is a description or chemical formula, but more commonly when the input is a CIF file. One reason may be that descriptions and chemical formulas resemble natural language, which LLMs can more easily interpret compared to CIF files. This may also indicate that when the input modality during inference differs significantly from the modalities encountered during pretraining, fine-tuning is necessary to achieve reasonable performance. Another key observation is that these models often generate the same property value for different inputs (i.e. hallucinate), contributing to their poor performance across multiple tasks.
Representing materials with their textual descriptions improves the performance of LLM-based property predictors compared to other representations. A significant performance improvement is observed when the input is a description compared to when it is a CIF file or a chemical formula. One of the possible reasons for this might be that LLMs are more adept at learning from natural language data. On the other hand, although material compositions appear more natural to LLMs compared to CIF files, they lack sufficient structural information. This is likely why LLMs with CIF files as input significantly outperform those using chemical formulas.
More advanced, general-purpose generative LLMs do not necessarily yield better results in predicting material properties. The results indicate that, despite being trained on substantially larger and higher-quality datasets, more advanced versions of generative LLMs show limited improvements in performance and validity of predictions for material properties. For instance, Llama 3 and 3.1 8b models were trained on over 15 trillion tokens—around eight times more data than the 2 trillion tokens used for the Llama 2 7b models. This finding highlights the ongoing challenges of leveraging LLMs in material property prediction and underscores the need for further research to harness the potential of these robust models in this domain.
The performance on energetic properties is consistently better across all datasets compared to other properties. This is consistent with the trend observed in the community benchmarks such as MatBench and JARVIS-Leaderboard, where energetic properties are among those that can be most accurately predicted. This is not surprising because energy is known to be relatively well predicted from e.g., compositions and atom coordination (bonding), which is inherently represented in GNNs and also presented in text descriptions.
Task-specific predictive LLM-based models excel with shorter textual descriptions, while CGCNN performs better on datasets with longer descriptions. For regression tasks, LLM-Prop outperforms CGCNN on only 4 out of 10 datasets (MP, JARVIS-DFT, JARVIS-QETB, and SNUMAT), and MatBERT outperforms CGCNN on just 2 out of 10 datasets (MP and JARVIS-QETB). In contrast, CGCNN achieves the best performance on 5 out of 10 datasets (GNoME, hMOF, Cantor HEA, OQMD, and OMDB). Further analysis reveals that CGCNN tends to perform better than LLM-based models on datasets with relatively longer textual descriptions, while LLM-based models excel on datasets with shorter descriptions. The performance gain on shorter descriptions may stem from LLM-based models’ ability to leverage more context from compact text, while CGCNN consistently benefits from training on the entire crystal structure.
The authors conclude that LLMs are increasingly being utilized in materials science, particularly for materials property prediction and discovery. However, the absence of standardized evaluation benchmarks has impeded progress in this field. They introduced LLM4Mat-Bench, a comprehensive benchmark dataset designed to evaluate LLMs for predicting properties of atomic and molecular crystals and MOFs. Their results demonstrate the limitations of general-purpose LLMs in this domain and underscore the necessity for task-specific predictive models and instruction-tuned LLMs tailored for materials property prediction. These findings emphasize the importance of using LLM4Mat-Bench to advance the development of more effective LLMs in materials science.
Due to computational constraints and the number of experiments, the authors were unable to conduct thorough hyperparameter searches for each property and dataset. The reported settings were optimized on the MP dataset and then fixed for other datasets. Additionally, they could not include results from SOTA commercial LLMs such as GPT-4o or Claude 3.5 Sonnet due to budget constraints. They also encountered issues with chat-based models, which sometimes failed to follow the output format, producing invalid or incomplete outputs. Extracting property values was therefore challenging. Furthermore, they did not include comparisons with dataset-specific retrieval-augmented generation (RAG) models, such as the recently developed LLaMP, a RAG-based model tailored for interaction with the MP dataset.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities:
1. Improved AI System: Task-Specific Predictive LLM for Materials Property Prediction
-
Improvements:
-
Architecture: Use a small, encoder-only transformer (e.g., T5-small, 35M parameters) instead of large generative models. This is based on the finding that LLM-Prop and MatBERT significantly outperform larger chat models.
-
Input Representation: Accept and process three distinct input modalities: chemical composition, CIF structure, and text description generated by Robocrystallographer. The system should prioritize text descriptions as they yield the best performance.
-
Training: Fine-tune the model on the LLM4Mat-Bench dataset. Use the provided train/validation/test splits to ensure reproducibility. For CIF inputs, implement xVal encoding to handle numerical values efficiently.
-
Evaluation: Use the MAD:MAE ratio for regression tasks and AUC for classification tasks, as defined in the paper.
-
Capabilities:
-
Accurate Property Prediction: Predict 45 distinct properties (e.g., band gap, formation energy, elastic moduli) with high accuracy, outperforming general-purpose LLMs by a significant margin (e.g., achieving MAD:MAE ratios above 5.0, which is considered
good
). -
Multi-Modal Input: Can predict properties from a simple chemical formula (e.g.,
NaCl
), a full CIF file, or a natural language description of the crystal structure. -
Efficient and Fast: Due to its small size, it can be trained and deployed much faster than large models (e.g., 2.5 days for 300K data points vs. 0.5 days for inference on 40K samples for a 7B model).
-
Reliable Output: It will not hallucinate or produce invalid outputs, a common failure mode of generative models, especially when processing CIF files.
2. Improved AI System: Reliable Property Extraction from Generative LLMs
3. Improved AI System: Benchmarking and Model Selection Tool
Abstract
Large language models (LLMs) are increasingly being used in materials science. However, little attention has been given to benchmarking and standardized evaluation for LLM-based materials property prediction, which hinders progress. We present LLM4Mat-Bench, the largest benchmark to date for evaluating the performance of LLMs in predicting the properties of crystalline materials. LLM4Mat-Bench contains about 1.9M crystal structures in total, collected from 10 publicly available materials data sources, and 45 distinct properties. LLM4Mat-Bench features different input modalities: crystal composition, CIF, and crystal text description, with 4.7M, 615.5M, and 3.1B tokens in total for each modality, respectively. We use LLM4Mat-Bench to fine-tune models with different sizes, including LLM-Prop and MatBERT, and provide zero-shot and few-shot prompts to evaluate the property prediction capabilities of LLM-chat-like models, including Llama, Gemma, and Mistral. The results highlight the challenges of general-purpose LLMs in materials science and the need for task-specific predictive models and task-specific instruction-tuned LLMs in materials property prediction.
Sources
- GPT-4 Technical Report
- Crystal Structure Generation with Autoregressive Large Language Modeling
- LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- The Llama 3 Herd of Models
- Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files
- Mistral 7B
- Probing out-of-distribution generalization in machine learning for materials
- ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
- LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions
- Gemma: Open Models Based on Gemini Research and Technology
- Gemma 2: Improving Open Language Models at a Practical Size
- What Information is Necessary and Sufficient to Predict Materials Properties using Machine Learning?
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- DARWIN Series: Domain Specific Large Language Models for Natural Science
Related papers
- AES-Debye: an Accurate, Efficient, and Scalable Engine for Debye Scattering Calculations
- Cooperative Quantum Optical Effects of Moir'e Exciton Superlattices
- Imaging Surface Magnetization in Altermagnetic MnTe Films
- Accidental accuracy and formal consistency in GW +BSE: Exact benchmarks and regime-dependent error cancellation
- Modifying van der Waals Materials via Cavity Vacuum Fluctuations
- Linear dichroic soft X-ray microscopy of ferroelectric stripe domains in epitaxial K 0.6 Na 0.4 NbO 3