Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery

arXiv:2502.14631 · cs.LG · Submitted 2025-02-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery".

Jane: The paper was written by Minh-Quyet Ha, Dinh-Khiet Le, Duc-Anh Dao, Tien-Sinh Vu, Duong-Nguyen Nguyen et al. from Japan Advanced Institute of Science and Technology and HPC SYSTEMS Inc. and National Institute for Materials Science.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone! Tom here, and I've got Jane with me. Jane, we're looking at a paper that's got one of those titles that sounds like it could be from three different fields at once: "Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery."

Jane: Tom, that title is a mouthful, but it's actually describing something really clever. These researchers are trying to discover new high-entropy alloys, which are basically metals made of five or more elements mixed together in roughly equal amounts. And they're using a mathematical framework called Dempster-Shafer theory to combine knowledge from two very different places.

Tom: Right, and those two places are what really caught my eye. One source is computational material datasets, which are basically huge tables of alloy compositions and their properties. The other source is domain knowledge distilled from scientific literature using large language models, specifically GPT-4o.

Jane: Exactly. So they're asking GPT-4o questions like "Can copper and manganese be substituted for each other?" and getting answers based on what the model learned from reading thousands of scientific papers. Then they're combining those answers with patterns found in the actual dataset.

Tom: And the key insight here is something called elemental substitutability. The idea is that if two elements behave similarly in alloys, you can swap one for the other and get similar properties. That's how they explore new compositions without having to test every single combination.

Jane: Which is huge, Tom. There are millions of possible alloy combinations out there. You can't just test them all in a lab. But if you know that nickel and cobalt are often interchangeable, you can focus your search on promising regions of that compositional space.

Tom: So the title is really saying: we're fusing knowledge from data and from language models, using evidence theory to handle uncertainty, all to find new high-entropy alloys. And Jane, I think the most exciting part is that this approach actually works better than traditional methods when you're dealing with elements that weren't in the training data.

Jane: That's the part I want to dig into. Because that's the difference between a model that just interpolates and a model that can actually extrapolate to new territory. Let's keep going and talk about what they actually found.

Summary: Tom: So Jane, we're still on "Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery," and I want to get into the actual results. They tested this on four datasets of quaternary alloys, which means alloys with four elements each.

Jane: Right, and two of those datasets were about phase stability, meaning whether the alloy forms a single solid solution or separates into multiple phases. The other two were about magnetic properties, specifically magnetization and Curie temperature. And they ran two types of experiments.

Tom: The first was cross-validation, where they varied the training set size from just one percent up to thirty percent of the data. And here's what's interesting: when the training data was tiny, the models that used only the language model knowledge actually did better than the ones using only the material dataset.

Jane: That makes sense, Tom. When you have almost no data, the domain knowledge from GPT-4o can fill in the gaps. It's like having an expert consultant who's read thousands of papers, even if you've only run a handful of experiments yourself.

Tom: But as the training data grew, the models using only the material dataset caught up and sometimes surpassed the language model ones. And the multi-source model, which combined both, stayed competitive throughout. It was never the absolute best at any single point, but it was consistently good.

Jane: And then came the extrapolation experiment, which is where things got really exciting. They removed all alloys containing a specific element from the training data, then tested on those alloys. So the model had never seen, say, osmium in any training example.

Tom: And the multi-source model hit an accuracy of zero point eight seven on the stability dataset, compared to zero point five zero for the model using only the material dataset. That's basically a coin flip versus a strong prediction. The language model knowledge was doing the heavy lifting there.

Jane: Because when you've never seen osmium, the dataset alone gives you nothing. But GPT-4o has read about osmium's chemical similarity to other transition metals, so it can reason about how osmium might behave in an alloy.

Tom: Exactly. And the AUC scores, which measure overall discriminative power, were zero point nine three for the multi-source model on that same dataset. The language model alone got zero point nine one, and the material dataset alone got zero point five zero, which is random chance.

Jane: So the summary is: combining sources helps, but the real magic happens when you need to reason about things you've never seen before. That's the difference between memorizing and understanding.

Improvements: Tom: Jane, we're back on "Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery," and I want to talk about what this paper actually improves over existing approaches. Because it's not just about accuracy numbers.

Jane: Right, Tom. The big improvement here is interpretability. Most machine learning models for materials are black boxes. You feed in a composition, you get out a prediction, but you have no idea why. This framework actually tells you which elements are substitutable for each other.

Tom: And they show that in a really visual way. They built hierarchical clustering trees of the elements based on substitutability. And the results match what metallurgists already know: copper, silver, and gold cluster together. Early transition metals like tantalum and niobium cluster together. Late transition metals like iron and cobalt cluster together.

Jane: But here's the improvement I find most compelling. They identified a set of fourteen transition metals that they call E. And when you form alloys exclusively from these elements, ninety-nine percent of them form high-entropy alloys. That's a stunning success rate.

Tom: Ninety-nine percent is remarkable. And they showed that nearly all high-entropy alloys in their dataset contain at least one element from this set. So these fourteen metals are basically the foundation of high-entropy alloy stability.

Jane: And they went further. They used this substitutability information to build what they call alloy maps, which are essentially visualizations of the compositional space. They colored regions based on whether alloys there are likely to form single phases or multiple phases.

Tom: And when they added osmium-based alloys to the map, which were completely absent from training, the map reorganized in a way that made physical sense. The osmium alloys clustered with the other transition metals, confirming that osmium behaves similarly to its neighbors in the periodic table.

Jane: So the improvement isn't just "we predict better." It's "we can show you why we predict what we predict, and those reasons align with physical intuition." That's huge for materials scientists who need to trust the model before they spend resources on actual experiments.

Tom: And there's a practical angle too. The paper notes that the multi-source model occasionally underperforms at intermediate training sizes, which suggests the evidence integration needs calibration. But that's a refinement problem, not a fundamental flaw.

Jane: Exactly. The framework is solid, but the weighting between sources might need to adapt dynamically. That's a natural next step for the authors.

Conclusion: Tom: Alright Jane, we've covered a lot of ground on "Synergistic Fusion of Multi-Source Knowledge via Evidence Theory for High-Entropy Alloy Discovery." Let me try to pull it together.

Jane: Please do, Tom. Because there's a lot here, and I want to make sure we capture the big picture.

Tom: So the core idea is that you can discover new high-entropy alloys by understanding which elements can substitute for each other. And you can learn that substitutability from two sources: computational datasets and language models that have read the scientific literature.

Jane: And the key finding is that combining both sources gives you robust performance, especially when you're trying to predict properties for alloys containing elements you've never seen before. The language model knowledge compensates for missing data.

Tom: Right. And the interpretability is what sets this apart. You're not just getting predictions. You're getting a map of elemental relationships that makes physical sense. You're getting a set of fourteen transition metals that are the backbone of high-entropy alloy stability.

Jane: And that's actionable. If you're designing a new alloy, you know which elements to start with. You know which substitutions are likely to preserve the single-phase structure. That's a huge head start.

Tom: There are limitations, of course. The datasets are computational predictions, not experimental results. And the language model knowledge might not align perfectly with every physical property, like magnetism, which the paper showed some gaps on.

Jane: But the framework is extensible. You could add more knowledge domains, more data sources, more physical properties. And the authors suggest adaptive weighting as a future direction, which could address the intermediate training size issues.

Tom: So we're saying goodbye to this paper, but the ideas are going to stick with us. Combining data-driven evidence with expert knowledge, using evidence theory to handle uncertainty, and making the whole process interpretable. That's a recipe that could accelerate materials discovery across many fields.

Jane: Absolutely, Tom. And I'm excited to see where this goes next. Maybe we'll see this applied to other material classes, or integrated into active learning loops where the model suggests which experiments to run next.

Tom: That would be something. For now, thanks for joining us, and we'll see you on the next paper.

Jane: Take care, everyone.

Minh-Quyet Ha, Dinh-Khiet Le, Duc-Anh Dao, Tien-Sinh Vu, Duong-Nguyen Nguyen, Viet-Cuong Nguyen, Hiori Kino, Van-Nam Huynh, Hieu-Chi Dam

Japan Advanced Institute of Science and Technology · HPC SYSTEMS Inc. · National Institute for Materials Science

cs.LG

Submitted: 2025-02-20

Updated: 2026-08-18

Comments: 13 pages, 7 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

Key concepts

High-Entropy Alloys
These are metals composed of five or more different elements mixed together in roughly equal amounts. Researchers use this technique to discover new materials with potentially superior properties.
Evidence Theory (Dempster-Shafer theory)
This is a mathematical framework used to combine knowledge from multiple, different sources while handling uncertainty. It allows researchers to fuse data patterns with expert knowledge effectively.
Elemental Substitutability
This key insight suggests that if two elements behave similarly within alloys, one can be swapped for the other without drastically changing the material's properties. This narrows down vast search spaces.
Interpretability
Unlike 'black box' models, this framework provides clear reasons for its predictions. It shows which elements are substitutable and why, aligning with existing physical knowledge.

Terminology

Summary

Summary

This paper introduces a hybrid framework that integrates multi-source knowledge from large language models (LLMs) and material datasets (MD) to support decision-making in high-entropy alloy (HEA) discovery. The framework centers around the principle of elemental substitutability, providing an interpretable approach to exploring compositional spaces. By aggregating evidence from both empirical data and domain knowledge, the framework effectively addresses challenges related to data scarcity, uncertainty, and exploration.

The methodology is built upon Dempster-Shafer theory (evidence theory) to model and combine substitutability evidence from multiple sources. For material datasets, the framework transforms materials data into substitutability evidence by considering pairs of alloys that share at least one common element. The intersection acts as the context for measuring similarity: if the labels of two alloys agree (both HEA or both ¬HEA), the non-overlapping element combinations are inferred to be substitutable; otherwise, they are non-substitutable. This evidence is represented by mass functions assigning probabilities to subsets of the frame of discernment Ωsim = similar, dissimilar, with a parameter α controlling the confidence in the evidence and the remaining mass assigned to the ambiguous set similar, dissimilar to encode epistemic uncertainty.

For domain knowledge, the framework leverages GPT-4o to distill insights from scientific literature across five key domains: Corrosion Science, Materials Mechanics, Metallurgy, Solid-State Physics, and Materials Science. A two-step prompt structure is used: first asking whether the LLM possesses sufficient knowledge to assess substitutability, and if yes, rating substitutability as High, Medium, or Low. The LLM responses are mapped to mass functions with a parameter β indicating confidence in the GPT response. The evidence from multiple sources is integrated using Dempster's rule of combination with a reliability-aware discounting step, where each source's discount factor γS is computed based on its macro-averaged F1 score in a 10-fold cross-validation setup.

To evaluate hypothetical candidates, the framework applies analogy-based inference. For a new alloy Anew, it identifies subsets Ct of known alloys that, when replaced by Cv, generate Anew. The similarity M[t, v] between Ct and Cv determines the mass assigned to HEA or ¬HEA based on the label of the known alloy, with the remaining mass assigned to the ambiguous set HEA, ¬HEA. Multiple pieces of evidence from different host alloys are combined using Dempster's rule to produce a final mass function supporting decision-making.

The framework is evaluated on four computational datasets of quaternary alloys: D0.9Tm and D1350K (stability predictions at 0.9 Tm and 1350 K, containing 14,950 alloys each), and DMag and DTC (magnetization and Curie temperature, containing 5,968 alloys each). Two complementary experiments are performed: cross-validation with training set sizes varying from 1% to 30%, and extrapolation experiments where alloys containing a specific element are excluded from training and used as the test set.

In cross-validation experiments, the results show that at smaller training sizes (approximately 1%–10%), logistic regression achieves the highest overall accuracy, surpassing evidential models. Among evidential models, single LLM-source models initially outperform MD-source models, likely because the LLM's domain-specific insights help mitigate data limitations. However, as training size exceeds 10%, MD-source models exhibit superior performance on magnetization and Curie temperature datasets, while reaching accuracy levels comparable to LLM-source models on alloy stability datasets. Multi-source models maintain robust performance across all training-set sizes, reflecting the flexibility gained by merging domain-based substitutability perspectives with empirical data. ROC curve analysis shows that multi-source and MD-source models exhibit similar performance and outperform other models, with LLM-source models achieving comparable results on stability datasets but lagging behind on magnetization and Curie temperature datasets.

In extrapolation experiments, the results demonstrate that MD-source models achieve relatively poor accuracies (0.47–0.56) across all datasets, while models incorporating LLM knowledge or multiple evidence sources attain significantly higher accuracies. On D0.9Tm, multi-source and LLM-source models reach 0.87 and 0.86 accuracy respectively, compared to 0.50 for MD-source models. On DMag and DTC, multi-source and LLM-source models surpass both LR-based and MD-source models by a large margin. The multi-source models achieve the best performance overall on all datasets, consistently outperforming LLM-source models by a modest but persistent gap. AUC analysis reveals that multi-source models achieve the highest scores (0.92–0.95) across the four datasets, followed closely by LLM-source models (0.90–0.95), while LR-based models peak around 0.84 and MD-source models hover at 0.50.

Further analysis of element substitutability reveals distinct clusters. The LLM-derived knowledge shows strong consensus across domains regarding coinage metals (Cu, Ag, Au), which consistently exhibit similar substitutability behavior. Early transition metals (blue-labeled) are clearly distinguished from late transition metals (orange-labeled). Aluminum, silicon, and arsenic cluster together, while transition metals from groups 4 and 5 (Ta, Nb, Hf, Ti, Zr) form another distinct cluster. Similar patterns are observed from the material datasets, with coinage metals forming a cohesive group but displaying a negative effect on phase stability. A set E of 14 transition metals (Fe, Co, Ir, Cu, Ni, Pt, Pd, Rh, Au, Ag, Ru, Os, Re, Mn, Ta, Ti, W, Mo, Cr, V, Hf, Nb, Zr) demonstrates consistent substitutability behavior and plays a pivotal role in HEA formation.

The visualization of compositional landscapes using t-SNE with a hybrid distance matrix reveals four distinct groups of HEAs in the dataset D0.9Tm. Group A stands out with 711 HEAs, all formed from a set of 13 critical transition metals (a subset of E), accounting for 99% of all possible quaternary combinations derived from these elements. Groups B and C are characterized by alloys with Nb, Ta, Ti, and Si combined with elements from the subset of E, while Group D comprises 250 palladium-based alloys. After integrating Os-based alloys, the updated map shows four new groups, with Group E containing 1,363 HEAs formed through interactions between Os-based alloys restricted to E and HEAs in Group A. Of the 1,001 quaternary alloys in E, 997 form HEAs, achieving a 99% stability rate.

The analysis of the impact of element set E shows that nearly 98% of all possible HEAs in the dataset contain at least one element from E. When alloys are formed exclusively from elements in this set, 99% form HEAs; however, the stability rate declines as fewer elements from E are included. This strong stabilizing effect persists even at elevated temperatures, with nearly 97% of all possible HEAs still containing at least one element from the set at 1350 K, though the stability rate drops to 61% when only elements from E are used.

The paper concludes that incorporating multi-source knowledge achieves strong performance across various interpolation scenarios, with the framework's strength particularly evident under extrapolation conditions. Models utilizing LLM-derived knowledge excel in data-sparse situations, while MD-based models eventually outperform LLM-based models as training data grows. The findings highlight the pivotal role of the core set of 14 transition metals in stabilizing single-phase HEAs, with their stabilizing effect persisting across various temperature ranges. The paper acknowledges challenges remain in optimizing the integration of multi-source evidence, noting that multi-source models occasionally underperform at intermediate training sizes, suggesting that evidence integration requires careful calibration. Future work should focus on adaptive frameworks that dynamically adjust evidence weighting based on real-time performance validation and expanding the framework to include diverse material properties such as mechanical strength or thermal stability.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Implement a Dempster-Shafer-based fusion layer that combines data-driven evidence (from material datasets) with LLM-derived domain knowledge, using reliability-weighted discounting.

What the improved system can do:

  • Dynamically weight evidence sources based on their historical predictive performance (F1-score in cross-validation)

  • Maintain robust predictions when one source fails (e.g., when training data lacks certain elements)

  • Achieve 0.87–0.92 accuracy in extrapolation scenarios where key elements are absent from training data, versus 0.50 for single-source data-driven models

Improvement: Build a substitutability matrix M[t,v] that quantifies how interchangeable two element combinations are, learned from both empirical data and LLM domain knowledge.

Improvement: Replace point predictions with mass functions over HEA, ¬HEA, uncertain, explicitly modeling epistemic uncertainty.

Improvement: Use structured two-step prompting (knowledge assessment → substitutability rating) to extract domain knowledge from LLMs across five scientific domains.

Improvement: Generate t-SNE embeddings using a hybrid distance matrix that integrates substitutability dissimilarity and Jaccard compositional distance.

Improvement: Implement a dynamic discounting mechanism that adjusts source reliability based on dataset characteristics and training size.

Improvement: Design the inference mechanism to operate on element combinations rather than fixed feature vectors, enabling generalization to unseen elements.

Related papers