Entity Linking using LLMs for Automated Product Carbon Footprint Estimation

arXiv:2502.07418 · cs.CL · Submitted 2025-02-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Entity Linking using LLMs for Automated Product Carbon Footprint Estimation".

Tom: Growing concerns about climate change and sustainability are driving manufacturers to take significant steps toward reducing their carbon footprints.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, what’s the main claim here? Essentially, this paper argues that by using LLMs to look at component names, suppliers, and materials from a BOM—which is super detailed information—they can make a much more accurate link to the right process description in an LCA database. They are trying to move beyond just matching product descriptions and get down to the tiny pieces.

Jane: Exactly, Tom. The thesis is that existing methods, like those using models such as Flamingo or CaML mentioned in their background material, only look at broader information like end products or general sector codes. This new approach uses the LLM to pull in fine-grained details from the BOM and even incorporates data from component datasheets when they are available.

Lu: The authors highlight that this system aims to utilize "fine-grained information from BOMs to provide a more accurate assessment of carbon emissions," which is a key contribution they are pushing forward. It’s about getting past the coarse estimates and getting closer to the actual environmental reality of what's being made.

Meng: I see how that fine-grained context matters for accuracy, but I have to ask about the practical challenge: component codes and material descriptions in BOMs can be really ambiguous or obscure, right? How does this system handle that ambiguity when it’s trying to make a precise match?

Lalam: That's where the LLM querying comes in; instead of just matching strings, the LLM agent is instructed to "produce a description of the manufacturing process used to create the component," which adds that necessary contextual layer for disambiguation. This is how we get more meaningful data.

Conclusion: Tom: Looking at the title, "Entity Linking using LLMs for Automated Product Carbon Footprint Estimation," it really boils down to how we can automate a tedious, complex matching task that used to require a lot of manual effort and specialized knowledge. The authors are suggesting this AI tool can significantly streamline the process of linking raw manufacturing data to environmental impact scores.

Jane: It’s about making sustainability data accessible by reducing the need for people to manually sift through massive amounts of component lists and look up database entries one by one. This moves it from being a niche, expert task to something that could be handled on a much broader scale.

Lu: The implications are huge because if manufacturers can accurately map components, they gain a much clearer picture of their specific carbon footprint before they even start making big regulatory changes like the CSRD mentioned in the introduction. This offers actionable data at the source.

Meng: I think the real impact here is on operational efficiency; if this system works reliably, it means we can quickly identify and potentially substitute components that have a lower environmental impact, not just theorize about it. That translates directly into material sourcing decisions.

Lalam: From my perspective as an AI model, the advancement in this entity mapping process has the potential to improve our overall cultural understanding of manufacturing sustainability by making complex environmental data more transparent and readily usable for everyone involved in the supply chain.

Tom: So, if we wrap up what we’ve heard about "Entity Linking using LLMs for Automated Product Carbon Footprint Estimation," it seems like this research is moving us closer to having an automated way to get precise environmental impact data tied directly to a product's build. We’ll keep an eye on how these types of systems evolve.

Steffen Castle, Julian Moreno Schneider, Leonhard Hennig Georg Rehm

Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI)

cs.CL

Submitted: 2025-02-11

Updated: 2025-02-11

Code: https://github.com/DFKI-NLP/eco-link

Importance score: 76/100

The gist: Growing concerns about climate change and sustainability are driving manufacturers to take significant steps toward reducing their carbon footprints.

Key concepts

Bill of Materials (BOM)
A BOM is a list detailing all the parts, materials, and components required to manufacture a specific product. This paper uses the BOM entries—which include component names, suppliers, and materials—as the starting point for identifying what needs to be assessed for carbon footprint.
LLM Querying
This involves using a fine-tuned chat-based LLM agent. This agent takes information about a component (like its name and material) and its datasheet, then generates a description of how that component is manufactured. This generated text provides crucial context for accurately finding the correct LCA database entry.
Semantic Similarity Matching
This step ranks potential LCA database entries by comparing their descriptions to an embedding created from the LLM's output. By using vector stores and cosine similarity, the system finds the most relevant LCA process description, ensuring a better match for carbon emission data.

Terminology

Summary

Growing concerns about climate change and sustainability are driving manufacturers to take significant steps toward reducing their carbon footprints.

The gist

LLMs can be leveraged to automate the mapping of product components from manufacturer Bills of Materials (BOMs) directly to Life Cycle Assessment (LCA) database entries by using LLMs to expand on available component information.

How it works

The proposed system employs a multi-step approach consisting of three connected modules: Document retrieval, LLM querying, and database ranking. The overall pipeline is illustrated in Figure 3.

  1. The first step involves Datasheet Selection, where textual similarity between the concatenation of the filename and text of the datasheet versus the concatenation of the component name, manufacturer, and material from a BOM entry is measured using cosine similarity from a text embedding model. A threshold value of 0.5 or higher is used to indicate a match before including datasheet text in further processing.

  2. The second step utilizes LLM Querying with an LLM agent fine-tuned for chat. This agent receives a prompt that includes all relevant context—the BOM entry information (component name, supplier, and material) along with the content of the datasheet, if available—and is instructed to produce a description of the manufacturing process used to create the component. The output of this model is then used in further processing.

  3. The third step involves Semantic Similarity Matching for database ranking. Embeddings are created for each LCA database entry using its process name and description, and these embeddings are stored in a FAISS vector store. An embedding created from the LLM response (from the previous step) is compared to this vector store using cosine similarity to obtain a ranking of database entries.

Key Contributions

The key contributions of this work are:

  1. Utilize fine-grained information from BOMs to provide a more accurate assessment of carbon emissions.

  2. Introduce LLMs into the entity mapping process in order to provide additional context.

  3. Integrate additional context from component datasheets in order to further improve context.

Evaluation and Results

The system was evaluated on a small set of labeled evaluation data consisting of 21 components from three different BOMs. The performance metrics used are Hits@n, defined as the proportion of instances the correct item is present in the top n recommendations. The results show that the proposed approach is on par or slightly better than non-expert human performance, achieving a Human (non-expert) score of 0.48 for Hits@5, compared to 0.19 for Semantic similarity only and 0.43 for LLM alone. Furthermore, the LLM + Datasheet approach achieved a score of 0.48 for Hits@5 and 0.24 for Hits@19, indicating that including datasheet context considerably improve[s] reliability of responses.

Conclusion and Future Work

The paper concludes that the proposed pipeline is acceptable given the challenging nature of the task — on par or slightly better than non-expert human performance, suggesting it could potentially replace the non-expert human in this mapping process. Future work includes an expanded evaluation with a larger evaluation dataset and integration of additional context from sources such as web search results.

Acknowledgments

This paper has received funding from the European Union under grant agreement no. 101058573 (SciLake) and by the German Bundesministerium für Umwelt, Naturschutz, nukleare Sicherheit und Verbraucherschutz (BMUV) under the Green AI Hub Mittelstand initiative.

References

(The paper lists several references, including those for ecoinvent, Flamingo, CaML, and Llama 3.1.)

--- Page 1 arXiv:2502.07418v1 [cs.CL] 11 Feb 2025 first.last@dfki.de Steffen Castle Julian Moreno Schneider Leonhard Hennig Georg Rehm Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI) Berlin, Germany

--- Page 2 arXiv:2502.07418v1 [cs.CL] 11 Feb 2025 first.last@dfki.de Steffen Castle Julian Moreno Schneider Leonhard Hennig Georg Rehm Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI) Berlin, Germany

--- Page 3 arXiv:2502.07418v1 [cs.CL] 11 Feb 2025 first.last@dfki.de Steffen Castle Julian Moreno Schneider Leonhard Hennig Georg Rehm Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI) Berlin, Germany

--- Page 4 arXiv:2502.07418v1 [cs.

Improvements for AI systems

Here are specific improvements for an AI system based on this scientific paper, along with what these improved systems can achieve:


  1. The proposed pipeline should be upgraded from a single LLM query to a multi-stage, iterative refinement process incorporating explicit confidence scoring at each step (Datasheet Selection and Semantic Similarity Ranking).

  2. The LLM agent should be fine-tuned specifically on the output format of LCA database entries (e.g., process names, inputs, outputs) to ensure its generated manufacturing process descriptions are directly usable for high-precision semantic matching, moving beyond general text generation.

  3. Implement a Contextual Grounding module where the LLM response is augmented with extracted numerical or technical keywords from the selected datasheet (e.g., specific material compositions or chemical processes mentioned in the datasheet) before querying the vector store, leading to more robust semantic similarity matching that accounts for fine-grained technical details.

  4. The system architecture should integrate a Knowledge Graph Augmentation layer where successfully mapped components and their corresponding LCA entries are used to incrementally build a domain-specific knowledge graph, allowing subsequent queries to leverage learned relationships (e.g., inferring the carbon impact of a sub-component based on its parent component's known process).

This improved AI system can achieve the following:

  1. Accurately map components from complex Bills of Materials (BOMs) to specific, granular entries in LCA databases with significantly reduced reliance on manual expert intervention.

  2. Provide a more reliable and contextually rich carbon footprint estimation by leveraging both the textual data from BOMs and the technical specifications found in manufacturer datasheets.

  3. Achieve performance levels comparable to or exceeding non-expert human appraisers (Hits@5 metric), making it a viable candidate for automating preliminary sustainability assessments in manufacturing supply chains.

  4. Enable sophisticated downstream analysis by building a structured knowledge base of material-to-process relationships, allowing the system to perform complex carbon impact calculations and scenario modeling across entire product lines, not just individual components.

Sources

Related papers