EulerESG: Automating ESG Disclosure Analysis with LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EulerESG: Automating ESG Disclosure Analysis with LLMs".
Tom: ESG reports are often published as long, heterogeneous PDF documents, making systematic analysis difficult and labor-intensive.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Hey Jane, so let's talk about the paper 'EulerESG: Automating ESG Disclosure Analysis with LLMs' a bit more; what do you think about how they framed the problem in their introduction?
Jane: I think it’s smart how they clearly laid out the challenges, pointing out that existing tools are either using brittle rules or just treating everything as plain text, which isn't sufficient for real compliance checking.
Lu: The authors really pinpointed the core issues: unstructured and heterogeneous formats, inconsistent terminology across companies, and the problem of promotional content muddying up the actual disclosures.
Meng: I wonder how they managed to build a system that could handle all those variations simultaneously without needing a huge manual rulebook for every single report type.
Lalam: From my perspective, it’s exciting because it shows that we don't need perfect parsing; we just need a system that can map messy inputs onto clean, structured outputs reliably.
The paper's summary: Tom: Moving on to what the paper actually proposes as their solution, the core idea seems to be this entire pipeline designed for high-fidelity metric extraction from these complex reports.
Jane: Right, they are building an automated tool that uses LLMs to extract structured ESG metrics in a way that significantly cuts down on the need for manual rule engineering by explicitly modeling those standards.
Lu: They detail a whole framework, starting with Standard Metric Identification and then refining those extracted metrics through semantic expansion to give them the context they need for better retrieval later on.
Meng: That multi-stage process sounds intensive; how do you make sure that the initial identification stage doesn't miss subtle but important metrics embedded in weird formatting?
Lalam: The summary emphasizes a dual-channel retrieval system, which I think is key because combining keyword searching with semantic retrieval should capture information that keyword searches would completely miss.
The paper's improvements: Tom: Now let's look at the specific improvements they detail in 'EulerESG: Automating ESG Disclosure Analysis with LLMs', because it sounds like they didn't just propose an idea, but a whole architecture.
Jane: They suggest several key enhancements, including fine-grained alignment of extracted disclosures against industry-specific requirements across over one hundred industries and multiple reporting frameworks.
Lu: The paper lays out a five-module system that handles everything from encoding the report content to the final LLM-driven analysis, which shows a really comprehensive approach to data processing.
Meng: I’m focusing on the dual-channel metrics retrieval part; combining keyword searches with semantic similarity matching using embeddings from those standard metric definitions is a clever way to improve recall.
Lalam: The LLM-driven analysis module that performs content reasoning and then classifies disclosures into fully discussed, partially discussed, or not discussed categories offers a much more nuanced view than just saying whether a metric exists or not.
Conclusion: Tom: So, to wrap up this discussion on 'EulerESG: Automating ESG Disclosure Analysis with LLMs', it seems the paper’s main implication is that we can move past brittle extraction methods toward a system that understands the structure and context of ESG reporting.
Jane: That really means investors and regulators can get much more consistent data, allowing for better benchmarking across different companies and sectors based on standardized metrics.
Lu: I think the way they integrated semantic expansion with LLM reasoning shows how powerful these models become when given structured knowledge about the standards they are trying to follow.
Meng: Practically speaking, it suggests that we can build tools that don't just pull text but actually perform meaningful comparisons and provide actionable insights for corporate teams.
Lalam: I’m really optimistic because this work lays the groundwork for a future where AI systems can genuinely help bridge the gap between complex regulatory language and actionable business intelligence.
UNSW Sydney
cs.CL, cs.AI, cs.CY
Submitted: 2025-11-18
Updated: 2026-10-03
Code: https://github.com/UNSW-database/EulerESG
Importance score: 83/100
The gist: ESG reports are often published as long, heterogeneous PDF documents, making systematic analysis difficult and labor-intensive.
Key concepts
- ESG-Standard-Specific Metrics Extraction
- This is the process of using LLMs to pull specific ESG data points from unstructured PDF reports. Instead of relying on manual rules, the system learns to identify and extract metrics directly from the text, even when the reports are noisy or formatted differently.
- Multi-Framework Standard Alignment
- The system can compare disclosures across different industries and regulatory frameworks simultaneously. It aligns extracted data with industry-specific requirements, allowing users to flexibly benchmark companies against various standards at once.
- Dual-Channel Metrics Retrieval
- This module uses two methods to find relevant information in a report: keyword searching for exact terms and semantic retrieval using vector embeddings. This combination ensures that both precise terminology and conceptually similar information are captured.
- Interactive LLM ESG Chatbot
- A chatbot interface allows users to ask natural language questions about the reports. It can define complex ESG terms, query policies not explicitly covered by standards, or summarize ambiguous disclosures for easy understanding.
Terminology
Summary
ESG reports are often published as long, heterogeneous PDF documents, making systematic analysis difficult and labor-intensive. EulerESG presents an LLM-powered system that automates ESG disclosure analysis by explicitly modeling underlying reporting standards, enabling high-fidelity metric extraction and interactive exploration of corporate sustainability data.
Contributions
The paper details several key contributions aimed at addressing the challenges of analyzing unstructured ESG reports:
-
ESG-Standard-Specific Metrics Extraction: The system designs an automated pipeline leveraging LLMs to extract structured ESG metrics from
unstructured, noisy, and highly variable corporate reports in PDF format,
significantly reducing the need for manual rule engineering. -
Multi-Framework Standard Alignment: The work covers
over 100 industries and multiple reporting frameworks.
A specific module enablesfine-grained alignment of extracted disclosures with industry-specific requirements, supporting flexible comparisons across companies, industries, and regulatory standards.
-
LLM-Powered Automated ESG Analysis: An
ESG agent that integrates LLM reasoning with prompt engineering
is designed to performcontext-aware entity recognition, metric classification, and compliance flagging.
This agent supports benchmarking and decision support through interactive interfaces.
System Design
The EulerESG architecture is decomposed into five specialized modules to manage the complexity of ESG analysis:
-
Standards-Based Metric Extraction & Expansion: This module preprocesses metrics through a three-stage refinement process: first, the
Standard Metric Identification (SMI) component extracts relevant metrics from an extensive collection of unstructured ESG standards
; second,Standard-Aligned Metrics Refinement
augments each metric with metadata; and finally,Semantic Expansion of Metric Definitions
enriches each metric with contextual information necessary for vector-based retrieval. -
Disclosure-Oriented Report Encoding: To handle heterogeneous formats, this module uses a three-stage approach:
Content Extraction is performed page by page
; followed bySegment-Level Structuring annotates the extracted content at the segment level
; and finally,Report Content Embedding transforms each textual segment into a vector representation using an embedding model,
specifically adopting the open-source BGE-M3 model. -
Dual-Channel Metrics Retrieval: This module combines two retrieval methods to find relevant information from reports:
Keyword retrieval searches directly for occurrences of metric terminology within the report,
and a complementarySemantic Retrieval mechanism, which uses the embeddings generated during the semantic expansion phase to identify relevant segments in the report.
Relevance scores are refined by applying vector similarity matching between extended metric definitions and report content using BGE-M3 embeddings, followed by re-ranking with the BGE-Reranker-v2-M3 model. -
LLM-Driven Disclosure Analysis: This module performs a
twophase reasoning process.
First,LLM-Guided Content Reasoning employs designed analytical prompts to examine retrieved content segments, evaluating both the indicators and the quality of disclosures based on structured criteria.
Second, results undergoStructured MultiClass Disclosure Classification,
categorizing each result into one of three predefined classes:(1) (Fully) Discussed/Disclosed, (2) Partially Discussed/Disclosed, or (3) Not Discussed/Disclosed.
-
Interactive LLM ESG Chatbot: This integrated chatbot allows users to interact with reports through natural language queries, supporting
on-demand definitions of ESG terms and metric codes,
contextual queries about metrics or policies not explicitly covered by standard frameworks,
and thesummarization of policy narratives or ambiguous disclosures.
Performance Analysis
The system was evaluated on a cross-industry case study using four large companies (BMW, Macquarie Group, P&G, and Dell) across twelve company-industry pairs aligned with SASB standards. The evaluation measured two primary dimensions: metric-level accuracy and running time. Across all tested LLM backends (Gemini-2.5-flash, Deepseek-V3.2-Exp, Claude-Haiku-4.5, GPT-5, and Qwen-3), all models achieved high metric accuracy (Table 1), with the highest average accuracy attained by GPT-5 at 0.95 across the tested pairs. The system demonstrated an efficiency-accuracy trade-off,
where while GPT-5 achieved the highest accuracy, models like Gemini-2.5-flash and Qwen-3 offer a balanced middle ground, combining near-top accuracy with substantially lower latency than GPT-5.
Demonstration and Utility
The EulerESG interface is designed as an interactive dashboard featuring a Disclosure Summary Panel,
a Metrics Table
presenting extracted data in a standardized format, and a PDF Viewer for cross-referencing. The system is positioned to serve diverse stakeholders: Regulators and Investors can use it to monitor ESG compliance, benchmark firms across sectors, and build consistent ESG scorecards.
Corporate teams can apply it for pre-publication self-assessment
and identifying "missing or weakly substantiated metrics.
Improvements for AI systems
Here are specific improvements for existing AI systems based on the EulerESG framework, along with what the improved system could achieve:
-
The current limitation is brittle, rule-based extraction or treating ESG reports as generic text without standard awareness.
-
Improvement: Implement a modular pipeline that explicitly models reporting standards (e.g., SASB, GRI) from the outset using a "Standards-Based Metric Extraction & Expansion" module.
-
What the improved system can do: It will move beyond simple keyword matching to perform high-fidelity extraction of structured metrics, even when the report uses non-standard terminology or different units, by mapping extracted data directly against predefined standard taxonomies.
-
Improvement: Integrate a Dual-Channel Metrics Retrieval mechanism that combines direct Keyword Retrieval with advanced Semantic Retrieval (using embeddings from standardized metric definitions).
-
What the improved system can do: It will significantly reduce false negatives and positives during information retrieval by ensuring that semantic similarity between a reported narrative and a formal metric definition is captured, leading to more comprehensive data capture than keyword search alone.
-
Improvement: Introduce an LLM-Driven Disclosure Analysis module that uses structured analytical prompts to classify disclosures into three granular classes: (1) Fully Discussed/Disclosed, (2) Partially Discussed/Disclosed, and (3) Not Discussed/Disclosed, distinguishing between quantitative disclosure and qualitative discussion.
-
What the improved system can do: It will provide a nuanced assessment of disclosure quality that goes beyond simple presence/absence flags, allowing users to immediately gauge the depth of commitment or risk mitigation described in the report.
-
Improvement: Develop an interactive LLM ESG Chatbot integrated with vector databases representing both ESG standards and extracted report content.
-
What the improved system can do: It will enable natural language querying for complex, contextual insights—such as
Compare the governance discussion on board diversity between Company X and Company Y,
orSummarize all climate resilience strategies mentioned in Q3 disclosures
—providing immediate, context-aware benchmarking and explanation without needing to parse raw documents. -
Improvement: Implement a dynamic LLM backend selection strategy based on required trade-offs (Efficiency vs. Accuracy).
-
What the improved system can do: It will allow users to select the optimal model for their task—e.g., selecting a smaller, faster model (like Gemini-2.5-flash) for high-throughput screening, or a larger, more capable model (like GPT-5) for deep investigative analysis—optimizing both runtime and required accuracy simultaneously.
Sources
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- GPT-4 Technical Report
- Advanced Unstructured Data Processing for ESG Reports: A Methodology for Structured Transformation and Enhanced Analysis
- Recent Advances in Generative AI and Large Language Models: Current Status, Challenges, and Perspectives
- A Scalable Framework for Table of Contents Extraction from Complex ESG Annual Reports
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering