LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis".
Jane: The paper was written by Authors not found in the provided text. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: Okay, so we've established that "LEMMA-RCA" is aiming for a massive, comprehensive view of system failures across many different types of data and industries. Jane, when we look at the summary section, what are the core components they say this dataset actually contains?
Jane: The summary really hammers home that they aren't just throwing random data at us. They’ve structured it to capture specific types of faults—both transient hiccups and persistent problems. It provides a framework for understanding failure diversity.
Meng: And I was paying close attention to the details on the fault types themselves, like out-of-memory errors or DDoS attacks mentioned in relation to the IT domain datasets. Knowing that they’ve mapped these specific, real-world faults is crucial; it moves this beyond theoretical examples.
Lu: The inclusion of both IT and OT domains, as they point out with SWaT and WADI, is a massive signal. It shows an intention to bridge the gap between purely digital systems and physical cyber-physical infrastructure. That's where the most complex failures happen in reality.
Lalam: The explicit mention of covering ten distinct fault types across these domains isn't just a number; it represents robustness in the model design. It tells us that the training data forces any AI system to learn a very rich, varied understanding of what 'normal' and 'failed' look like across contexts.
Tom: So, they’re saying this dataset acts as a foundational source of truth for simulating complex real-world operational failures. Meng, does the structured nature of including these specific faults make it practical for immediate use by an engineer?
Meng: It makes it highly valuable because we don't have to spend months generating failure cases just to test a new model. They’ve done that heavy lifting of curating and validating the fault scenarios using industry-standard tools like Prometheus and CloudWatch.
Jane: Exactly, Meng. It gives us benchmarks—and not just academic ones—that mirror what we actually deal with when things go wrong in production environments every day. It validates the entire evaluation platform aspect of the paper.
Lu: I think the real power here is that they are providing a comparative analysis against other existing benchmarks, like Petshop. This isn't just another dataset; it's designed to *improve* the state-of-the-art by showing where its performance leads compared to established baselines.
Lalam: What this summary accomplishes, in my view, is de-risking the process of building advanced RCA tools. By standardizing such a rich and diverse dataset, they allow future researchers to focus purely on algorithmic improvement rather than data preparation headaches. That's huge for accelerating AI adoption in critical infrastructure.
Tom: Wow, so we’ve seen the scope, and now we know it’s structured for maximum realism. But if the dataset is so comprehensive already, Jane, what improvements or advanced applications does the paper suggest building on top of this foundation?
Improvements: Tom: We've talked about how massive and detailed "LEMMA-RCA" is, covering so many types of faults across IT and OT. Jane, moving into the suggestions for improvement, what advanced research directions does the paper really push us toward?
Jane: The focus seems to shift from *just having* the data to *how we should use* it. It suggests that we need to move beyond simple fault classification and start modeling the temporal evolution of failures—how one small issue leads inevitably to a bigger one.
Meng: From an engineering standpoint, the paper highlights the need for continuous validation and quality assurance in how new data is collected. They emphasize that every fault scenario needs to be meticulously validated against real-world conditions, which speaks to maintaining dataset integrity over time.
Lu: I was really intrigued by the potential for using this framework to build predictive models rather than just reactive ones. If we can model the sequence of events leading up to a known failure type, we could potentially predict the failure before any alarm is even triggered—that’s true preventative maintenance intelligence.
Lalam: Lu touched on prediction, and I think that's where the culture shift happens. We move from being able to *diagnose* (telling us what broke) to proactively *preventing* (telling us what will break). This shifts the focus of AI in operations from forensic accounting to anticipatory care.
Tom: So, it’s about predictive intelligence. Meng
Paper discussion segment 3: Tom: So, if we're talking about what LEMMA-RCA really brings to the table beyond just having a big dataset, it’s how it forces us to connect these different types of data streams into one cohesive picture for root cause analysis.
Jane: Exactly! It’s not enough just to have metrics and logs separately; the power comes from linking a specific unusual trace in an application call directly to a corresponding spike in CPU usage reported by Prometheus.
Lu: That linkage across modalities is revolutionary because it mimics how actual human experts troubleshoot things—you never look at just one dashboard widget when something breaks, you build a story from all the available data points.
Meng: But Lu, if we’re talking about linking traces to metrics and logs across different domains like cloud services *and* physical industrial control systems, what's the overhead on the ingestion side? Can this pipeline actually run in a real-time incident response scenario without collapsing under its own weight?
Lalam: That concern Meng raises about infrastructure is huge, but thinking about the impact, this level of data synthesis means that future AI systems won't just point to a variable; they'll write a narrative explaining *why* the entire system behaved badly.
Tom: Right, Lalam touched on the narrative part; it shifts RCA from being a list of correlated symptoms to being an actual story of failure, which is what we desperately need in complex microservice environments.
Jane: And because they covered both IT and OT systems, it shows that the core principles of diagnosing failure are universal, whether you’re dealing with a failing cloud API or a malfunctioning industrial pump.
Lu: I think the biggest implication is that this moves the field past simple correlation detection and into genuine causal inference across heterogeneous systems, which opens up entire new areas for predictive maintenance modeling.
Meng: Causal inference is one thing, but for me, the practical leap would be if this framework allowed us to train models that could *proactively* adjust system parameters before a documented fault pattern even started appearing.
Lalam: If we can reliably build these cross-domain causal models, the cultural shift will be incredible; it means moving from reactive firefighting to preemptive system self-healing, fundamentally increasing trust and uptime in critical infrastructure everywhere.
Tom: So, if we nail this combination of multi-modality and multi-domain scope, it suggests that future monitoring tools need to act less like dashboards and more like integrated diagnostic detectives.
Conclusion: Tom: So, summing up everything we’ve heard today, it really feels like this paper changes the entire landscape of how we diagnose system failures.
Jane: Absolutely, Tom. It's not just another dataset; it's a comprehensive framework that finally connects the dots between those diverse types of operational data—metrics, logs, traces—in a way that actually helps engineers find the *why* behind an outage.
Tom: Exactly! The ability to handle multi-modal, multi-domain failures means we’re moving past simple anomaly detection and straight into true causal reasoning.
Lu: And what I find so thrilling is that this dataset structure allows us to train models that don't just predict failure, but actually reason about the underlying physical or logical interaction that caused the failure in the first place.
Meng: But Lu, even with all those sophisticated models, somebody has to build the infrastructure around it. My question is, does this level of labeled data make it possible to automate much of the triage process for L1 support teams?
Jane: That’s a really practical point, Meng. Because the root causes are so granular—down to specific code regions—it gives operators actionable steps instead of just pointing them toward a general alert.
Lu: It shifts the focus from "something is broken" to "this specific interaction between service A and resource B broke." That’s a paradigm shift in observability, I tell you.
Lalam: Speaking of paradigm shifts, I think the most profound implication here isn't just for the engineers looking at dashboards; it's for how we build institutional knowledge within companies.
Meng: So you’re saying this could be used to teach junior staff how to think like senior SREs, right?
Lalam: Precisely. By having such a richly labeled dataset of failures, organizations can actually codify their collective troubleshooting wisdom and improve the overall operational culture by making knowledge explicit.
Tom: Wow, Lu, Meng, Lalam—that really puts into perspective that this isn't just a research paper; it's an operational playbook for the future.
Jane: It’s incredible how much effort went into creating a dataset that is so representative of real-world fault scenarios across both IT and OT domains.
Tom: We gotta wrap up, but I think we can say that "LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis" is going to be a massive resource for the entire industry.
Lu: It's genuinely paving the way for next-generation AI systems that can truly understand complex physical and digital interactions.
Meng: For us engineers, this means we can finally build tools that are reliable enough to trust with critical infrastructure.
Lalam: And for humanity, it means more resilient systems and a deeper understanding of technological failure modes across the board.
Jane: Thanks so much to everyone for joining us on the radio today; we'll catch you next time when we break down another fascinating paper!
Authors not found in the provided text.
cs.AI, cs.LG
Submitted: 2024-06-08
Updated: 2026-08-25
Code: https://github.com/amazon-science/petshop-root-cause-analysis
Project page: https://lemma-rca.github.io
Importance score: 91/100
The gist: This paper introduces LEMMA-RCA, a "large-scale, open-source dataset" designed to address the scarcity of diverse, high-quality data for Root Cause Analysis (RCA).
Key concepts
- Root Cause Analysis (RCA)
- The process of identifying the fundamental reason(s) why a system failure or incident occurred. The dataset supports this by providing detailed fault scenarios across various domains, helping pinpoint the 'why' behind an outage.
- Multi-modal/Multi-domain
- Refers to the dataset's scope, covering diverse data types (metrics, logs, traces) and operational environments (IT and OT). This allows AI to link failures across both purely digital systems and physical infrastructure.
- Causal Inference
- A key advancement discussed, moving beyond simple correlation. It involves training AI models to understand the actual cause-and-effect relationship between different system events, leading to deeper understanding of failure mechanisms.
- Predictive Maintenance
- Using advanced AI models to forecast potential failures before they happen. Instead of reacting after an alarm triggers, this approach aims to predict issues by modeling the sequence of events that lead up to a known fault.
Terminology
Summary
This paper introduces LEMMA-RCA, a large-scale, open-source dataset
designed to address the scarcity of diverse, high-quality data for Root Cause Analysis (RCA). By providing a multi-modal and multi-domain benchmark, the authors aim to bridge this gap
and facilitate the development of more robust RCA techniques
for complex, real-world systems.
The limitations of existing datasets
Progress in the field of RCA is currently hindered by the lack of large-scale, open-source datasets tailored for RCA.
The authors note that existing public datasets often fail to meet the needs of modern research due to several specific deficiencies:
-
NeZha is
confined to one domain
and contains significant missing monitoring data. -
PetShop is limited by a
small size
and low system complexity. -
Sock-Shop utilizes
synthetic
injected faults rather than real-world failures. -
Datasets like ITOps and Murphy are
not public,
preventing fair comparisons between methods.
Dataset composition and modalities
LEMMA-RCA encompasses real-world applications such as IT operations and water treatment systems,
involving hundreds of system entities. The dataset is organized into two primary domains:
-
IT Domain: Includes the
Product Review Platform
andCloud Computing Platform,
which feature microservice architectures with various system pods and nodes. -
OT Domain: Includes the
Secure Water Treatment (SWaT)
andWater Distribution (WADI)
sub-datasets, which capture the monitoring status of sensors and actuators.
The dataset is multi-modal, providing textual system logs with millions of event records and time series metric data with more than 100,000 timestamps.
Data preprocessing and feature extraction
To handle unstructured log data, the authors transform it into a timeseries format
using log parsing tools. They extract three distinct feature types to enhance analysis:
-
X 1L: The
occurrence frequency of each log template.
-
X 2L: The frequency of
abnormal logs associated with system failures,
such as those containing keywords like 'error' or 'critical.' -
X 3L: A feature derived from a
TF-IDF based method
combined with Principal Component Analysis (PCA).
For the OT domain, the authors convert discrete status labels into a continuous format
by employing anomaly detection algorithms, such as Support Vector Data Description and Isolation Forest,
to generate system Key Performance Indicators (KPIs).
Experimental evaluation and findings
The authors evaluate six baseline algorithms, including PC-based, CIRCA, epsilon-Diagnosis, RCD, BARO, and Nezha. Their empirical study reveals several critical insights:
-
Multi-modal input—combining both metric and log data—significantly enhances the performance of RCA methods
compared to using a single modality alone. -
The PC algorithm and epsilon-Diagnosis
perform worst,
likely because theystruggle to capture long-term dependencies in large-scale datasets.
-
In the OT domain, CIRCA
consistently achieves the best overall performance,
though the authors note thatfleeting events
in the SWaT and WADI datasets remain a significant challenge for all current methods.
Improvements for AI systems
1. Cross-Modal Temporal Fusion Transformer (CMTFT)
-
Improvement: Replace simple feature concatenation of logs (X L) and metrics with a transformer-based architecture that utilizes cross-attention mechanisms to learn the explicit alignment between log template frequency spikes and metric volatility.
-
Capability: The system can detect
silent
orsubtle
failures—such as the Cryptojacking or Silent Pod Degradation scenarios described in the paper—where individual metrics might stay within nominal bounds but exhibit abnormal temporal correlations with specific log event patterns.
2. Multi-Domain Foundation Model for System Observability (MFM-SO)
-
Improvement: Implement a self-supervised pre-training regime using the massive, heterogeneous data from both IT (Microservices) and OT (Water Treatment/Distribution) domains provided in LEMMA-RCA to learn universal representations of
systemic anomaly.
-
Capability: The system can perform zero-shot or few-shot root cause localization in entirely new industrial environments (e.g., a smart manufacturing plant or a power grid) by transferring learned causal patterns from the IT/OT datasets to the new domain.
3. Neuro-Symbolic LLM Diagnostic Agent
-
Improvement: Integrate the structured
Golden Signal
log features (X 2L) and PCA-reduced TF-IDF components (X 3L) into a structured prompting framework for Large Language Models, constrained by the system's physical/logical dependency graph. -
Capability: Instead of merely outputting a ranked list of entities (Top- K), the system can provide human-readable, causal explanations (e.g.,
The Database pod is the root cause because an external storage saturation event triggered a spike in 'connection timeout' logs, which subsequently cascaded to high latency in the Rating service
).
4. Topology-Aware Spatio-Temporal Graph Neural Network (TST-GNN)
-
Improvement: Incorporate the system entity interdependencies (e.g., Pod to Node to Cluster) and the temporal granularity of 1-second metric intervals into a dynamic Graph Neural Network that treats faults as evolving perturbations on a graph.
-
Capability: The system can effectively distinguish between
source
entities andsymptom
entities in complex cascading failures, preventing the common error where downstream services are incorrectly flagged as root causes due to their high error rates.
5. Proactive Online Multi-Modal Anomaly Predictor
-
Improvement: Shift from retrospective RCA to predictive diagnosis by training a model on the continuous time-series KPI data and log streams to identify
pre-fault
trajectories. -
Capability: The system can issue early warnings before a system failure occurs (e.g., detecting the gradual resource consumption of a Malware Attack or Cryptojacking script) by identifying the infinitesimal shifts in the joint distribution of metrics and logs that precede a KPI breach.
Sources
- Constructing Large-Scale Real-World Benchmark Datasets for AIOps
- Multi-modal Causal Structure Learning and Root Cause Analysis
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection