LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis

arXiv:2406.05375 · cs.AI, cs.LG · Submitted 2024-06-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis".

Jane: The paper was written by Authors not found in the provided text. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: Okay, so we've established that "LEMMA-RCA" is aiming for a massive, comprehensive view of system failures across many different types of data and industries. Jane, when we look at the summary section, what are the core components they say this dataset actually contains?

Jane: The summary really hammers home that they aren't just throwing random data at us. They’ve structured it to capture specific types of faults—both transient hiccups and persistent problems. It provides a framework for understanding failure diversity.

Meng: And I was paying close attention to the details on the fault types themselves, like out-of-memory errors or DDoS attacks mentioned in relation to the IT domain datasets. Knowing that they’ve mapped these specific, real-world faults is crucial; it moves this beyond theoretical examples.

Lu: The inclusion of both IT and OT domains, as they point out with SWaT and WADI, is a massive signal. It shows an intention to bridge the gap between purely digital systems and physical cyber-physical infrastructure. That's where the most complex failures happen in reality.

Lalam: The explicit mention of covering ten distinct fault types across these domains isn't just a number; it represents robustness in the model design. It tells us that the training data forces any AI system to learn a very rich, varied understanding of what 'normal' and 'failed' look like across contexts.

Tom: So, they’re saying this dataset acts as a foundational source of truth for simulating complex real-world operational failures. Meng, does the structured nature of including these specific faults make it practical for immediate use by an engineer?

Meng: It makes it highly valuable because we don't have to spend months generating failure cases just to test a new model. They’ve done that heavy lifting of curating and validating the fault scenarios using industry-standard tools like Prometheus and CloudWatch.

Jane: Exactly, Meng. It gives us benchmarks—and not just academic ones—that mirror what we actually deal with when things go wrong in production environments every day. It validates the entire evaluation platform aspect of the paper.

Lu: I think the real power here is that they are providing a comparative analysis against other existing benchmarks, like Petshop. This isn't just another dataset; it's designed to *improve* the state-of-the-art by showing where its performance leads compared to established baselines.

Lalam: What this summary accomplishes, in my view, is de-risking the process of building advanced RCA tools. By standardizing such a rich and diverse dataset, they allow future researchers to focus purely on algorithmic improvement rather than data preparation headaches. That's huge for accelerating AI adoption in critical infrastructure.

Tom: Wow, so we’ve seen the scope, and now we know it’s structured for maximum realism. But if the dataset is so comprehensive already, Jane, what improvements or advanced applications does the paper suggest building on top of this foundation?

Improvements: Tom: We've talked about how massive and detailed "LEMMA-RCA" is, covering so many types of faults across IT and OT. Jane, moving into the suggestions for improvement, what advanced research directions does the paper really push us toward?

Jane: The focus seems to shift from *just having* the data to *how we should use* it. It suggests that we need to move beyond simple fault classification and start modeling the temporal evolution of failures—how one small issue leads inevitably to a bigger one.

Meng: From an engineering standpoint, the paper highlights the need for continuous validation and quality assurance in how new data is collected. They emphasize that every fault scenario needs to be meticulously validated against real-world conditions, which speaks to maintaining dataset integrity over time.

Lu: I was really intrigued by the potential for using this framework to build predictive models rather than just reactive ones. If we can model the sequence of events leading up to a known failure type, we could potentially predict the failure before any alarm is even triggered—that’s true preventative maintenance intelligence.

Lalam: Lu touched on prediction, and I think that's where the culture shift happens. We move from being able to *diagnose* (telling us what broke) to proactively *preventing* (telling us what will break). This shifts the focus of AI in operations from forensic accounting to anticipatory care.

Tom: So, it’s about predictive intelligence. Meng

Paper discussion segment 3: Tom: So, if we're talking about what LEMMA-RCA really brings to the table beyond just having a big dataset, it’s how it forces us to connect these different types of data streams into one cohesive picture for root cause analysis.

Jane: Exactly! It’s not enough just to have metrics and logs separately; the power comes from linking a specific unusual trace in an application call directly to a corresponding spike in CPU usage reported by Prometheus.

Lu: That linkage across modalities is revolutionary because it mimics how actual human experts troubleshoot things—you never look at just one dashboard widget when something breaks, you build a story from all the available data points.

Meng: But Lu, if we’re talking about linking traces to metrics and logs across different domains like cloud services *and* physical industrial control systems, what's the overhead on the ingestion side? Can this pipeline actually run in a real-time incident response scenario without collapsing under its own weight?

Lalam: That concern Meng raises about infrastructure is huge, but thinking about the impact, this level of data synthesis means that future AI systems won't just point to a variable; they'll write a narrative explaining *why* the entire system behaved badly.

Tom: Right, Lalam touched on the narrative part; it shifts RCA from being a list of correlated symptoms to being an actual story of failure, which is what we desperately need in complex microservice environments.

Jane: And because they covered both IT and OT systems, it shows that the core principles of diagnosing failure are universal, whether you’re dealing with a failing cloud API or a malfunctioning industrial pump.

Lu: I think the biggest implication is that this moves the field past simple correlation detection and into genuine causal inference across heterogeneous systems, which opens up entire new areas for predictive maintenance modeling.

Meng: Causal inference is one thing, but for me, the practical leap would be if this framework allowed us to train models that could *proactively* adjust system parameters before a documented fault pattern even started appearing.

Lalam: If we can reliably build these cross-domain causal models, the cultural shift will be incredible; it means moving from reactive firefighting to preemptive system self-healing, fundamentally increasing trust and uptime in critical infrastructure everywhere.

Tom: So, if we nail this combination of multi-modality and multi-domain scope, it suggests that future monitoring tools need to act less like dashboards and more like integrated diagnostic detectives.

Conclusion: Tom: So, summing up everything we’ve heard today, it really feels like this paper changes the entire landscape of how we diagnose system failures.

Jane: Absolutely, Tom. It's not just another dataset; it's a comprehensive framework that finally connects the dots between those diverse types of operational data—metrics, logs, traces—in a way that actually helps engineers find the *why* behind an outage.

Tom: Exactly! The ability to handle multi-modal, multi-domain failures means we’re moving past simple anomaly detection and straight into true causal reasoning.

Lu: And what I find so thrilling is that this dataset structure allows us to train models that don't just predict failure, but actually reason about the underlying physical or logical interaction that caused the failure in the first place.

Meng: But Lu, even with all those sophisticated models, somebody has to build the infrastructure around it. My question is, does this level of labeled data make it possible to automate much of the triage process for L1 support teams?

Jane: That’s a really practical point, Meng. Because the root causes are so granular—down to specific code regions—it gives operators actionable steps instead of just pointing them toward a general alert.

Lu: It shifts the focus from "something is broken" to "this specific interaction between service A and resource B broke." That’s a paradigm shift in observability, I tell you.

Lalam: Speaking of paradigm shifts, I think the most profound implication here isn't just for the engineers looking at dashboards; it's for how we build institutional knowledge within companies.

Meng: So you’re saying this could be used to teach junior staff how to think like senior SREs, right?

Lalam: Precisely. By having such a richly labeled dataset of failures, organizations can actually codify their collective troubleshooting wisdom and improve the overall operational culture by making knowledge explicit.

Tom: Wow, Lu, Meng, Lalam—that really puts into perspective that this isn't just a research paper; it's an operational playbook for the future.

Jane: It’s incredible how much effort went into creating a dataset that is so representative of real-world fault scenarios across both IT and OT domains.

Tom: We gotta wrap up, but I think we can say that "LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis" is going to be a massive resource for the entire industry.

Lu: It's genuinely paving the way for next-generation AI systems that can truly understand complex physical and digital interactions.

Meng: For us engineers, this means we can finally build tools that are reliable enough to trust with critical infrastructure.

Lalam: And for humanity, it means more resilient systems and a deeper understanding of technological failure modes across the board.

Jane: Thanks so much to everyone for joining us on the radio today; we'll catch you next time when we break down another fascinating paper!

Authors not found in the provided text.

cs.AI, cs.LG

Submitted: 2024-06-08

Updated: 2026-08-25

Code: https://github.com/amazon-science/petshop-root-cause-analysis

Project page: https://lemma-rca.github.io

Importance score: 91/100

The gist: This paper introduces LEMMA-RCA, a "large-scale, open-source dataset" designed to address the scarcity of diverse, high-quality data for Root Cause Analysis (RCA).

Key concepts

Root Cause Analysis (RCA)
The process of identifying the fundamental reason(s) why a system failure or incident occurred. The dataset supports this by providing detailed fault scenarios across various domains, helping pinpoint the 'why' behind an outage.
Multi-modal/Multi-domain
Refers to the dataset's scope, covering diverse data types (metrics, logs, traces) and operational environments (IT and OT). This allows AI to link failures across both purely digital systems and physical infrastructure.
Causal Inference
A key advancement discussed, moving beyond simple correlation. It involves training AI models to understand the actual cause-and-effect relationship between different system events, leading to deeper understanding of failure mechanisms.
Predictive Maintenance
Using advanced AI models to forecast potential failures before they happen. Instead of reacting after an alarm triggers, this approach aims to predict issues by modeling the sequence of events that lead up to a known fault.

Terminology

Summary

This paper introduces LEMMA-RCA, a large-scale, open-source dataset designed to address the scarcity of diverse, high-quality data for Root Cause Analysis (RCA). By providing a multi-modal and multi-domain benchmark, the authors aim to bridge this gap and facilitate the development of more robust RCA techniques for complex, real-world systems.

The limitations of existing datasets

Progress in the field of RCA is currently hindered by the lack of large-scale, open-source datasets tailored for RCA. The authors note that existing public datasets often fail to meet the needs of modern research due to several specific deficiencies:

  • NeZha is confined to one domain and contains significant missing monitoring data.

  • PetShop is limited by a small size and low system complexity.

  • Sock-Shop utilizes synthetic injected faults rather than real-world failures.

  • Datasets like ITOps and Murphy are not public, preventing fair comparisons between methods.

Dataset composition and modalities

LEMMA-RCA encompasses real-world applications such as IT operations and water treatment systems, involving hundreds of system entities. The dataset is organized into two primary domains:

  • IT Domain: Includes the Product Review Platform and Cloud Computing Platform, which feature microservice architectures with various system pods and nodes.

  • OT Domain: Includes the Secure Water Treatment (SWaT) and Water Distribution (WADI) sub-datasets, which capture the monitoring status of sensors and actuators.

The dataset is multi-modal, providing textual system logs with millions of event records and time series metric data with more than 100,000 timestamps.

Data preprocessing and feature extraction

To handle unstructured log data, the authors transform it into a timeseries format using log parsing tools. They extract three distinct feature types to enhance analysis:

  1. X 1L: The occurrence frequency of each log template.

  2. X 2L: The frequency of abnormal logs associated with system failures, such as those containing keywords like 'error' or 'critical.'

  3. X 3L: A feature derived from a TF-IDF based method combined with Principal Component Analysis (PCA).

For the OT domain, the authors convert discrete status labels into a continuous format by employing anomaly detection algorithms, such as Support Vector Data Description and Isolation Forest, to generate system Key Performance Indicators (KPIs).

Experimental evaluation and findings

The authors evaluate six baseline algorithms, including PC-based, CIRCA, epsilon-Diagnosis, RCD, BARO, and Nezha. Their empirical study reveals several critical insights:

  • Multi-modal input—combining both metric and log data—significantly enhances the performance of RCA methods compared to using a single modality alone.

  • The PC algorithm and epsilon-Diagnosis perform worst, likely because they struggle to capture long-term dependencies in large-scale datasets.

  • In the OT domain, CIRCA consistently achieves the best overall performance, though the authors note that fleeting events in the SWaT and WADI datasets remain a significant challenge for all current methods.

Improvements for AI systems

1. Cross-Modal Temporal Fusion Transformer (CMTFT)

  • Improvement: Replace simple feature concatenation of logs (X L) and metrics with a transformer-based architecture that utilizes cross-attention mechanisms to learn the explicit alignment between log template frequency spikes and metric volatility.

  • Capability: The system can detect silent or subtle failures—such as the Cryptojacking or Silent Pod Degradation scenarios described in the paper—where individual metrics might stay within nominal bounds but exhibit abnormal temporal correlations with specific log event patterns.

2. Multi-Domain Foundation Model for System Observability (MFM-SO)

  • Improvement: Implement a self-supervised pre-training regime using the massive, heterogeneous data from both IT (Microservices) and OT (Water Treatment/Distribution) domains provided in LEMMA-RCA to learn universal representations of systemic anomaly.

  • Capability: The system can perform zero-shot or few-shot root cause localization in entirely new industrial environments (e.g., a smart manufacturing plant or a power grid) by transferring learned causal patterns from the IT/OT datasets to the new domain.

3. Neuro-Symbolic LLM Diagnostic Agent

  • Improvement: Integrate the structured Golden Signal log features (X 2L) and PCA-reduced TF-IDF components (X 3L) into a structured prompting framework for Large Language Models, constrained by the system's physical/logical dependency graph.

  • Capability: Instead of merely outputting a ranked list of entities (Top- K), the system can provide human-readable, causal explanations (e.g., The Database pod is the root cause because an external storage saturation event triggered a spike in 'connection timeout' logs, which subsequently cascaded to high latency in the Rating service).

4. Topology-Aware Spatio-Temporal Graph Neural Network (TST-GNN)

  • Improvement: Incorporate the system entity interdependencies (e.g., Pod to Node to Cluster) and the temporal granularity of 1-second metric intervals into a dynamic Graph Neural Network that treats faults as evolving perturbations on a graph.

  • Capability: The system can effectively distinguish between source entities and symptom entities in complex cascading failures, preventing the common error where downstream services are incorrectly flagged as root causes due to their high error rates.

5. Proactive Online Multi-Modal Anomaly Predictor

  • Improvement: Shift from retrospective RCA to predictive diagnosis by training a model on the continuous time-series KPI data and log streams to identify pre-fault trajectories.

  • Capability: The system can issue early warnings before a system failure occurs (e.g., detecting the gradual resource consumption of a Malware Attack or Cryptojacking script) by identifying the infinitesimal shifts in the joint distribution of metrics and logs that precede a KPI breach.

Sources

Related papers