Prognostics for Autonomous Deep-Space Habitat Health Management under Multiple Unknown Failure Modes
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Prognostics for Autonomous Deep-Space Habitat Health Management under Multiple Unknown Failure Modes".
Tom: The gist: We propose an unsupervised prognostics framework for Remaining Useful Life (RUL) prediction that jointly identifies latent failure modes and selects informative sensors using unlabeled run-to-failure data.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and who wrote this stuff: Prognostics for Autonomous Deep-Space Habitat Health Management under Multiple Unknown Failure Modes. It sets the stage perfectly for what they’re trying to achieve.
Jane: The authors are Peters, Mohanty, Fang, Robinson, and Gebraeel from places like Georgia Tech and UC Davis. They’ve been working on these kinds of long-term autonomous systems for a while now.
Lu: What's interesting is that the context they set up with the Artemis program and the Gateway habitat gives you a real sense of what's being discussed here—it grounds this research in real future missions.
Meng: It sounds like this isn't just theoretical stuff; it’s directly relevant to how we design and monitor these systems for long-duration space missions.
Tom: Right, and what they focus on is that deep space habitats are incredibly complex engineering systems, or CES, made up of tightly integrated parts that have to keep going without anyone handy.
Jane: And the paper is focused on how this complexity leads to multiple failure modes that we don't fully understand yet.
Lu: They’re looking at how these different subsystems—like life support or power generation—can all break down in ways that create distinct sensor signatures, even though they are all part of the same habitat.
Meng: So the authors are tackling the issue where you have a bunch of sensors, and each sensor might be important for a different kind of failure.
Tom: Exactly. They’re aiming to move past just detecting *a* problem and instead predicting *how long* it will take before the habitat actually fails, using this multi-mode understanding.
The paper's summary: Jane: So what’s the actual breakdown of their method? Basically, they propose a two-stage approach: first, an offline sensor selection step and then an online diagnosis and prognostics step.
Tom: That’s the structure. The offline part is where they use historical data without knowing the failure modes yet to pick out the best sensors and label those modes.
Lu: They use a method called Expectation-Maximization, or EM, which lets them jointly cluster unlabeled failure events while simultaneously picking informative sensors for each identified mode.
Meng: Jointly doing both tasks—clustering and selection—without needing anyone to pre-label the failures is a big deal for autonomous systems where labeling data is nearly impossible.
Tom: Precisely. After that, in the online phase, they use real-time data to figure out which failure mode is active and then predict the remaining useful life based only on those relevant sensors.
Jane: It sounds like they build a system that learns from the past in a way that lets it adapt to the present situation without needing constant human intervention.
The paper's improvements: Tom: Now, let’s look at what they actually improved compared to previous ideas. They suggest using a feature extraction method called Covariate-Adjusted Functional Principal Component Analysis, or CA-FPCA.
Lu: That’s clever because the covariates in that method represent the underlying failure modes themselves, which are unknown when you start out.
Jane: So instead of just looking at raw sensor numbers, they adjust the features based on what they *think* those unknown failure modes might be, even if it's just an initial guess from clustering.
Meng: That’s the key for onboard deployment—they need compact and informative representations that work even when you can't download massive datasets back to Earth.
Tom: And they use those CA-FPC scores in a Mixture of Gaussian Regression model, or MGR, to estimate things like how likely a certain failure mode is happening.
Jane: The optimization step involves fitting an Expectation–Maximization algorithm to minimize this incomplete-data log-likelihood, which helps them find the best sensor weights and the best failure mode assignments simultaneously.
Conclusion: Tom: So to wrap up, these authors have shown a way to use unsupervised learning to handle deep space habitat health management under multiple unknown failure modes by separating selection and prediction into offline and online stages.
Jane: It really shows that you don't need perfect prior knowledge of every single failure mode to get a reliable prognosis for complex systems like the DSHs.
Lu: The ability to select informative sensors based on failure mode awareness without expert labeling is something I think opens up so much possibility for truly autonomous long-term monitoring.
Meng: From an engineering standpoint, having a method that can dynamically diagnose the active fault and then only use the specific sensors needed for that fault in real time makes the system far more efficient to run on limited power.
Lalam: I think what this means culturally is that we are moving toward AI systems that are truly self-aware in their diagnosis, not just reacting to pre-programmed alerts, which really shifts how we trust complex machinery.
Tom: It’s a solid paper about making autonomous health management smarter by letting the data tell us where to look and what sensors matter most. That’s it for this one.
University of Texas at Rio Grande Valley · Georgia Institute of Technology · North Carolina State University · University of California Davis
stat.ML, cs.LG, cs.SY, eess.SY, stat.AP
Submitted: 2024-11-19
Updated: 2026-10-08
Importance score: 89/100
The gist: The gist: We propose an unsupervised prognostics framework for Remaining Useful Life (RUL) prediction that jointly identifies latent failure modes and selects informative sensors using unlabeled
Key concepts
- Unsupervised Prognostics
- This approach uses historical data from a habitat without pre-labeled failure event information. The goal is to automatically discover hidden patterns, such as different ways the system can fail (failure modes), and use this discovery to predict how much longer the system will operate before failure.
- Offline Sensor Selection
- This initial phase uses Expectation-Maximization (EM) algorithms to determine which sensors are most useful for detecting each potential failure mode. It clusters unlabeled data and optimizes sensor subsets so that the selected sensors provide the best information for predicting different types of failures.
- Online Diagnosis and Prognostics
- Once deployed, this phase uses real-time data from the habitat. It first classifies the current active failure mode by comparing signals to known patterns. Then, it uses only the relevant sensors to fit a regression model that predicts the remaining operational life (RUL) based on how those selected sensors are behaving in real-time.
Terminology
Summary
The gist: We propose an unsupervised prognostics framework for Remaining Useful Life (RUL) prediction that jointly identifies latent failure modes and selects informative sensors using unlabeled run-to-failure data.
Problem Context
Deep-space habitats (DSHs) are safety-critical systems that must operate autonomously for long periods, often beyond the reach of ground-based maintenance or expert intervention Critical DSH subsystems, including environmental control and life support, power generation, and thermal control, are monitored by many sensors and can degrade through multiple failure modes. These failure modes are often unknown, and informative sensors may vary across modes, making accurate RUL prediction challenging when historical failure data are unlabeled. The complexity of a DSH is caused not only by the tight integration of its subsystems but also by subsystem heterogeneity.
Framework Overview
The proposed methodology consists of two key stages: an offline sensor selection step and an online diagnosis and prognostics step. This approach is designed to align with two mission phases: during early deployment, the habitat operates under limited supervision to collect historical sensor data for initializing the model; after this setup phase, autonomous health monitoring and prediction begin. The main contributions include developing a feature extraction methodology that fuses multivariate sensor data into compact and informative representations suitable for autonomous onboard deployment, proposing a failure-mode-aware sensor selection approach using an Expectation-Maximization algorithm that jointly clusters unlabeled failure events and selects informative sensors for each failure mode without requiring expert labeling, and presenting an integrated online framework that uses real-time data to (1) diagnose the active failure mode and (2) predict RUL using a mode-specific regression model.
Offline Sensor Selection
The offline phase involves two parts: Feature Extraction and Optimization.
The Feature Extraction step uses a covariate-adjusted functional principal component analysis (CA-FPCA) method to extract informative features from high-dimensional, time-varying signals. In this setting, the covariates represent the underlying failure modes, which are unknown a priori. The process involves first performing a K-means clustering on sensor signals and using the resulting cluster labels as initial covariates. After feature extraction, these features are called CA-FPC scores, which are then used as predictors in a mixture of Gaussian regression (MGR) model.
The Optimization step involves fitting the MGR-ASGL model to estimate parameters such as failure mode probabilities and sensor selection weights. The optimization minimizes the negative incomplete-data log-likelihood (IDLL) using an Expectation–Maximization (EM) algorithm. This step results in optimal failure mode labels for each sample, and an optimal subset of informative sensors for each failure mode.
Online Diagnosis and Prognostics
Once the offline phase is complete, the habitat enters autonomous mode where real-time data is used to diagnose the active failure mode and predict RUL.
The Diagnosis step first applies multivariate functional principal component analysis (MFPCA) to signals from all selected sensors to extract compact features called MFPC scores, which represent the joint behavior of all sensors. The active failure mode is then classified by finding its K-nearest neighbors and assigning the most common failure mode among them.
The Prognostics step uses only the informative sensors for that diagnosed mode to predict RUL. This involves fitting a weighted functional regression model where time-varying MFPC-scores are used as predictors. The prediction is obtained by expanding the coefficient function using eigenfunctions obtained from applying MFPCA on the training set.
Validation and Results
The proposed methodology was validated through two case studies. The first case study used a simulated dataset designed to reflect key telemetry challenges, including high sensor count, variable signal-to-noise ratios, and unlabeled failure modes. The second case study evaluated the method on the NASA C-MAPSS turbofan engine degradation dataset. Results demonstrated that the MGR-ASGL model successfully identifies mode-specific sensor subsets without access to failure mode labels. The framework showed strong performance in clustering failure modes, identifying informative sensors, and predicting RUL with low relative error. Specifically, the methodology improved early-life prediction accuracy and produced interpretable results by selecting a compact set of informative sensors on the NASA C-MAPSS benchmark.
Conclusion
The work demonstrates the effectiveness of an unsupervised prognostic framework for DSHs with high-dimensional sensor data and multiple degradation modes. The proposed method enables RUL-based prognostics in systems with high-dimensional sensor data and unknown failure modes. This framework is particularly relevant for autonomous DSH environments where expert labeling is constrained by communication delays. The research assumes a linear relationship between log(TTF) and the sensor signals and that each system experiences a single failure mode with the number of modes K known a priori. Future work plans to relax these assumptions by incorporating nonlinear models and more flexible failure-mode structures.
Acknowledgments
This effort is supported by NASA under grant number 80NSSC19K1052 as part of the NASA Space Technology Research Institute (STRI) Habitats Optimized for Missions of Exploration (HOME) ‘SmartHab’ Project. The opinions, findings, conclusions, or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Aeronautics and Space Administration.
Log-Likelihood Definitions
The log-likelihood function for the MGR model is given by equation (25), which is called as the incomplete-data log-likelihood
(IDLL) since we do not know the cause of the observed failure in the data. The “complete-data log-likelihood" (CDLL) given by equation (26) assumes knowledge of the underlying causes of the observed failure in the data. Lemma 1 shows that the expectation of the CDLL w.r.t distribution g is a lower bound for the IDLL i.e., Eg [lC (theta;Y, Z)] ≤ l(theta;Y). This inequality holds because ln(.) is a concave function and the second term on the right side of the inequality is the negative entropy of the g distribution which is positive.
References
[70] Städler, N., Bühlmann, P., and van de Geer, S., “L1-Penalization for Mixture Regression Models,” TEST, Vol. 19, No. 2, 2010.
[75] Murphy, K. P., Machine learning: a probabilistic perspective, MIT press, 2012.
[68] Karhunen, K., Über lineare Methoden in der Wahrscheinlichkeitrechnung, Kirjapaino oy. sana, 1947.
[72] Müller, H.-G., and Zhang, Y., “Time-varying functional regression for predicting remaining lifetime distributions from longitudinal trajectories,” Biometrics, Vol. 61, No. 4, 2005.
[73] Fang, X., Zhou, R., and Gebraeel, N., “An adaptive functional regression-based prognostic model for applications with missing data,” Reliability Engineering & System Safety, Vol. 133, 2015.
[74] Saxena, A., Goebel, K., Simon, D.
Improvements for AI systems
-
textbfIdentify Latent Failure Modes via EM Clustering in Offline Phase:
We develop a novel Expectation-Maximization (EM) algorithm that simultaneously labels failure modes and selects the most informative sensors for each mode.
This allows the system tojointly cluster unlabeled failure events and select informative sensors for each failure mode without requiring expert labeling.
-
textbfEmploy Mode-Specific Sensor Selection:
We propose a failure-mode-aware sensor selection approach using an Expectation–Maximization algorithm that jointly clusters unlabeled failure events and selects informative sensors for each failure mode.
This ensures that the online phase, wherea functional regression model uses these scores to predict the RUL,
is powered only by sensors relevant to the diagnosed fault. -
textbfReal-Time Failure Mode Diagnosis: "First, we apply multivariate functional principal component analysis (MFPCA) to signals from all selected sensors to extract a compact set of features, called MFPC scores, that represent the joint behavior of all sensors. Then we classify the active failure mode by finding its K-nearest neighbors and assigning the most common failure mode among them.
This enables
real-time diagnosis" in the online phase before RUL prediction. -
textbfAdaptive RUL Prediction:
Once the active failure mode is diagnosed, we recalculate the MFPC scores using only the informative sensors for that mode. A functional regression model uses these scores to predict the RUL.
This results in amode-specific regression model
that provides a more accurate prognosis than general models. -
textbfRobust Feature Extraction: "To extract informative features for prognostics, we use a covariate-adjusted functional principal component analysis (CA-FPCA) method. CA-FPCA models the variation in each sensor signal while accounting for external covariates [which represent underlying failure modes].
This provides
compact and informative representations suitable for autonomous onboard deployment" by modeling time-varying signals across unknown mode covariates.
Sources
- A review on competing risks methods for survival analysis
- L1-Penalization for Mixture Regression Models
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey