Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction

arXiv:2608.06993 · cs.LG, cs.AI · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Lifetime prediction of new cryocoolers".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper summary: Tom: All right, this one grabbed me immediately — predicting how long a satellite cryocooler keeps working.

Jane: And the twist is the data situation. Only 1,305 unlabeled telemetry sequences, plus 95 labeled ones from destructive lifetime tests.

Tom: Ninety-five samples would starve any large foundation model. So the authors went the opposite direction entirely.

Jane: They built a family of small encoders — CNN1D, LSTM, GRU, Transformer — trained unsupervised to reconstruct the telemetry.

Lu: I love the capacity-control angle. Embeddings range from 2 dimensions to 512, so the model gets matched to the data you actually have.

Meng: And those embeddings feed two very different tasks — binary classification of lifetime class and one-class anomaly detection.

Lalam: The bigger story is that you can get foundation-model-like behavior — reusable, task-agnostic representations — from tens of thousands of parameters.

Tom: They even built a dimension-aware search, da-NAS, to choose the embedding size automatically.

Jane: And the empirical pattern is clean. Small data wants small embeddings.

Lu: With only four to eight training samples, a 2D or 4D embedding still beats random guessing.

Meng: The bigger Transformer shines when labels are plentiful, then collapses fastest when they vanish.

Lalam: That is exactly the trade-off industrial teams face every day.

Tom: The first page maps the whole pipeline in one picture, from raw sensor readings to clean reusable vectors.

Jane: So let's start there — with the preprocessing that makes everything else possible.

Page 1 of the paper: Tom: We've been circling the big idea — small-data representation models. Page one shows how the messy telemetry gets cleaned up.

Jane: The first filter is brutally simple. Anything outside minus 50 to plus 100 degrees Celsius in housing temperature gets flagged.

Lu: That catches sensor malfunctions and corrupted readings before they poison the model.

Meng: Then entire sequences with missing values get dropped. You cannot train on holes in the data.

Tom: And instead of standard scaling, they reach for a Robust Scaler.

Jane: Good call, because the telemetry is full of outliers and non-Gaussian distributions.

Lu: Robust scaling shrugs off extreme values, which keeps training stable.

Meng: The picture labels that the heavy lifting comes next — unsupervised sequence-to-sequence training on unlabeled data.

Tom: Reconstruction loss pushes the encoder to keep temporal and structural patterns in a compact latent space.

Jane: And a NAS step optimizes both architecture and embedding size.

Lu: Out the other end come fixed-size, task-agnostic vectors. Ready for any downstream job.

Meng: The bottom row shows what those jobs are — binary lifetime classification and one-class anomaly detection.

Tom: Class imbalance handling is baked into that evaluation from the start.

Jane: Which raises a natural question — once the data is clean and embedded, what can you actually predict? The contributions section tells us.

Page 2 of the paper: Tom: We've seen the pipeline on page one. The next stretch of the paper lays out what's genuinely new.

Jane: Six contributions. The first is the paradigm itself — FSD-RM, a small-data representation model designed for satellite telemetry.

Lu: The central claim — generalization can emerge far below the scale that pretrained giants demand.

Meng: They're not copying GPT-style scale. They approximate its functional properties instead.

Tom: Task-agnostic representations, cross-task generalization, but under tight resource constraints.

Jane: Contribution two is the family itself — the four encoders, parameterized by embedding capacity.

Lu: Capacity scaling is the trick. You can dial complexity up or down depending on your data.

Meng: Contribution three — a single pretrained representation supports both binary fault classification and one-class anomaly detection, with no encoder retraining.

Tom: Same vectors, two different jobs. The task-agnostic promise holds up.

Jane: Then there's multi-regime robustness. They test under shrinking downstream training subsets, all the way down to ultra-low-sample scenarios.

Lu: That mirrors real manufacturing, where degradation-stage examples are nearly absent.

Meng: Contribution five is da-NAS — a lightweight search over embedding dimension with progressive scheduling and early stopping.

Tom: Different from classic NAS, which chases depth and width.

Jane: And the final contribution is system-level — the whole pipeline working as one, not a single optimized model.

Lu: Representation learning, capacity scaling, cross-task reuse, and NAS — unified for small-data industrial settings.

Meng: That framing matters when you are actually deploying this in a factory.

Tom: But before deployment, you need to understand the data. The problem specification section shows just how messy it really is.

Page 3 of the paper: Tom: The contributions promise a lot. Now the paper gets concrete about the data and why it's hard.

Jane: The numbers again — 1,305 unlabeled sequences for representation learning, only 95 with ground-truth lifetime labels.

Lu: That is the small-N regime, and it is brutal.

Meng: The downstream task is binary classification around a lifetime threshold. At or below the threshold is standard; above is long-lifetime.

Tom: The paper evaluates three thresholds — 10,000 hours, 15,000 hours, and 20,000 hours.

Jane: And the class balance shifts with each one. At 10k it's 63 standard versus 32 long. At 20k it's 84 versus 11.

Lu: So the imbalance goes from mild 2:1 to severe 7:1. Realistic manufacturing statistics.

Meng: The rare class is the one you actually care about. That's the painful part.

Tom: They call it positive-unlabeled, because most of the telemetry carries no label at all.

Jane: Which is exactly why unsupervised representation learning makes sense. You cannot supervise with 95 samples.

Lu: The other obstacles — heterogeneous sensor modalities, variable sequence lengths, domain shift across test regimes.

Meng: Different test conditions shift the distributions, and measurement noise makes everything worse.

Lalam: From a program perspective, this is the typical aerospace reality — low production volumes, niche signals, and no public benchmark to lean on.

Tom: So the paper is blunt: standard supervised learning and large-scale deep learning both struggle here.

Jane: That sets up the framework's design — but where does the data actually come from? The acquisition section explains it.

Page 4 of the paper: Tom: The problem section lays out the challenge. The acquisition section shows how the telemetry is actually gathered.

Jane: Post-production testing runs through several phases. Run-in lasts 150 hours, with data logged every minute.

Lu: Then a 15-minute noise test at 1-second resolution, measuring vibration frequencies.

Meng: Next comes ESS — environmental stress screening — at room temperature, at minus 40, at plus 71, then a post-ESS room temperature pass.

Tom: If the device survives, it goes through an acceptance test procedure, both before and after the life test.

Jane: The life test runs continuously, with ESS repeated every 500 hours to keep verifying reliability.

Lu: That is a serious amount of hardware in the loop.

Meng: The ESS tests track 10 core telemetry features — temperatures, bus voltage, motor current, RPM, heater power.

Tom: The noise test contributes 33 frequency-domain features — spectral power from 20 hertz to 20 kilohertz, plus RPM and total band power.

Jane: And here's a key design choice — they never fuse time-domain and frequency-domain data into one input.

Lu: Separate preprocessing pipelines per modality. That prevents cross-modal interference.

Meng: Each encoder trains on its own modality, so embeddings stay consistent within each test regime.

Tom: The filtering rules are strict too — bad temperatures out, NaN sequences out, Robust Scaler applied.

Jane: Clean data, per-modality pipelines, and then the embedding machinery takes over.

Lu: Which is precisely where da-NAS enters the story.

Page 5 of the paper: Tom: We've seen how the data is collected and cleaned. Now the search machinery that picks the embedding size.

Jane: Four components — dimension controller, trial optimizer, cross-dimensional stop policy, and scoreboard registry.

Lu: The dimension controller walks through embedding sizes in order, from 2 to 512.

Meng: The trial optimizer uses Optuna to sample hyperparameters and measure validation loss per trial.

Tom: The scoreboard stores the best configuration for each dimension and publishes it as a target.

Jane: Then the stop policy decides when to quit.

Lu: There's a two-regime strategy. Low dimensions — 2, 4, 8, 16 — get full exploration, up to 4,000 trials each.

Meng: No early stopping down there. They want a complete map of the low-dimensional landscape.

Tom: From 32 upward, the Beat-Lower-Dimension rule activates.

Jane: Once the validation loss at the current dimension matches or beats the previous dimension's best, that triggers a BLD event.

Tom: Then a short patience window — five trials — before termination.

Lu: There's also a relative improvement threshold. A trial has to beat the current best by 10 percent to count as significant.

Meng: That stops tiny oscillations from killing the search early.

Lalam: They ran this on the Leonardo pre-exascale supercomputer with A100 GPUs — so the search is heavy offline, but the deployed encoder stays light.

Tom: That offline-online split is a big deal for industrial adoption.

Jane: Now, once the embedding is chosen, what do you actually do with it? The downstream task section answers that.

Page 6 of the paper: Tom: The search picks the embedding. The next section defines how those embeddings get judged.

Jane: Two downstream tasks. First, binary classification of lifetime category.

Lu: They run seven classifiers — Naive Bayes, Logistic Regression, Random Forest, SVM with an RBF kernel, KNN, MLP, and XGBoost.

Meng: That spread avoids architectural bias. If every classifier works, the embedding is genuinely useful.

Tom: Class weighting and balanced sampling handle the imbalance.

Jane: And the headline metric is ROC-AUC — threshold-free, insensitive to imbalance, measuring ranking quality.

Lu: The second task is one-class anomaly detection. Train only on standard-lifetime samples.

Meng: A linear One-Class SVM. The linear kernel keeps it simple and avoids overfitting with tiny training sets.

Tom: The contamination parameter ν stays constrained between 0.01 and 0.20.

Jane: So the model is conservative, flagging only clear deviations from the normal distribution.

Lu: Again ROC-AUC on test labels, so both tasks stay comparable.

Meng: One task tests supervised discrimination. The other tests unsupervised anomaly sensitivity.

Tom: A representation that works on both — that's the reusable, task-agnostic behavior they're chasing.

Jane: And the model choices are deliberate. No Mamba, no large pretrained transformers, because the dataset is far too small and imbalanced.

Lu: The selected architectures balance interpretability, stability, and compute.

Meng: Which brings us to the actual numbers. The binary classification results are on the next page.

Page 7 of the paper: Tom: Downstream tasks defined. Now the results — and they're revealing.

Jane: CNN1D, LSTM, and GRU all found their sweet spot at 16-dimensional embeddings.

Lu: And they stayed stable across classifiers — ROC-AUC between 0.78 and 0.83, as the paper reports.

Meng: That consistency is the fingerprint of classifier-agnostic features.

Tom: The Transformer needed 128 dimensions to peak, reaching 0.84 with Logistic Regression.

Jane: But it also sank to 0.56 with Naive Bayes. Far more sensitive.

Lu: Classic attention behavior in small-data settings — high ceiling, unstable floor.

Meng: Logistic Regression turned out to be the most robust lightweight classifier across the board.

Tom: Then they push into harder imbalance scenarios — 15,000 hours and 20,000 hours.

Jane: Absolute PR-AUC falls as the minority class gets rarer. But the baselines fall even harder.

Lu: At 10k, PR-AUC lands around 1.9 to 2.4 times the prevalence baseline.

Meng: At 20k, it's 4.2 to 4.5 times the baseline. The signal survives.

Lalam: That's the message that matters for reliability engineers — the model beats random even where positives are almost absent.

Tom: LSTM plus Logistic Regression held the most stable behavior across all scenarios.

Jane: Transformer led under mild imbalance but degraded badly under severe skew.

Lu: And CNN1D proved most resilient when the anomaly class nearly vanished.

Meng: The paper's own recommendation — LSTM-LR for general use, Transformer-LR for moderate imbalance, CNN1D-LR for extreme scarcity.

Tom: Which sets up the harder question — can those same embeddings work with no labels at all?

Page 8 of the paper: Tom: The binary results are solid. The one-class experiments are where it gets genuinely tense.

Jane: Training only on standard-lifetime samples, then hunting for anomalies.

Lu: At 10,000 hours, GRU with a 512-dimensional embedding did best — 0.74 ROC-AUC.

Meng: The others trailed, between 0.62 and 0.67.

Tom: But as the normal training set grows, the picture flips.

Jane: At 20,000 hours, CNN1D with just a 4-dimensional embedding hit 0.84.

Lu: LSTM at 64 dimensions reached 0.79. More normal data means a tighter boundary.

Meng: Meanwhile GRU collapsed to 0.54 in that same scenario. Embedding geometry matters.

Tom: The Transformer stayed moderate at 0.68.

Jane: Then the paper checks cross-task consistency — the same embedding used for anomaly detection and binary classification.

Lu: CNN1D at 4 dimensions — 0.84 one-class, 0.81 binary at 20k. Consistent.

Meng: And every experiment runs through a 60-fold repeated holdout, so those numbers are stable.

Tom: The same vectors serve margin-based and discriminative models. That's the task-agnostic claim demonstrated, not just asserted.

Jane: One pretrained backbone supports multiple operational questions.

Lu: For manufacturing, that's huge — the representation is not retrained for every new task.

Meng: So the remaining question is practical — how do you pick the embedding size in a real production setting, and what does it cost?

Page 9 of the paper: Tom: The one-class results show the same embeddings stretch across tasks. The final results section gives practical rules for choosing capacity.

Jane: The rule of thumb is explicit — match embedding size to training volume.

Lu: With 50 or more samples, a broad range works, and higher capacities like 64 to 128 can help.

Meng: Around 19 samples, 8 to 16 dimensions generalize better.

Tom: And with only 4 to 8 samples, ultra-compact 2 to 4 dimensions are surprisingly effective.

Jane: Still above baseline. That's the headline for industrial practice.

Lu: The cost side is stark. CNN1D trains in under half a second per epoch.

Meng: Transformer training can stretch toward roughly 25 seconds per epoch at larger dimensions.

Tom: Parameter counts tell the story — about 50,000 for CNN1D, up to 300,000 for the Transformer.

Jane: Foundation models run to millions or billions of parameters. This is orders of magnitude smaller.

Lu: Inference stays in the sub-millisecond to low-millisecond range.

Meng: That makes real-time quality assessment on a test line feasible.

Lalam: And the da-NAS search cost is paid once, offline. The deployed system stays small and fast.

Tom: So you get foundation-model-like flexibility without the data-center bill.

Jane: That's the practical bridge into cryocooler production.

Lu: The paper wraps up by connecting all of this to deployment and future work.

Meng: And the conclusion pulls the whole argument together.

Conclusion: Tom: We've covered pipeline, data, search, and results. So where does this leave us?

Jane: The core message — non-destructive lifetime prediction from standard telemetry, without destroying expensive hardware.

Lu: That replaces costly destructive testing with a model running on data already being collected.

Meng: And it's built for low-volume manufacturing. Ninety-five labeled samples is the real-world reality.

Tom: The future directions are sensible. First, uncertainty quantification and physics-informed priors.

Jane: For safety-critical aerospace, calibrated confidence matters as much as raw accuracy.

Lu: Second, semi-supervised, positive-unlabeled, and active learning strategies.

Meng: Squeeze more from the limited labels without running more destructive tests.

Tom: Third, transferability to other cryocooler types, manufacturers, and aerospace components.

Jane: The framework was built on one cooler, but the paradigm is general.

Lalam: ESA's RASCOSA project funding shows this is grounded in actual mission needs, not just academic curiosity.

Meng: The whole approach fits the push for advanced eye in satellite reliability.

Tom: And that's the throughline — foundation-model ambitions adapted to tiny industrial datasets.

Jane: Small models, capacity control, and honest evaluation under real constraints.

Lu: The paper gives the community a reproducible template for that.

Meng: Anyone working with scarce sensor data should give it a careful read.

Tom: Great discussion, everyone. We'll close the book on cryocoolers — next paper coming up soon.

Gregor Molan, Grafika Jati, Francesco Barchi, Andrea Acquaviva, Aljaž Osterman, Martin Molan

cs.LG, cs.AI

Submitted: 2026-08-07

Updated: 2026-08-10

Comments: 48 pages

Journal ref: Reliability Engineering and System Safety 277 (2027) 113105

DOI: 10.1016/j.ress.2026.113105

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 45/100

The gist: The paper addresses the challenge of cryocooler lifetime prediction in satellite-based thermal imaging systems, where reliable lifetime prediction is essential for mission planning.

Key concepts

Small-data representation model
A machine learning model designed to learn useful data representations from very few labeled examples. Instead of relying on massive datasets, it uses unsupervised training on unlabeled data to create compact, reusable vectors that can then be applied to different tasks like classification or anomaly detection.
Embedding size
The number of dimensions in the vector representation of data. The paper shows that with small datasets, smaller embeddings (like 2 or 4 dimensions) work better than larger ones, because they avoid overfitting. Larger embeddings (like 128 or 512) only help when more labeled data is available.
One-class anomaly detection
A method to identify unusual data points by training only on normal examples. Here, the model learns what standard-lifetime cryocoolers look like and flags anything that deviates from that pattern as anomalous, which helps predict failures without needing labeled failure examples.
da-NAS
A dimension-aware neural architecture search that automatically selects the best embedding size for a given dataset. It tests different sizes from 2 to 512, using a stop policy to save time, and helps match model complexity to the amount of available data.

Terminology

Summary

The paper addresses the challenge of cryocooler lifetime prediction in satellite-based thermal imaging systems, where reliable lifetime prediction is essential for mission planning. The authors note that "large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack." This motivates the proposal of a new paradigm.

The work is motivated by the industrial context of cryocooler quality assurance: "Traditional quality assurance relies heavily on destructive lifetime testing, where coolers are operated continuously until failure to estimate durability, typically targeting a minimum operational threshold of 8,000 hours for space-grade equipment. The authors emphasize that this process is costly, time-consuming, and results in the destruction of high-value components."

The authors identify several fundamental challenges in cryocooler lifetime prediction:

  1. Limited labeled dataset: Labeled data from lifetime tests is scarce due to the destructive nature of the process, which restricts the applicability of conventional supervised learning approaches. The dataset contains only 1,305 unlabeled and 95 labeled sequences.

  2. Severe class imbalance: The dataset exhibits strong skew, with long-lifetime units significantly rarer than standard units.

  3. Heterogeneous sensor modalities: Cryocooler telemetry comprises both time-domain operational signals and frequency-domain measurements with distinct statistical properties.

  4. Variable-length sequences: The telemetry data consists of sequences of differing durations.

  5. Domain shift across test regimes: Data collected under different testing conditions exhibits distributional differences, combined with measurement noise, which can degrade generalization performance.

The core contribution is a framework called FSD-RM (Family of Small-Data Representation Models), defined as "a set of heterogeneous encoders—CNN1D, LSTM, GRU, and Transformer—pretrained to learn a generic representation of multivariate time-series data. These models share the same objective but differ in architectural principles, capacity, and inductive bias."

The authors explicitly contrast their approach with large foundation models: "Rather than aiming to replicate the scale of existing foundation models, our objective is to approximate their functional properties—such as task-agnostic representations and cross-task generalization—within a resource-constrained, domain-specific setting."

The framework employs an unsupervised sequence-to-sequence representation learning architecture designed to encode multivariate time-series data into a compact latent embedding. The four encoder variants are:

  • Seq2SeqCNN1D: A lightweight convolutional encoder–decoder, Seq2SeqCNN1D, designed for sequence reconstruction and embedding extraction from multivariate time series using stacked 1D convolutions to produce compact latent representations.

  • Seq2SeqLSTM: Uses recurrent layers to capture long-range dependencies and uses sequence reconstruction as a self-supervised learning objective.

  • Seq2SeqGRU: Provides a parameter-efficient recurrent alternative while retaining the ability to model long-range dependencies — GRUs reduce model complexity compared to LSTMs, making them well suited for small-data scenarios.

  • Seq2SeqTransformer: Employs self-attention mechanisms to model complex temporal dependencies without recurrence — Transformers capture long-range interactions through attention mechanisms and enable scalable representation learning with flexible capacity control.

All encoders use temporal pooling to produce a compact embedding vector that captures the input's overall temporal dynamics (e.g., global average pooling for CNN1D, attention-based pooling for Transformer).

The authors introduce a NAS framework da-NAS which "organizes the search as a dimension-wise hierarchical process. It sequentially explores embedding dimensions (from 2 to 512), leveraging cross-dimensional knowledge transfer to efficiently find the optimal model configuration."

da-NAS operates under two distinct regimes:

  • Full Exploration Regime (D ∈ 2, 4, 8, 16): The optimizer performs an exhaustive search of up to 4000 trials per dimension with early stopping disabled.

  • Target-Guided Regime (D ≥ 32): The framework activates the –beat lower dim flag, enabling cross-dimensional early stopping once the validation loss LD matches or surpasses the best result of the previous dimension TD−1.

The Cross-Dimensional Stop Policy (CDSP) uses a 10% relative threshold for significant improvement and Plateau Detection: When no improvement beyond δ occurs for a fixed patience window (10 trials for the smallest dimension or 5 trials after a BLD), the study terminates automatically.

The preprocessing pipeline includes three stages:

  • Filtering Invalid Data: Values of T housing outside −50°C or above 100°C are flagged as sensor malfunction, corrupted measurements, or outliers.

  • Filtering NaN Values: The preprocessing also includes a step to remove sequences that contain missing values (NaN).

  • Normalization: The authors selected Robust Scaler, as it demonstrated greater resilience to outliers and non-Gaussian feature distributions commonly observed in our sensor data, resulting in more stable training dynamics.

The data comes from a structured post-production testing process: "Run-in is the initial phase, lasting 150 hours with data collected every minute... the Noise Test runs for 15 minutes with a 1-second resolution, measuring vibration frequencies... The ESS (Environmental Stress Screening) includes four phases: RT (Room Temperature), LT (Low Temperature, -40°C), HT (High Temperature, 71°C), and Post ESS RT... Finally, the Life Test involves continuous operation, with ESS tests repeated every 500 hours."

The telemetry includes time-domain features (T housing, T ambient, DC Bus, I motor, P heater, DCVolts, as well as high-level control variables like Error Reg, CD, and RPM) and frequency-domain NoiseTest features (33 frequency-domain features, including spectral power across multiple frequency bands (20 Hz to 20,000 Hz), RPM, and total band power). Importantly, time-domain and frequency-domain telemetry are not fused into a single unified input representation. Instead, each modality is processed using a modality-specific preprocessing pipeline.

The downstream task "formulates lifetime estimation as a binary classification problem defined by a threshold T. Units with lifetime ≤ T are labeled as class 0 (standard lifetime), whereas those exceeding T are labeled as class 1 (long lifetime units)." Three thresholds are evaluated:

  • 10,000-hour threshold: 63 samples in class 0 and 32 samples in class 1 (mild imbalance 2:1)

  • 15,000-hour threshold: 73 samples in class 0 and 22 samples in class 1 (medium imbalance 3:1)

  • 20,000-hour threshold: 84 samples in class 0 and 11 samples in class 1 (severe imbalance 7:1)

Classification is preferred over regression because it closely mirrors operational decision-making, classification models are generally more robust to noise and outliers, in many industrial contexts, labels are limited to categorical or censored data, threshold-based classification outputs are easier to interpret, and regression on single-signal time series often struggles to model the complex temporal and multivariate patterns underlying system degradation.

Two downstream tasks validate the representations:

  1. Binary Classification: Embeddings are used as input to a diverse set of classical machine learning classifiers—Naive Bayes, Logistic Regression, Random Forest, SVM (RBF), KNN, MLP, and XGBoost—to avoid architectural bias and to ensure classifier-agnostic evaluation. Performance is measured with ROC–AUC, as it is threshold-free, insensitive to class imbalance, and reflects ranking quality rather than fixed decision boundaries.

  2. One-Class Classification: The second downstream task evaluates whether a single encoder can support unsupervised anomaly detection using only class 0 (short-lifetime) samples for training. This uses a One-Class SVM with a linear kernel with the contamination parameter ν... constrained within 0.01 ≤ ν ≤ 0.20.

The authors list six core contributions:

  1. Small-Data Representation Modeling for Cryocooler Satellite Telemetry—demonstrating that generalization can emerge under extremely small-data constraints, far below the scale required by existing large-scale pretrained time-series models.

  2. Capacity-Scalable Family of Small-Data Representation Models—the FSD-RM encoder family parameterized by embedding capacity enabling controlled scaling of representational complexity.

  3. Cross-Task Transferability of Learned Representations—a single pretrained representation supports multiple heterogeneous downstream tasks—binary fault classification and one-class anomaly detection—without retraining the encoder.

  4. Multi-Regime Robustness Under Data Scarcity and Class Rarity—spanning from full-data availability down to ultra-low-sample scenarios.

  5. Dimension-Aware NAS for Optimal Embedding Capacity—da-NAS selects the optimal embedding dimension using a progressive dimension schedule, a Beat-Lower-Dimension improvement rule, and a simple dimension-aware early stopping criterion.

  6. System-Level Design Beyond Single-Model Optimization—a system-level framework that integrates family-based representation learning, capacity scaling, cross-task reuse, and dimension-aware NAS into a unified pipeline.

With the best embedding dimension per encoder at the 10k-hour threshold, the results show:

  • The CNN1D, LSTM, and GRU encoders consistently produced stable, low-dimensional representations, with the optimal embedding dimension converging at 16 for all three models, achieving ROC–AUC values between 0.78 and 0.83 across all downstream classifiers.

  • "The Transformer encoder achieved the highest peak ROC–AUC (0.84 using Logistic Regression), but only when using a substantially larger embedding dimension (128). Performance varied considerably across classifiers, ranging from strong (Logistic Regression) to notably weak (Naive Bayes)."

  • Across all architectures, Logistic Regression emerged as the most robust and lightweight classifier.

Using encoder paired with Logistic Regression across scenarios:

  • 10,000h (mild): The Transformer achieves the highest scores (F1-Macro = 0.74, ROC–AUC = 0.84, PR–AUC = 0.76).

  • 15,000h (medium): The Transformer again yields the strongest results (F1-Macro = 0.68, ROC–AUC = 0.73, PR–AUC = 0.58).

  • 20,000h (severe): CNN1D unexpectedly achieves the highest ROC–AUC (0.84)... PR–AUC values fall within 0.44–0.47, which remains meaningful given the severe imbalance: the baseline PR–AUC, equal to the positive prevalence, is only 0.105.

The authors note: PR–AUC decreases monotonically with increasing imbalance severity, but relative to the baseline PR–AUC dictated by test-set prevalence... all models yield higher performance, reaching 4.2–4.5× baseline in the severe scenario.

The results show that anomaly separability is influenced by both the encoder's embedding geometry and the statistical quality of the normal-only training set. Key findings: in the 10,000h scenario, only the GRU(512) encoder produces a sufficiently compact representation to support a well-defined one-class boundary, achieving a ROC–AUC of 0.74; in the 15,000h scenario, CNN1D(4) [achieves] stronger separability (0.72); and in the 20,000h scenario, CNN1D(4) attains the highest separation (0.84).

Across all scenarios, the representations produced by CNN1D (4D), LSTM (64D), GRU (512D), and Transformer (8D) remain transferable, achieving consistent performance on both tasks despite their differing objectives. The authors conclude: The consistent performance observed across all encoders indicates that the latent features encode task-agnostic structure, allowing them to support both density-based and discriminative classifiers.

Using LSTM with Logistic Regression, the results show a clear dependency between embedding capacity and data size: with full training data (76 samples), dimensions approximately 8D up to 128D—achieves high performance (ROC–AUC ≈ 0.75–0.85), but as the training set becomes smaller (e.g., 19 samples or fewer), smaller embeddings (4D–16D) achieve higher performance than larger ones, and in the most extreme low-data scenarios (8 or 4 samples)... small embeddings yet still remain the most stable and continue to achieve ROC–AUC values meaningfully above the random baseline of 0.50.

Across training fractions, the Transformer degrades more sharply than lower-capacity encoders. In contrast, CNN1D (16) and LSTM (16) exhibit much more graceful degradation. In severe imbalance (20k), CNN1D and LSTM encoders again show the highest robustness, achieving 0.70–0.62 at 10% and 5% [training fractions], while Transformer remains most affected by compounding scarcity and skew.

The authors provide practical guidelines: "embedding capacity should be matched to the available training volume. When the training set is relatively large (e.g., ≥ 50 samples), a wide range of embedding sizes can be utilized effectively... for moderate data regimes (e.g., 19 samples), smaller embedding sizes (e.g., 8D–16D) consistently yield superior generalization... In extremely low-data settings (e.g., 4–8 samples), ultra-compact embeddings (2D–4D) become not only necessary but surprisingly effective."

The proposed models are lightweight: the evaluated encoders range from approximately 50K to 300K parameters compared to millions to billions of parameters in large-scale models. CNN1D architectures consistently exhibit the lowest computational cost, maintaining training times per epoch below 0.5 seconds for most configurations, while Transformer models... training time per epoch increasing to approximately 25 seconds for larger embedding dimensions. The da-NAS overhead is incurred only once: the deployed encoder operates with the same lightweight footprint as the underlying FSD-RM model.

The authors decline direct empirical comparison with foundation models: "We do not include empirical results for TimesFM and UniTS, as applying or fine-tuning these models in this data regime—without large-scale domain-relevant pre-training—is unlikely to yield a controlled, comparable evaluation or results that are scientifically meaningful or directly comparable. They note in Table 1 that models like TimesFM, UniTS, PatchTST, and Mamba are designed for large-scale pretraining and have limited applicability in the small-data regime, whereas FSD-RM is specifically tailored to a constrained industrial setting with limited data availability."

The paper concludes that "the proposed framework leverages unsupervised sequence-to-sequence representation learning and a dimension-aware neural architecture search (da-NAS) to enable capacity-scalable embeddings. The resulting representations demonstrate cross-task generality, supporting both binary classification and one-class anomaly detection without retraining."

The practical impact: The proposed approach enables non-destructive lifetime prediction using standard telemetry data, reducing reliance on costly and time-consuming lifetime testing.

Future work is organized along three directions: "(1) extend the framework with uncertainty quantification and physics-informed priors... (2) investigate semi-supervised, positive–unlabeled, and active learning strategies to better exploit the limited labeled data... (3) evaluate the transferability of the proposed framework to other cryocooler types, manufacturers, and broader aerospace components."

The paper was "Accepted for publication in Reliability Engineering & System Safety" and was supported by ESA's RASCOSA project (RotAry Stirling CryOcoolers for Space Applications), with HPC resources from EuroHPC on the Leonardo Booster at CINECA, Italy.

Improvements for AI systems

Here are the specific improvements to AI systems that can be drawn from this paper, along with the resulting capabilities:


  1. Operate effectively in extreme small-data regimes without large-scale pretraining.

The system can replace the pretrain on massive data, then fine-tune paradigm with a family of lightweight encoders (CNN1D, LSTM, GRU, Transformer) trained via unsupervised sequence-to-sequence reconstruction. It learns generic, task-agnostic time-series representations from as few as 1,305 unlabeled sequences and 95 labeled samples — far below the data requirements of foundation models.

  1. Auto-select the optimal embedding capacity for a given dataset via dimension-aware NAS.

The system can use a hierarchical, cross-dimensional search process (da-NAS) that systematically explores embedding dimensions from 2 to 512, transferring knowledge from lower dimensions to higher ones. Cross-dimensional early stopping (using a beat lower dimension rule with a 10% improvement threshold) reduces search overhead while still finding the best capacity. The result: the system automatically matches model complexity to available data, avoiding overfitting from overly large embeddings.

  1. Gracefully degrade as training data shrinks, instead of collapsing.

By following the paper's capacity-scaling guidelines, the system can dynamically choose ultra-compact embeddings (2D–4D) for extremely small datasets (4–8 samples) and moderate embeddings (8D–16D) for medium data (≈19 samples). This yields ROC–AUC meaningfully above 0.50 even with only 4 training samples, whereas high-capacity models fail.

  1. Achieve cross-task transfer from a single pretrained encoder.

One learned representation can support both binary classification (e.g., lifetime thresholding) and one-class anomaly detection (e.g., identifying abnormal units) without retraining the encoder. This works because the unsupervised objective produces embeddings that are simultaneously discriminative and density-friendly — enabling both discriminative classifiers (Logistic Regression, SVM, XGBoost) and density-based detectors (One-Class SVM) to perform well.

  1. Maintain robust performance under severe class imbalance.

The system can handle positive-class prevalence as low as 10.5% (i.e., a 7:1 imbalance) while still achieving PR–AUC 4.2–4.5× the baseline prevalence. It does this by using threshold-free evaluation (ROC–AUC, PR–AUC) and by selecting embeddings that preserve rare-class structure — with CNN1D and LSTM proving more robust than Transformer in skewed, sparse regimes.

  1. Deploy lightweight models with near-real-time training and inference.

The system can run on edge or resource-constrained hardware because its encoders have only 50K–300K parameters (versus millions/billions in foundation models). CNN1D trains in under 0.5 seconds per epoch, while Transformer takes ≈25 seconds per epoch for larger embeddings — still far cheaper than any large-model fine-tuning. The NAS overhead is incurred once; the deployed model remains lightweight.

  1. Handle heterogeneous sensor modalities without forced fusion.

The system processes time-domain telemetry and frequency-domain spectra separately with modality-specific preprocessing, then learns representations independently for each. This avoids the distributional mismatch that occurs when mixing 1 Hz temporal signals with 20 kHz spectral features in a single input — improving generalization across domain shifts in testing conditions.

  1. Resist outliers and non-Gaussian sensor noise via robust normalization.

The system uses a Robust Scaler (instead of standard or min-max scaling), making training stable when telemetry contains sensor malfunctions, corrupted measurements, or extreme values. It also filters physically invalid values (e.g., housing temperature outside −50°C to 100°C) before learning.

  1. Provide practical, uncertainty-aware guidance for embedding selection.

The system can expose a capacity-suitability diagnostic: given the training set size and class imbalance, it recommends an embedding dimension range and predicts the expected performance degradation under further data scarcity. This makes model selection transparent and repeatable in industrial settings.

  1. Enable non-destructive lifetime prediction for high-value assets.

The system can replace destructive lifetime testing by predicting whether a unit will exceed a lifetime threshold (10k/15k/20k hours) using only standard telemetry collected during run-in, noise tests, and environmental stress screening. This saves cost and time, and avoids destroying expensive components.

  • Given a small labeled set (e.g., 95 labeled examples) and a larger unlabeled set (e.g., 1,305 sequences), it can pretrain four encoder variants, run da-NAS to pick the best embedding dimension, and then use the frozen embeddings to both classify long vs. standard lifetime and flag anomalous units — all in one pipeline.

  • In an industrial setting with only 19 labeled samples, it can still achieve ROC–AUC ≈ 0.75–0.85 by choosing an 8D–16D LSTM embedding, while a Transformer with large embedding would degrade below 0.6.

  • In a severely imbalanced scenario (11 positive samples out of 95), it can identify the rare long-lifetime units with PR–AUC 4.5× the random baseline, using CNN1D(4) embeddings and a One-Class SVM for anomaly detection.

  • It can transfer the same pretrained representations to a different downstream task (e.g., fault detection vs. quality control) without re-running the expensive representation learning step.

  • It can run entirely on modest computing infrastructure — no GPU cluster required — making it practical for SMEs, test labs, and on-site quality assurance.

In summary, the improved AI system is a small-data, capacity-aware, cross-task time-series learning framework that outperforms large pretrained models in resource-constrained industrial settings by matching model capacity to data availability and by learning reusable representations through unsupervised reconstruction.

Abstract

Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning, but they rely on abundant, diverse data that industrial and scientific domains often lack. We therefore propose the FSD-RM (Family of Small-Data Representation Models) paradigm as a practical alternative for limited, domain-specific telemetry. Rather than relying on large-scale pretraining, we focus on capacity-controlled representation learning using established encoder architectures (CNN1D, LSTM, GRU, Transformer), selected for their suitability in small-data settings and interpretability. These encoders are trained unsupervised on multivariate telemetry data and integrated into a two-stage pipeline for downstream lifetime prediction. To systematically examine architectural trade-offs under data constraints, we employ dimension-aware neural architecture search (NAS) to jointly optimize model capacity and input dimensionality. Experiments on cryocooler telemetry show that the proposed approach achieves competitive predictive performance while reducing training cost and model complexity. The contribution lies in combining established representation learning techniques within a coherent, NAS-driven framework tailored to small-data regimes, with explicitly defined parameter settings and design choices. The results indicate that effective representation learning can be achieved without large-scale pretraining when appropriate inductive bias and capacity control are applied.

Sources

Related papers