TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

arXiv:2608.11951 · cs.LG, cs.AI · Submitted 2026-08-12 · Read on arXiv

Karim Aly, Alexei Sharpanskykh, Jacco Hoekstra

Delft University of Technology

cs.LG, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Preprint submitted to journal

Code: https://github.com/karimyehia92/TailBooster

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: TailBooster is a dual-layer generative framework designed to address two complementary failure modes of conventional deep generative models applied to mixed-type tabular aviation records: the

Terminology

Summary

TailBooster is a dual-layer generative framework designed to address two complementary failure modes of conventional deep generative models applied to mixed-type tabular aviation records: the systematic under-representation of distributional tails and the production of operationally invalid synthetic instances. The framework combines IQR-based extreme subset extraction with dedicated Tabular Variational Autoencoder (TVAE) generative models and autoencoder-based operational cleaning, and was evaluated across five distinct dimensions on publicly available U.S. domestic flight records.

The dual-layer design brackets a TVAE generative stage with two anomaly detection layers: a statistical layer that precedes generation and a deep learning layer that follows it. The framework takes as input the full historical dataset together with two user-defined feature lists: (i) the target features whose distributional tails are to be augmented, used by the statistical layer to isolate extreme subsets, and (ii) the operationally correlated features, i.e., features whose joint values are governed by operational constraints, used by the deep learning layer to characterise operational feasibility. The process starts with the statistical layer, which performs IQR-based extreme-subset extraction prior to generation, isolating one extreme subset per target feature to supply a tail-concentrated training signal to dedicated generative models. A TVAE is then trained on the full dataset and another TVAE on each extreme subset; synthetic records are then sampled from all models, and the resulting records are filtered to retain only those whose origin–destination airport pairs appear in the historical data. After generation, the deep learning layer applies pre-trained autoencoders to these candidate synthetic records, discarding samples that fall outside the empirical operational envelope learned from the historical data. This data-driven cleaning step enforces operational validity without requiring hand-crafted domain rules to be available.

The pipeline yields three synthetic or augmented datasets—Naïve Synthetic, Augmented Synthetic, and Augmented Real—which, together with the original Real dataset, are assessed across five complementary dimensions: diversity, statistical similarity, fidelity, operational validity, and regression utility on extreme subsets.

The three preservation checks confirmed that the augmentation and cleaning processes did not degrade the qualities already achieved by conventional generation: diversity was maintained across all real clusters, and both statistical similarity and fidelity improved relative to the Naïve Synthetic baseline. The two primary improvement targets were both met. The data-driven operational cleaning layer markedly reduced the proportion of operationally implausible synthetic records relative to conventional generation. Targeted extreme-value augmentation consistently improved the predictability of extreme events: across six regression models, training on the Augmented Synthetic reduced Mean Absolute Error (MAE) by approximately 47–49% on extreme Air Time (min) and 29–57% on extreme Arrival ∆T (min) relative to training on Naïve Synthetic data, while training on the Augmented Real consistently outperformed training on real historical records alone, confirming that the improvements are a property of the augmented data rather than of any particular predictive algorithm.

The results demonstrate that TailBooster is beneficial both for practitioners with access to historical flight records, by enriching the representation of extreme-value regions through augmentation of real data with operationally valid synthetic extremes, and for those without such access, by providing augmented synthetic data that substantially outperforms conventionally generated alternatives. Being fully data-driven and generative-model-agnostic, the framework extends naturally to any domain where extreme-event prediction is operationally critical and domain-specific rules governing operational validity are unavailable.

Improvements for AI systems

Improvements to AI Systems:

  1. Extreme-Value Augmentation Module: Integrate TailBooster’s IQR-based extreme subset extraction and dedicated TVAE training into any generative model pipeline. This enables AI systems to generate synthetic data that over-samples rare, high-impact events (e.g., equipment failures, fraud spikes, rare disease cases) without manual threshold tuning.

  2. Operational Validity Autoencoder Filter: Add a post-generation autoencoder layer trained on historical joint feature distributions to reject synthetic samples that violate implicit operational constraints (e.g., impossible combinations of flight origin–destination, sensor readings, or transaction patterns). This makes AI systems produce only physically/logically feasible outputs, eliminating the need for hand-coded rules.

  3. Dual-Layer Anomaly Bracket: Implement a statistical pre-filter (for tail isolation) and a deep-learning post-filter (for validity) around any core generative model (VAE, GAN, diffusion). This improves AI systems’ robustness to distribution shift and reduces hallucination rates in structured data generation tasks.

  4. Extreme-Event Regression Enhancer: Use the augmented real dataset (real + operationally valid synthetic extremes) as training input for predictive models. This yields AI systems with 29–57% lower MAE on extreme-value predictions (e.g., arrival delays, energy grid overloads, financial tail risks) compared to training on original data alone.

  5. Model-Agnostic Data Enrichment Service: Package the framework as a reusable preprocessing layer for any tabular AI system. The improved system can automatically generate balanced training sets that preserve diversity, improve statistical similarity, and boost fidelity for downstream classifiers or regressors, especially in domains with scarce extreme-event records.

What the Improved AI System Can Do:

  • Generate synthetic tabular data that accurately represents rare operational events (e.g., extreme flight delays, rare medical complications) with high fidelity and validity.

  • Predict extreme outcomes with significantly higher accuracy (up to 57% error reduction) by training on augmented data.

  • Automatically clean synthetic outputs to ensure they adhere to hidden operational rules, reducing invalid predictions in production.

  • Operate without domain-specific rule engineering, making it adaptable to new industries (logistics, finance, healthcare) where extreme events are critical but rare.

Abstract

Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such events are rare in historical records, leaving insufficient training signal for machine learning models. Synthetic data augmentation offers a principled solution, but conventional generative models under-represent distributional tails and give no guarantee against operationally infeasible instances, such as a short air time paired with a long flight distance. No existing approach addresses both limitations for mixed-type tabular records. We propose TailBooster, a dual-layer generative framework combining generative modelling with two anomaly detection layers. A statistical layer extracts extremes via the interquartile range, supplying tail-concentrated training signal to dedicated generative models, here a Tabular Variational Autoencoder. A deep learning layer then applies autoencoder-based cleaning, discarding synthetic records that violate the operational envelope learned from historical data. The framework was evaluated on US flight records across five dimensions: diversity, statistical similarity, fidelity, operational validity, and utility, the latter two being the primary improvement targets. Data-driven cleaning markedly improved operational validity, while targeted augmentation enhanced utility for extreme-event prediction. Across six regression algorithms, training on the framework's records reduced Mean Absolute Error by 47-49% on extreme air time and 29-57% on extreme arrival delay prediction relative to conventional synthetic data, with comparable gains when real records were enriched with synthetic extremes. Being fully data-driven and model-agnostic, TailBooster extends to domains where extreme-event prediction is critical and domain-specific rules are unavailable.

Sources

Related papers