ADEPT: A Unified Framework for Deep Learning Test Adequacy

arXiv:2608.12144 · cs.SE, cs.LG · Submitted 2026-08-12 · Read on arXiv

Yidi Kao, Shawn Burnham, Tommi Rose Fahy, Ali Ghanbari

Auburn University

cs.SE, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Proceedings of 35th ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2026)

DOI: 10.1145/3837729.3840488

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 72/100

The gist: ADEPT: A Unified Framework for Deep Learning Test Adequacy This paper presents ADEPT, an extensible framework for running deep learning (DL) test adequacy metrics.

Terminology

Summary

ADEPT: A Unified Framework for Deep Learning Test Adequacy

This paper presents ADEPT, an extensible framework for running deep learning (DL) test adequacy metrics. The motivation is that, over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.g., neuron activation behavior, latent feature coverage, decision-boundary exploration, etc. However, these metrics are typically released as independent research prototypes with substantially different installation and preprocessing requirements, execution workflows, and configuration mechanisms. These complications make them quite difficult to reproduce, compare, and adopt in research work and practical deployment alike.

The paper presents the engineering details of ADEPT, a framework that integrates representative adequacy techniques, including neuron-coverage-based metrics, surprise adequacy, input distribution coverage, boundary coverage, and source- and model-level mutation score, under a consistent execution workflow. ADEPT provides a template-based metric interface with well-defined extension points for integrating new adequacy metrics. Furthermore, it provides YAML-based configuration management, preprocessing-cache reuse, and structured result reporting, making it easy to use in any research and development workflows. ADEPT is designed for researchers and practitioners who wish to reproduce and apply adequacy metrics without spending days or weeks implementing missing tooling or configuring disparate research prototypes.

The framework consists of four main components: (1) a template-based metric interface that provides access to representative adequacy metrics and enables new metrics to be integrated through well-defined extension points, (2) metric-specific processing modules that perform the preprocessing required by each adequacy technique and generate the corresponding intermediate artifacts, (3) a cache management component that stores and retrieves reusable intermediate artifacts across repeated evaluations, and (4) an adequacy scoring component that computes the final adequacy score for a given test dataset and generates a final report containing the score and execution metadata.

The metric interface provides a unified entry point for executing different DL test dataset adequacy metrics. Each metric is implemented as a plug-in module that follows the same framework-level abstraction while preserving its own processing and scoring logic. ADEPT integrates several representative DL test adequacy techniques: neuron coverage and its variations (NC-series), two flavors of surprise adequacy (LSA and DSA), input distribution coverage (IDC), deep boundary coverage (DBC), and source- and model-level mutation score (SLMS and MLMS).

NC-series metrics measure adequacy based on neuron activation behavior observed during model execution. NC and Top-k Neuron Coverage (TKNC) assess whether individual neurons are activated above a threshold or ranked among the top-k most active neurons for a given input. k-multisection Neuron Coverage (KMNC), Neuron Boundary Coverage (NBC), and Strong Neuron Activation Coverage (SNAC) characterize coverage relative to the activation ranges observed on the training set.

LSA and DSA assess test adequacy by measuring how surprising test inputs are relative to the training data based on their activation traces. IDC measures adequacy in a learned latent feature space, employing a variational autoencoder (VAE) to construct a compact representation of the input distribution and applying combinatorial interaction testing to assess feature diversity within a test dataset. DBC measures test adequacy from a decision-space perspective, focusing on how test inputs cover the model’s decision boundaries. SLMS and MLMS evaluate test datasets based on their ability to distinguish mutant models from the original model. The underlying intuition is that more effective test datasets should be capable of exposing behavioral differences introduced by injected faults. ADEPT supports both MLMS, which applies mutation operators directly to trained DL models, and SLMS, which introduces faults before the training process to generate mutants.

Different adequacy metrics require different forms of preprocessing before adequacy scores can be computed. For NC-based metrics that rely on training-set activation statistics (KMNC, NBC, SNAC), ADEPT performs neuron profiling on the training dataset. For LSA and DSA, activation traces are extracted and used for surprise estimation. IDC relies on pre-trained VAE models to construct latent feature representations, while DBC performs decision-boundary extraction to characterize model behavior near classification boundaries. ADEPT carries out MLMS à la DeepMutation++ and SLMS à la DeepCrime.

To avoid redundant computation, ADEPT provides a cache management component that stores and retrieves reusable artifacts generated during metric-specific processing. Examples of cached artifacts include neuron profiles for some NC-series metrics, activation traces for LSA/DSA, decision-boundary models for DBC, and generated mutant models for SLMS and MLMS. Cached artifacts are associated with the target model, dataset, metric, and creation time. When compatible cached artifacts are available, ADEPT allows users to choose either to reuse the existing artifacts and proceed directly to scoring, or to regenerate them.

The adequacy scoring component uses the artifacts produced by the processing stage to compute the final score according to the selected metric. Coverage-based metrics report the proportion of covered adequacy elements, LSA/DSA derives scores from activation trace distributions, IDC measures latent feature interaction coverage, DBC quantifies decision-boundary coverage, and SLMS/MLMS reports the classic mutation score, computed as the proportion of killed mutants over all generated mutants, i.e., MS = killed mutants / all mutants. For each run, ADEPT produces the metric adequacy score together with lightweight execution metadata, including preprocessing time, scoring time, and cache utilization information.

ADEPT is implemented in Python and can be used through a command-line interface. The framework currently accepts NumPy (.npy) datasets as input and supports Keras-based DNN models for metrics that require access to the tested model. Users can execute an adequacy metric by specifying the target model, dataset name, test inputs, optional training data, and a metric-specific configuration file. Metric-specific options are specified through YAML configuration files, allowing users to adjust processing and scoring behavior without modifying the source code. If a configuration file is not provided, ADEPT falls back to default values defined in each metric module. After each run, ADEPT generates a structured JSON report containing the final adequacy score, important configuration, preprocessing time, scoring time, and cache utilization information.

The related works section organizes DL test adequacy metrics into four categories: structural adequacy metrics (e.g., NC by Pei et al. and its variants in DeepGauge, DeepCT, MC/DC-style adequacy), input-space adequacy metrics (e.g., SA and IDC), boundary-based metrics (e.g., DeepBoundary), and mutation-based metrics (e.g., DeepMutation++, DeepMutation, and DeepCrime). The paper notes that existing available implementations, such as DeepXplore, IDC, and DeepCrime, only focus on specific metrics rather than providing a unified workflow, and many proposed adequacy metrics remain difficult to run in practice due to lack of public implementations, broken packages, limited maintenance, or complicated environment setup.

The paper concludes that ADEPT simplifies the execution, comparison, and deployment of diverse adequacy techniques by providing extensible mechanisms for configuration, metric processing, and reporting. As future work, the authors plan to incorporate more metrics in the framework. The research is partially supported by NSF grant #2446393. ADEPT is publicly available at https://zenodo.org/records/21682100, and a demo video is available at https://aub.ie/ADEPT video.

Improvements for AI systems

Improvements to AI Systems:

  1. Automated Test-Suite Quality Assessment for AI Models
  • Integrate ADEPT as a built-in validation layer in AI development pipelines. The system can automatically compute multiple adequacy metrics (neuron coverage, surprise adequacy, mutation score, etc.) on any test dataset before model deployment.

  • Capability: The AI system can flag under-tested regions (e.g., low neuron activation coverage or low decision-boundary coverage) and suggest targeted test inputs to improve robustness, reducing the risk of silent failures in production.

  1. Adaptive Test-Input Generation via Adequacy Feedback
  • Use ADEPT’s scoring outputs as a reward signal for generative models (e.g., GANs or diffusion models) that create new test inputs. The generator iteratively produces inputs that maximize coverage gaps identified by ADEPT (e.g., uncovered neurons or low surprise adequacy).

  • Capability: The AI system can autonomously explore rare or adversarial input regions, improving model generalization and detecting edge-case bugs that manual testing would miss.

  1. Cross-Model Comparison and Regression Testing
  • Leverage ADEPT’s unified interface to compare adequacy scores across model versions (e.g., before/after retraining or quantization). The system can automatically detect if a new model version reduces test coverage on critical data distributions.

  • Capability: The AI system can enforce a “coverage regression gate” in CI/CD pipelines, blocking deployment if adequacy drops below a threshold, ensuring that updates do not silently degrade test effectiveness.

  1. Resource-Efficient Adequacy Monitoring
  • Use ADEPT’s cache management to reuse intermediate artifacts (e.g., neuron profiles, activation traces) across multiple evaluations. The AI system can schedule periodic adequacy checks without recomputing expensive preprocessing, enabling real-time monitoring of test suites in large-scale systems.

  • Capability: The AI system can continuously track test adequacy during model drift or data shifts, alerting developers when coverage degrades due to changing input distributions.

  1. Explainable Adequacy Reports for Debugging
  • Integrate ADEPT’s JSON reports into an AI-driven debugging assistant. The assistant can parse adequacy scores and metadata (e.g., which neurons are uncovered, which mutants survive) to generate human-readable explanations of model weaknesses.

  • Capability: The AI system can automatically diagnose why a model fails on certain inputs (e.g., “low boundary coverage near class X”) and recommend specific retraining data or architectural changes, accelerating root-cause analysis.

  1. Multi-Metric Ensemble for Robustness Scoring
  • Combine ADEPT’s diverse metrics (structural, input-space, boundary, mutation) into a single composite robustness score using a learned weighting scheme. The AI system can optimize these weights based on downstream task performance (e.g., accuracy on real-world validation sets).

  • Capability: The AI system can provide a holistic, calibrated measure of test-suite quality that correlates better with real-world model reliability than any single metric, enabling more informed decisions on when to trust a model.

Abstract

Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.g., neuron activation behavior, latent feature coverage, decision-boundary exploration, etc. However, these metrics are typically released as independent research prototypes with substantially different installation and preprocessing requirements, execution workflows, and configuration mechanisms. These complications make them quite difficult to reproduce, compare, and adopt in research work and practical deployment alike. In this paper, we present the engineering details of ADEPT, a framework that integrates representative adequacy techniques, including neuron-coverage-based metrics, surprise adequacy, input distribution coverage, boundary coverage, and source- and model-level mutation score, under a consistent execution workflow. ADEPT provides a template-based metric interface with well-defined extension points for integrating new adequacy metrics. Furthermore, it provides YAML-based configuration management, preprocessing-cache reuse, and structured result reporting, making it easy to use in any research and development workflows. ADEPT is designed for researchers and practitioners who wish to reproduce and apply adequacy metrics without spending days or weeks implementing missing tooling or configuring disparate research prototypes. A demo video is available at https://aub.ie/ADEPT video.

Sources

Related papers