CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

arXiv:2608.10506 · cs.AR, cs.LG, cs.PF · Submitted 2026-08-11 · Read on arXiv

Linh Nguyen, Zhixin Pan

Florida State University

cs.AR, cs.LG, cs.PF

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening Summary This paper presents CARB, a characterization-guided framework for predicting CNN inference

Terminology

Summary

CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening

Summary

This paper presents CARB, a characterization-guided framework for predicting CNN inference costs (energy, latency, and peak memory) and enabling efficient deployment screening on GPU platforms. The work is motivated by the observation that existing estimation approaches—relying on FLOPs, latency measurements, or single-device profiling as energy proxies—fail to capture the non-linear interactions between architectural design and hardware load.

The authors conduct a large-scale empirical workload characterization study of 13,419 CNN configurations on two GPU platforms (NVIDIA RTX 5090 and RTX 3080) under controlled GPU telemetry. The study reveals several key findings:

  1. Energy, latency, and memory are not interchangeable. Under high computational demand, energy scales 35.4× while latency scales only 11.2× across batch-size regimes—a 3× divergence. FP16 reduces memory by 1.7× but energy by only 1.1×. No single architectural knob determines cost in isolation.

  2. Cross-GPU transferability is target-dependent. Energy and latency follow non-unit transfer slopes (1.61× and 2.09× respectively) and require platform-specific models, while memory transfers with slope ≈1 across the two tested platforms.

  3. These insights enable efficient deployment screening. CARB, a cascade ensemble guided by these findings, supports a two-stage screening workflow that reduces a 3072-configuration space to a Pareto-prioritized shortlist before any hardware profiling, with accuracy (R2 ≈ 0.99) sufficient for reliable budget classification.

Data Collection Methodology

The dataset is constructed by benchmarking CNN configurations under controlled hardware conditions. GPU clocks are locked (graphics clock: 1350 MHz; memory clock: 5001 MHz) to reduce DVFS-induced variance, and cuDNN benchmark mode is disabled. Between configurations, Python garbage collection, CUDA cache clearing, device synchronization, and a 1-second cooling interval are applied. All experiments use PyTorch v.2.8.0.

The search space uses a ResNet-style architecture covering both basic-block (depths 8–152) and bottleneck-block (depths 50–200) variants, with width multipliers from 0.1 to 4.0, input resolutions from 32 to 128, batch sizes from 1 to 256, and precision formats FP32 and FP16.

For each configuration, latency is the mean inference time per batch over 100 runs after 10 warmup passes. Energy is recorded via NVML cumulative energy counter differences. Peak memory is measured via max memory allocated. GPU utilization is sampled using NVML over a 2-second window during a training-mode forward-backward pass.

Key Characterization Findings

Energy Characterization: Depth and batch size are the dominant energy drivers within the basic family, spanning 13.8× and 15.9× respectively. The bottleneck family shows a narrower depth range (3.4×) but batch size remains dominant at 11.3×. Width multiplier shows asymmetry: within the basic family, energy scales 7.3× while memory scales 14.8×; within the bottleneck family, energy scales 4.7× while memory scales 14.8×. The bottleneck family draws 1.8× more energy than basic on average.

Energy–Latency Divergence: Among configurations in the high-SM-utilization, large-batch tier, energy scales 35.4× relative to the low-SM, small-batch baseline, while latency scales only 11.2× under the same conditions—a 3× divergence. The High-SM/Low-SM energy ratio grows monotonically from 3.29× at small batches to 35.4× at large batches.

Energy Efficiency: Precision causes up to 1.37× variation in J/GFLOP for the same architecture. Batch size has an even stronger effect: J/GFLOP falls from 14.81 at bs 1–4 to 2.52 at bs 17–64—a 5.9× improvement—before plateauing at 2.85 for bs 65–256.

Cross-GPU Generalizability: Energy follows the fit y = 1.61x + 52.8 (slope > 1), latency follows y = 2.09x − 2.7, while memory follows y = 0.92x + 6.7 (slope ≈ 1). This confirms that energy and latency require platform-specific models, while memory transfers well across the two tested platforms.

CARB Prediction Framework

CARB (Context-Aware Regime-corrected Blended ensemble) jointly estimates peak memory, energy, and latency. It is built around three principles: physical coupling between targets, hardware-load interaction features, and regime-specific residual correction.

Feature Engineering: Multiplicative interaction features are constructed from architectural parameters and live GPU telemetry, including batch x sm = batch size × avg sm util, flops x sm = log(1 + flops) × avg sm util, act x mem = log(1+total act MB) × avg mem util, and util product = avg sm util × avg mem util. Heavy-tailed columns are log-transformed, and categorical variables are one-hot encoded.

Specialist Ensemble with Cascade Prediction: The core is a blended ensemble of three structurally diverse base learners per target: XGBoost (800 trees, depth 9), LightGBM (800 trees, 127 leaves), and ExtraTrees (400 trees). The three targets are predicted via cascade:

(1) ŷ mem = f mem(x)

(2) ŷ energy = f energy(x, ŷ mem)

(3) ŷ latency = f latency(x, ŷ mem, ŷ energy)

Validation-calibrated blend weights are found by grid search: Energy is dominated by LightGBM (weight 0.5), latency by ExtraTrees (weight 0.5), and peak memory weights are approximately balanced.

Residual Correctability Analysis: For the high-stress regime (avg SM util > 65%), Spearman correlations between residuals and hardware utilization are: peak memory: ρ = 0.53, energy: ρ = 0.004, latency: ρ = 0.017. This indicates that only memory residuals contain learnable structure in the high-load regime. Instead, residual standard deviation analysis identifies the low-batch regime (batch size ≤ 8) as the common axis of elevated, structured error across all three targets.

Regime-Specific Residual Correctors: A LightGBM corrector is trained per target on the specialist's residuals, restricted to the low-batch regime. The corrected prediction is: ŷ corrected = ŷ specialist + λ·r̂, where λ = 0.9 is a damping factor.

Results and Baselines

CARB achieves R2 ≈ 0.99 across all three targets:

  • Peak Memory: R2 = 0.997, MAE = 6.93 MB

  • Energy: R2 = 0.993, MAE = 25.98 J

  • Latency: R2 = 0.992, MAE = 1.58 ms

Compared to baselines:

  • FLOPs-only linear regression achieves R2 < 0.38 across all targets

  • Latency-as-energy proxy achieves high R2 (0.993) but MAE of 71.37 J—CARB reduces energy MAE by 63.6%

Leave-one-batch-tier-out evaluation (training on batch size > 8, testing on batch size ≤ 8) yields R2 = 0.956, 0.991, and 0.968 for memory, energy, and latency respectively.

Feature Importance Insights

  • Peak memory is dominated by raw architectural size features (max activation MB, param size MB, flops per param log, compute intensity, total activation MB, flops)

  • Energy is led by raw compute and activation volume (flops, total activation MB) followed by flops per param log and compute intensity

  • For latency, pred energy J (the upstream cascade prediction) ranks first, ahead of every raw architectural feature

  • Three features appear in the top six for all three targets: flops, total activation MB, and flops per param log

Runtime Feature Ablation

Removing all telemetry features reduces R2 by at most 0.0012 in the overall regime across all three targets. The architecture-only model (M3) achieves R2 of 0.9963, 0.9892, and 0.9904 for memory, energy, and latency respectively—within noise of the full-telemetry model. This confirms that hardware telemetry is not required for reliable prediction of any of the three cost targets.

Two-Stage Deployment Screening

The framework uses a two-stage screening workflow:

Stage 1 (Ranking-based screening): Uses a single GPU's CARB model (RTX 5090) to rank all candidates and narrow the pool, exploiting rank-preservation across GPUs. For the example scenario, all 3072 candidates are scored and ranked by predicted energy; the top-100 configurations (shortlist energy range 38.2–42.6 J on the RTX 5090) are retained. Spearman rank correlation between the two GPU models' energy rankings is ρ = 0.95.

Stage 2 (GPU-specific prediction): Re-scores only the shortlisted candidates with the target GPU's model (RTX 3080), applying device-specific energy and latency constraints. For the example deployment budget of 75 J energy and 500 MB peak memory, 100% of shortlisted configurations pass the budget. After Pareto extraction, 3072 candidates are reduced to seven priority configurations—a 99.8% reduction in profiling load.

Screening Validation

On real held-out test feature vectors, all matched configurations are correctly classified against the deployment budget (3/3 correct budget decisions), with a mean absolute error of 2.49 J and RMSE of 2.70 J. Across the broader test set of 1,828 RTX 3080 configurations, CARB achieves 95.8% budget classification accuracy, with 2.1% false accepts and 2.1% false rejects.

Conclusion

The paper demonstrates that common deployment proxies such as FLOPs or latency are insufficient for reliable energy-aware optimization, particularly under high computational demand where energy and latency decouple substantially. Memory behavior generalizes across the two tested platforms more readily than energy or latency, motivating platform-specific modeling for accurate deployment estimation. CARB achieves R2 ≈ 0.99 across all three targets while enabling rapid Pareto-guided pre-deployment filtering through a two-stage screening workflow, demonstrating the importance of direct energy modeling for practical resource-aware CNN deployment.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Energy-Aware Model Selection and Deployment Optimization
  • AI systems can now predict energy consumption of CNN models with R2 ≈ 0.99, not just latency or FLOPs.

  • The system can automatically rank and filter thousands of model configurations (e.g., 3072 → 7) based on energy budgets before any hardware profiling, reducing deployment search time by 99.8%.

  • It can enforce hard energy constraints (e.g., ≤75 J per inference) with 95.8% classification accuracy, minimizing false accepts/rejects.

  1. Cross-Platform Transferable Memory Prediction
  • The system can predict peak memory usage on new GPU platforms using a model trained on one GPU (slope ≈ 0.92), enabling zero-shot memory estimation for deployment on unseen hardware.

  • It avoids unnecessary re-profiling for memory constraints, saving time and compute resources.

  1. Regime-Aware Correction for Small-Batch Inference
  • For batch sizes ≤ 8 (common in edge/real-time inference), the system applies a specialized residual corrector, improving prediction accuracy where generic models fail.

  • This enables reliable cost estimation for latency-sensitive, low-throughput deployment scenarios (e.g., autonomous driving, robotics).

  1. Cascade Prediction for Joint Cost Estimation
  • The system predicts peak memory first, then uses that prediction to improve energy and latency estimates, capturing physical coupling (e.g., memory-bound energy costs).

  • This yields more consistent multi-objective optimization, avoiding conflicting predictions that could mislead Pareto-based trade-off analysis.

  1. Hardware-Telemetry-Free Cost Estimation
  • The system can predict all three costs (energy, latency, memory) using only architectural features (e.g., FLOPs, activation size, parameter count) with R2 > 0.99, removing the need for live GPU telemetry.

  • This enables cost prediction during model design phases (before hardware is available) and for cloud/edge platforms where telemetry is inaccessible.

  1. Non-Linear Interaction Modeling for Architectural Search
  • The system captures non-linear interactions between batch size, SM utilization, and architecture (e.g., energy scales 35.4× vs. latency 11.2× under high load), enabling more accurate neural architecture search (NAS) that optimizes for energy efficiency rather than just speed.

  • It can identify energy-optimal configurations that would be missed by FLOPs- or latency-based proxies.

  1. Platform-Specific Energy and Latency Calibration
  • The system automatically applies platform-specific transfer slopes (energy: 1.61×, latency: 2.09×) when moving between GPUs, allowing accurate re-ranking of configurations for target hardware without full re-benchmarking.

  • This reduces the cost of multi-platform deployment by 50% (only shortlist needs re-scoring).

  1. Budget-Aware Pareto Front Extraction
  • The system can generate a Pareto-optimal set of configurations under multiple constraints (e.g., energy ≤ 75 J, memory ≤ 500 MB) in a single pass, enabling rapid design-space exploration for green AI or edge deployment.

  • It reduces the number of physical experiments needed from thousands to a handful (e.g., 7 priority configs).

  1. Improved Energy Efficiency in AI Training and Inference Pipelines
  • By predicting J/GFLOP variations (up to 5.9× across batch sizes), the system can recommend optimal batch sizes and precision formats (FP16 vs. FP32) to minimize energy per unit of compute, reducing data center power costs and carbon footprint.

  • It can dynamically adjust inference batching to operate in the most energy-efficient regime (e.g., batch 17–64) when latency constraints allow.

  1. Robustness to Hardware Variance
  • The system’s controlled benchmarking methodology (locked clocks, cuDNN disabled) and residual correction make predictions robust to DVFS-induced noise, enabling reliable cost estimates in production environments with fluctuating power states.

Sources

Related papers