NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space

arXiv:2608.13293 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Eleftherios Mylonas, Angelos Kouprizas, Michael Birbas, Alexios Birbas

University of Patras

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 6 pages, 6 figures, accepted for presentation to the 39th IEEE International System-on-Chip Conference, Heidelberg, Germany, September 30 - October 2, 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper proposes a three-stage hardware-aware Neural Architecture Search (NAS) pipeline for edge AI deployment on Coarse-Grain Reconfigurable Array (CGRA)-based accelerators, and presents an

Terminology

Summary

This paper proposes a three-stage hardware-aware Neural Architecture Search (NAS) pipeline for edge AI deployment on Coarse-Grain Reconfigurable Array (CGRA)-based accelerators, and presents an empirical study of the effects of INT4 Post-Training Quantization (PTQ) on the NAS-Bench-201 Pareto space.

The three-stage pipeline consists of: "a hardware-agnostic Pareto rank surrogate frontend on NAS-Bench-201, a quantization bridge with Pareto-aware filtering and feedback control, and an evolutionary Domain Space Exploration (DSE) backend on CGRA4ML for optimal hardware mapping." Stage I is a hardware-agnostic frontend based on a redesign of HW-PR-NAS, using a Pareto rank surrogate trained on accuracy and FLOPs objectives, with Random Search (RS) and Multi-Objective Evolutionary Algorithm (MOEA) as search strategies, returning ten candidate architectures. Stage II is a quantization bridge that applies INT4 PTQ via Brevitas, performs Pareto re-ranking, and filters out dominated architectures, with a feedback loop to Stage I if no survivors remain. Stage III is a hardware-aware backend using CGRA4ML, with a QONNX-to-QKeras translation layer, that performs DSE over four configuration parameters (PE array rows, columns, weight SRAM depth, AXI width) using an evolutionary algorithm with a three-term normalised scalar fitness function that jointly minimises latency, PE idle ratio, and array area.

The empirical study characterises how INT4 PTQ perturbs the NAS-Bench-201 Pareto space. The quantization scheme was selected by testing 325 PTQ configurations, with PTQ Id 6 (4-bit activations, 4-bit weights, layerwise quantization, asymmetric activation quantization) chosen. The study found that Models at lower computational scales (7–47 MFLOPs) exhibit considerably higher sensitivity to INT4 quantization, with median accuracy drops reaching approximately 5% at the 7–15 MFLOPs range, while sensitivity diminishes with increasing complexity, attributed to the dominance of the quantization-robust nor conv 3x3 operation at higher scales.

The study reports that a 0% front survival rate confirms complete reorganisation of the efficient frontier under INT4, while the 21.73% dominance flip rate indicates that one in five FP32 dominance relationships breaks down. The ground-truth Kendall's Tau rank correlation between FP32 and INT4 is 0.6655, and Pareto rank sensitivity is 0.2404. Despite this full structural reorganisation, the FP32 zero-shot surrogate outperforms a dedicated INT4-trained surrogate in Pareto space coverage, with hypervolume ratios of 1.12 and 1.07 for RS and MOEA respectively, corresponding to relative improvements of 12.26% and 6.77%. The surrogate transferability analysis shows KT of 0.8352 on the original FP32 domain, 0.7219 when transferred zero-shot to INT4, and 0.8219 after fine-tuning on quantized data.

The DSE results show that three Pareto front architectures converge to the same optimal CGRA4ML configuration (16×66 PEs, 576 weight SRAM depth), with clock counts of 53,605, 118,690, and 74,995 cycles at 250 MHz, and PE utilization of 42.61%, 45.8%, and 41.4% respectively.

Improvements for AI systems

Improvements to AI Systems:

  1. Quantization-Aware Surrogate Training with Pareto-Rank Preservation
  • Improvement: Replace standard accuracy/FLOPs surrogates with a Pareto-rank-aware loss that explicitly penalizes rank flips under INT4 quantization. Use the paper’s finding (21.73% dominance flip rate, KT=0.6655) to train surrogates that predict quantization-robust Pareto dominance rather than raw accuracy.

  • What the improved system can do: Generate neural architectures that maintain their Pareto-optimality after INT4 PTQ, reducing the need for post-hoc re-ranking and hardware re-mapping. It can also predict which FP32-optimal architectures will remain optimal under quantization, avoiding wasted search iterations.

  1. Adaptive Quantization Sensitivity Prediction
  • Improvement: Integrate a sensitivity predictor (based on the paper’s finding that 7–47 MFLOPs models suffer 5% accuracy drops) into the NAS search loop. This predictor can flag high-risk architectures early, triggering either (a) a switch to INT8 for those sub-networks or (b) a bias toward quantization-robust operations (e.g., nor conv 3x3) at low computational scales.

  • What the improved system can do: Automatically adjust bit-width per layer or per architecture based on predicted sensitivity, achieving higher accuracy under INT4 constraints without manual tuning. It can also guide the search toward architectures that are inherently quantization-friendly, improving deployment success rates on edge hardware.

  1. Feedback-Controlled Quantization Bridge with Surrogate Fine-Tuning
  • Improvement: Use the paper’s finding that fine-tuning the surrogate on quantized data (KT improves from 0.7219 to 0.8219) to implement a dynamic feedback loop: if the front survival rate is 0% (as observed), the system automatically fine-tunes the surrogate on INT4 data and re-runs Stage I search, rather than simply filtering. This turns a failure case into an adaptive learning opportunity.

  • What the improved system can do: Self-correct its search strategy when quantization disrupts the Pareto front, avoiding dead-ends and improving the chance of finding viable architectures. It can also learn to anticipate quantization effects over time, reducing the number of feedback iterations needed.

  1. Hardware-Mapping Co-Search with Quantization-Aware Fitness
  • Improvement: Extend the DSE backend to jointly optimize architecture and CGRA mapping using a fitness function that includes quantization-induced latency/area penalties (e.g., based on the observed PE idle ratios of 41–46% and optimal configs like 16×66 PEs). This allows the system to trade off accuracy, quantization robustness, and hardware efficiency in a single pass.

  • What the improved system can do: Produce end-to-end optimized solutions where the chosen architecture and hardware configuration are jointly robust to INT4 quantization, minimizing latency and area while maintaining accuracy. It can also predict which hardware configs are most forgiving to quantization-induced accuracy loss.

  1. Zero-Shot Surrogate Transfer with Rank-Correlation Awareness
  • Improvement: Leverage the finding that the FP32 zero-shot surrogate outperforms an INT4-trained surrogate (hypervolume ratio 1.12 for RS) by designing a dual-surrogate ensemble: one trained on FP32, one on INT4, with a meta-learner that weights them based on predicted rank correlation (KT=0.6655). This avoids the cost of full INT4 training while improving Pareto coverage.

  • What the improved system can do: Achieve better Pareto front coverage under quantization with minimal extra training data, by intelligently combining the strengths of both surrogates. It can also flag when the FP32 surrogate is likely to fail (low KT) and switch to the INT4-trained one.

  1. Quantization-Robust Architecture Search via Operation Bias
  • Improvement: Use the paper’s observation that quantization sensitivity diminishes with complexity due to nor conv 3x3 dominance to implement a complexity-aware operation bias in the NAS search space. The system can explicitly favor nor conv 3x3 at higher FLOPs and mix in other operations at lower FLOPs, based on a learned sensitivity map.

  • What the improved system can do: Generate architectures that are naturally more robust to INT4 PTQ across a range of computational budgets, reducing accuracy drops without requiring post-hoc quantization-aware retraining.

Abstract

Edge AI deployment demands neural architectures that are simultaneously accurate, computationally efficient, and hardware-deployable - a challenge addressed by hardware-aware Neural Architecture Search (NAS). While recent works incorporate quantization directly into the NAS loop, these approaches expand search complexity and tightly couple architecture and quantization design. The simpler post-search quantization strategy has received little analytical attention: the effects of Post-Training Quantization (PTQ) on the NAS-discovered Pareto structure remain uncharacterised, and no framework combines quantized architecture mapping onto reconfigurable accelerators with automated hardware exploration. This paper addresses both gaps. First, a three-stage pipeline is proposed: a hardware-agnostic Pareto rank surrogate frontend on NAS-Bench-201, a quantization bridge with Pareto-aware filtering and feedback control, and an evolutionary Domain Space Exploration (DSE) backend on CGRA4ML for optimal hardware mapping. Second, an empirical study characterises how INT4 PTQ perturbs the NAS-Bench-201 Pareto space through formal stability metrics on ground-truth data for all 15,625 architectures, and demonstrates that an FP32 zero-shot surrogate outperforms a dedicated INT4-trained surrogate in Pareto space coverage across two standard search strategies.

Sources

Related papers