FORGE: Forward-Only Test-Time Adaptation for Integer-Only Vision Models on Microcontrollers
cs.CV, cs.AR, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 16 pages, 5 figures, 9 tables. Published in Transactions on Machine Learning Research (2026). OpenReview: https://openreview.net/forum?id=A45I5p25dd. Code and checkpoints: https://github.com/Rehan000/forge-tta
Journal ref: Transactions on Machine Learning Research, 2026
Code: https://github.com/Rehan000/forge-tta
License: http://creativecommons.org/licenses/by/4.0/
The gist: Vision models deployed on microcontrollers (MCUs) are quantized to integer-only arithmetic and run in inference-only runtimes that do not carry the machinery backpropagation needs: the standard tool
Terminology
Abstract
Vision models deployed on microcontrollers (MCUs) are quantized to integer-only arithmetic and run in inference-only runtimes that do not carry the machinery backpropagation needs: the standard tool for adapting a model to the distribution shift (sensor noise, blur, lighting) it meets in the field. Existing forward-only test-time adaptation (TTA) methods either run only on server- or edge-GPU-class models (not true microcontroller integer execution), or require the batch-normalization (BN) layers that integer deployment fuses away. We present a forward-only TTA method that operates on deployed, BN-folded, integer-only convolutional networks. The key observation is that fusing BN into the preceding convolution, a mandatory step for integer inference, destroys the statistics that normalization-based adaptation relies on. We restore adaptation by re-normalizing each folded convolution's per-channel output to its clean training statistics, using only forward-pass estimates. The method (i) recovers most of gradient-based TENT's accuracy gain (+20.9 vs. +24.9 points) and matches forward-only BN adaptation, while being the only method that runs on a folded integer-only model; (ii) needs to adapt only 3 of 21 layers (selected without seeing the test corruptions) to recover 93% of the benefit; (iii) survives single-sample streaming with a batch-size-scaled momentum; and (iv) generalizes across three datasets (up to 200 classes) and two architectures. We validate bit-exact int8 convolution execution and deploy on an ESP32-S3, where, measured with a Nordic PPK2 power profiler, the forward-only adaptation (a lightweight fp32 recalibration around the int8 convolutions) costs only 8.3 mJ (6.8% of inference energy) and 21.9 ms on the deployed SIMD-optimized model: forward-only adaptation is cheap on a real microcontroller.
Sources
- Test-Time Model Adaptation for Quantized Neural Networks
- LeanTTA: A Backpropagation-Free and Stateless Approach to Quantized Test-Time Adaptation on Edge Devices
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs
- Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
- Evaluating Prediction-Time Batch Normalization for Robustness under Covariate Shift
- A White Paper on Neural Network Quantization
- Test-Time Model Adaptation with Only Forward Passes
- Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models