Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision

arXiv:2609.13947 · cs.LG, cs.AI, cs.AR · Submitted 2026-09-12 · Read on arXiv

cs.LG, cs.AI, cs.AR

Submitted: 2026-09-12

Updated: 2026-09-12

Comments: Under submission at IEEE Transactions on Emerging Topics in Computing

License: http://creativecommons.org/licenses/by/4.0/

The gist: In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor.

Terminology

Abstract

In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the logic chip integrated with a CMOS image sensor (CIS) is tightly constrained in compute and memory, limiting conventional deep neural network partitioning. We present OASIS, a distributed in-sensor vision framework that uses a lightweight encoder to generate compact, task-relevant representations before off-chip transmission. The encoder is trained end-to-end using task, entropy, and reconstruction objectives, while the decoder is used only during training. OASIS supports two complementary deployment paths. The first applies 4-bit quantization and Huffman coding while preserving the spatial structure required by classification and dense-prediction tasks. The second uses Sobol-based hyperdimensional computing (HDC) to transform the encoder latent into a fixed-dimensional binary hypervector for associative-memory classification. For the SwinViT-based VWW model, mapping a 3 times3 times8 latent to a 64-dimensional hypervector provides an additional 1.77 times communication reduction with less than one percentage point of accuracy loss relative to the 128-dimensional configuration, yielding an overall 18, 816 times reduction compared with raw 8-bit image transmission. We implement the digital near-sensor pipeline on an AMD Xilinx Zynq UltraScale+ FPGA and characterize it using direct board-level power measurements and Vivado post-implementation analysis, together with circuit-simulated CIS models and a 7-nm ASIC projection. Across visual wake-word classification, hand tracking, and eye tracking, OASIS reduces total system energy by approximately 2 times - 4.5 times while maintaining competitive accuracy, demonstrating a practical hardware-algorithm co-design path for communication-efficient in-sensor vision.

Related papers