Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision
| Source: arXiv AI
Tags: edge-AI, in-sensor-computing, FPGA, compression, computer-vision, hyperdimensional-computing
OASIS reduces in-sensor vision system energy by 2-4.5x through a lightweight encoder that compresses image representations before off-chip transmission, achieving 18,816x compression versus raw 8-bit pixels for visual wake-word detection on a Xilinx FPGA with under 1% accuracy loss.
Details
In-sensor computing processes image data near the CMOS sensor before transmission, reducing bandwidth and power. The challenge: sensors' integrated logic chips are tightly constrained in compute and memory, limiting how deep a neural network can run on-chip. OASIS addresses this with a hardware-algorithm co-design: a lightweight encoder trained end-to-end generates compact, task-relevant representations, which are transmitted instead of raw pixels. A decoder used only during training handles reconstruction and is never deployed on the constrained chip. OASIS supports two deployment paths. The first uses 4-bit quantization with Huffman coding to preserve spatial structure needed by classification and dense-prediction tasks. The second uses Sobol-based hyperdimensional computing to transform the encoder latent into a fixed-dimensional binary hypervector for associative-memory classification. For visual wake-word detection using SwinViT, mapping a 3x3x8 latent to a 64-dimensional hypervector achieves 18,816x compression versus raw 8-bit images with less than one percentage point of accuracy loss. System energy reduction of 2-4.5x is validated on an AMD Xilinx Zynq UltraScale+ FPGA with direct board-level power measurements and a 7nm ASIC projection, across visual wake-word classification, hand tracking, and eye tracking tasks.