Minimalist Visual Inertial Odometry

arXiv:2605.19990 · cs.RO, cs.CV, cs.LG · Submitted 2026-05-19 · Read on arXiv

cs.RO, cs.CV, cs.LG

Submitted: 2026-05-19

Updated: 2026-08-26

Comments: This work has been submitted to the IEEE for possible publication

Code: https://github.com/pastifra/four-pixel-vio

License: http://creativecommons.org/licenses/by/4.0/

The gist: Visual-Inertial Odometry (VIO), which is critical to mobile robot navigation, uses cameras with a large number of pixels.

Terminology

Abstract

Visual-Inertial Odometry (VIO), which is critical to mobile robot navigation, uses cameras with a large number of pixels. Capturing and processing camera images requires significant resources. This work presents a minimalist approach to planar odometry, showing that just four visual sensors (pixels) and an IMU provide robust motion estimation for differential-drive robots. Our key insight is that four downward-facing photodiodes that sense the world through optical Gabor masks produce signals that encode speed. Based on this, we jointly optimize the mask parameters alongside a Temporal Convolutional Network (TCN) using a physically-grounded simulator. The resulting model decodes speed from the four photodiode measurements. Pairing these estimates with an IMU's angular speed yields a continuous planar trajectory. We validate our approach with a prototype sensor mounted on a differential drive robot. Across diverse indoor and outdoor terrains, our system closely tracks the reference trajectories without any real-world fine-tuning. Our work shows that minimalist sensing enables efficient and accurate planar odometry.

Related papers