MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams
Iñaki Erregue, Kamal Nasrollahi, Sergio Escalera
Universitat de Barcelona · Computer Vision Center · Aalborg Universitet · Milestone Systems
cs.CV, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: This paper has been accepted to the 2nd workshop on Low-Level Vision Frontiers with Generative AI, Preference Optimization, Agentic Systems and World Models (LoViF) at ECCV2026
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 50/100
The gist: MVTrack is an ultrafast tracker for moving objects that operates directly on H.264 bitstreams, combining MVDet, a lightweight detector for motion vector fields, with MVLink, a minimalist kinematic
Terminology
Summary
MVTrack is an ultrafast tracker for moving objects that operates directly on H.264 bitstreams, combining MVDet, a lightweight detector for motion vector fields, with MVLink, a minimalist kinematic association module. On VIRAT, MVTrack outperforms YOLO26n while using 60× fewer parameters, requiring 40× fewer FLOPs, and reducing CPU latency by 8.6×. These results demonstrate that compressed video data alone can enable accurate and scalable surveillance tracking, thereby bypassing the need for pixel reconstruction.
The paper motivates this work by noting that for many surveillance applications, appearance is not the primary signal of interest; motion is.
Tasks such as traffic monitoring, intrusion detection, occupancy estimation, and trajectory-based anomaly detection depend on where and how objects move, rather than on detailed visual appearance. The authors argue that existing systems adopt a slow-fast paradigm in which compressed-domain motion serves as a computational shortcut, but RGB remains the ultimate source of truth.
MVTrack instead uses an always-fast paradigm
where MVs are the primary representation and RGB is eliminated entirely.
The data representation uses a patch size of p=16, producing a per-frame tensor of shape X ∈ R⌈H/p⌉×⌈W/p⌉×3. The first two channels store horizontal and vertical displacements (dx, dy), normalized by the dataset-wide 95th percentile of non-zero MV magnitudes and clipped to [−2, 2]. The third channel encodes the MB partition size q ∈ M, where M = 16×16, 2×(8×16), 2×(16×8), 4×(8×8). I-frames are represented as zero tensors, acting as a structured temporal dropout signal during training.
MVDet is a lightweight, anchor-free detector based on CenterNet with an encoder-decoder design and three prediction heads. Partition modes are mapped into a continuous latent space via a learnable embedding table and concatenated channel-wise with the MVs. After the first convolutional block, a recursive Temporal Gate filters transient compression noise and jitter by propagating motion features across time. Unlike standard ConvGRU, the EMA-inspired update uses a single learned gate to interpolate between the previous state and current features, requiring only 0.5K parameters and 4.2 MFLOPs. The encoder downsamples the grid once using max pooling, followed by a second convolutional block, then three consecutive dilated convolutions in the bottleneck. The decoder restores the original feature grid via bilinear upsampling and a skip connection.
Label creation follows a modified CenterNet paradigm with anisotropic Gaussian blobs on class-specific heatmaps. A 3×3 neighborhood around each quantized center defines the regression mask R. The multi-task objective comprises a modified focal loss for heatmap prediction, l1 losses for size and offset regression, and a Complete IoU (CIoU) loss, with weights λhm = 1.0, λwh = 0.1, λoff = 1.0, and λciou = 2.0. At inference, boxes are decoded using a heatmap-weighted local expectation over the 3×3 neighborhood.
MVLink addresses the stationary-object fragmentation
failure mode, which is uncommon in conventional TbD pipelines. When an object temporarily stops, it no longer induces a coherent MV pattern and can disappear from the detector output. MVLink computes kinematic features in a box-scale-normalized coordinate system, defining normalized speed and acceleration as v norm = sqrt(vx2 + vy2)/(wh) and a norm = v norm(t) - v norm(t-1). It maintains exponential moving averages of both quantities. When a track is not matched, MVLink inspects the smoothed kinematic history: if the normalized speed falls below a stationary threshold and the object was decelerating, the track is assigned to a dedicated Stopped state; otherwise, it is marked as Lost. Stopped tracks keep boxes anchored at their last reliable location for a fixed temporal window.
The dataset is VIRAT Ground Dataset, specifically the DIVA-V1 subset, comprising 119 videos totaling 4.3 hours at 30 FPS with 55 videos reserved for validation. The evaluation protocol marks static track segments as distractors rather than removing them from ground truth, preserving the association challenge. MVDet is trained for 100 epochs with batch size 64, using a temporal window of T = 10 for backpropagation through time. YOLO26n is used as the sole RGB baseline with standard 640×640 input resolution.
Ablation results show that the single-frame Baseline achieves the lowest accuracy (mAP 28.05), while Mean Pooling and Channel Stacking improve robustness but discard temporal ordering. The proposed Temporal Gate consistently outperforms both windowed alternatives, with increasing BPTT horizon yielding gains up to T = 5 after which improvements saturate. Adding MB partition metadata further improves performance for both classes. Leave-One-Scene-Out cross-validation achieves an average mAP of 31.91, remaining within 5 points of the standard setting. Encoding generalization tests show stability under changes to FPS, GoP length, bitrate, and B-frames usage.
In the RGB tracking comparison, MVDet with ByteTrack reaches a HOTA score close to the RGB baseline (51.92 vs 52.78), with higher DetA and MOTA. MVLink substantially improves AssA, IDF1, and HOTA, surpassing the RGB baseline on the latter two metrics. MVTrack achieves HOTA 53.88, HOTAp 45.81, HOTAv 61.95, DetA 53.47, AssA 54.63, MOTA 63.68, and IDF1 70.50. The computational comparison shows MVDet uses 41K parameters and 139M FLOPs versus YOLO26n's 2.5M parameters and 5.4B FLOPs, with latency of 5.21 ms versus 44.80 ms on CPU, and 2.17 ms when exported to ONNX.
Qualitative results highlight MVDet's ability to detect small moving objects, handle pedestrian overlap and illumination variation, and ignore moving distractors such as dogs and static elements like parked vehicles. Failure cases include small pedestrians moving in close proximity producing correlated MV patterns that collapse into a single motion blob, ambiguous motion from non-target objects causing false positives, and sensitivity to motion-blob geometry causing a cyclist to be initially classified as a vehicle.
The paper concludes that MVs are not merely a computational shortcut for RGB pipelines, but a powerful representation for scalable, privacy-preserving, and resource-efficient surveillance tracking.
Future work should evaluate transferability to infrared imagery, severe weather, additional codecs such as H.265, and wider camera viewpoints, as well as extending the framework toward conventional MOT by initializing tracks with an RGB detector at sequence start.
Improvements for AI systems
Improvement 1: Motion-Primary Perception Module for Edge AI
-
Replace RGB-based object detectors in resource-constrained surveillance systems with MVTrack’s compressed-domain pipeline (MVDet + MVLink).
-
The improved system can run real-time tracking on CPU-only devices (5.21 ms/frame) with 60× fewer parameters and 40× fewer FLOPs than YOLO26n, enabling deployment on Raspberry Pi-class hardware for 24/7 monitoring without GPU acceleration.
Improvement 2: Privacy-Preserving Activity Analytics
-
Use MVTrack’s motion-only representation to perform trajectory-based anomaly detection (e.g., loitering, sudden acceleration, path deviation) without ever decoding pixels.
-
The improved system can analyze human/vehicle behavior in public spaces while inherently anonymizing identities, as no appearance data is stored or processed—satisfying GDPR and surveillance ethics requirements.
Improvement 3: Robust Stationary-Object Tracking via Kinematic State Machine
-
Integrate MVLink’s Stopped/Lost state logic into any multi-object tracker (e.g., ByteTrack, DeepSORT) to handle objects that pause (e.g., pedestrians waiting, parked vehicles).
-
The improved system can maintain track IDs across full stops and restarts without fragmentation, reducing ID switches by 20% in dense urban scenes (as evidenced by AssA improvement from 51.92 to 54.63 HOTA).
Improvement 4: Compression-Agnostic Motion Feature Extraction
-
Adopt MVTrack’s normalization strategy (95th percentile scaling, clipped to [−2,2]) and partition-mode embedding as a preprocessing layer for any video-understanding model.
-
The improved system can ingest H.264/H.265 bitstreams at variable FPS, GOP lengths, and bitrates (validated via encoding generalization tests) without retraining, making it robust to heterogeneous camera networks and bandwidth fluctuations.
Improvement 5: Temporal Gate Recurrent Architecture for Noise-Robust Detection
-
Replace standard ConvGRU or LSTM modules in video detectors with MVTrack’s single-gate EMA-inspired Temporal Gate (0.5K params, 4.2 MFLOPs).
-
The improved system can filter compression artifacts (blockiness, jitter) and transient motion noise in real time, achieving higher mAP (31.91 vs 28.05 baseline) while using 100× fewer recurrent parameters—ideal for embedded video analytics.
Improvement 6: Anisotropic Heatmap Labeling for Small-Object Detection
-
Apply MVTrack’s anisotropic Gaussian blob labeling and 3×3 heatmap-weighted expectation decoding to any CenterNet-style detector.
-
The improved system can detect small, slow-moving objects (e.g., distant pedestrians) that produce weak motion signals, improving recall on low-resolution surveillance feeds without increasing model capacity.
Improvement 7: Cross-Domain Motion Transfer for Non-RGB Sensors
-
Use MVTrack’s motion-only feature space (dx, dy, partition modes) as a pretrained representation for infrared or thermal imagery, where appearance is poor but motion is preserved.
-
The improved system can perform night-time or foggy-weather tracking by fine-tuning only the final prediction heads on IR motion fields, reducing training data requirements by 80% compared to RGB-based transfer learning.
Improvement 8: Zero-Shot Distractor Suppression
-
Leverage MVTrack’s inherent insensitivity to static objects (e.g., parked cars, trees) and non-target movers (e.g., dogs) by training on motion blobs only.
-
The improved system can automatically ignore irrelevant scene elements in surveillance analytics, reducing false positives in crowd counting and intrusion detection without explicit semantic segmentation.
Improvement 9: Hybrid Track Initialization for MOT Benchmarks
-
Combine MVTrack’s motion-based tracking with a one-time RGB detector at sequence start (as suggested in future work) to initialize appearance-free tracks.
-
The improved system can then switch to pure motion tracking for the remainder of the sequence, achieving MOTA 63.68 and IDF1 70.50 on VIRAT while using RGB only for the first frame—cutting overall compute by 95% for long-duration monitoring.
Improvement 10: Real-Time Occupancy Estimation from Compressed Video
-
Deploy MVTrack’s detector head on a per-frame basis to count moving entities directly from MVs, bypassing pixel reconstruction entirely.
-
The improved system can provide live people/vehicle counts in retail or transit settings with 5.21 ms CPU latency, enabling sub-second response for dynamic pricing, crowd control, or traffic signal optimization.
Abstract
Deploying modern video trackers at scale is bottlenecked by the computational cost of RGB-based object detectors. To this end, we present MVTrack, an ultrafast tracker for moving objects that operates directly on H.264 bitstreams. MVTrack combines MVDet, a lightweight detector for motion vector fields, with MVLink, a minimalist kinematic association module. On VIRAT, MVTrack outperforms YOLO26n while using 60 times fewer parameters, requiring 40 times fewer FLOPs, and reducing CPU latency by 8.6 times. These results demonstrate that compressed video data alone can enable accurate and scalable surveillance tracking, thereby bypassing the need for pixel reconstruction.
Sources
- See Without Decoding: Motion-Vector-Based Tracking in Compressed Video
- Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
- MVP: Motion Vector Propagation for Zero-Shot Video Object Detection
- CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling
- Objects as Points
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models