Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
cs.RO, cs.AI, cs.CV, cs.LG
Submitted: 2025-08-15
Updated: 2026-09-16
Comments: \c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Journal ref: IEEE Robotics and Automation Practice (2026)
Code: https://github.com/nasa-jpl/visual-perception-engine
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving