EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
cs.CV
Submitted: 2026-09-30
Updated: 2026-10-02
Code: https://github.com/Ropedia/EgoTools
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
- EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
- EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
- Grounded Question-Answering in Long Egocentric Videos
- Learning Task-Oriented Grasping for Tool Manipulation from Simulated Self-Supervision
- LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
- Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
- Ego4D: Around the World in 3,000 Hours of Egocentric Video
- ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
- EgoTaskQA: Understanding Human Tasks in Egocentric Videos
- MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
- LLaVA-OneVision: Easy Visual Task Transfer
- SAW-Bench: Learning Situated Awareness in the Real World
- Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
- LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction
- EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models