GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
cs.CV, cs.AI
Submitted: 2026-07-15
Updated: 2026-08-28
Terminology
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- SAM 3: Segment Anything with Concepts
- TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification
- iPay: Integrated Payment Action Recognition via Multimodal Networks and Adaptive Spatial Prior Learning
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- A Deep Neural Network Approach to Fare Evasion
- ActionCLIP: A New Paradigm for Video Action Recognition
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models