Prompt, Probe, Train, or Annotate? Single-camera sports video understanding in amateur settings
cs.CV, cs.AI
Submitted: 2026-09-23
Updated: 2026-09-23
Terminology
Sources
- SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes
- MultiSports: A Multi-Person Video Dataset of Spatio-Temporally Localized Sports Actions
- A General Framework for Jersey Number Recognition in Sports Video
- Towards long-term player tracking with graph hierarchies and domain-specific features
- SoccerNet Game State Reconstruction: End-to-End Athlete Tracking and Identification on a Minimap
- SoccerNet 2025 Challenges Results
- SoccerNet 2026 Challenges Results
- Revisiting Feature Prediction for Learning Visual Representations from Video
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
- VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
- What Does Vision Tool-Use Reinforcement Learning Really Learn? Disentangling Tool-Induced and Intrinsic Effects for Crop-and-Zoom
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models