SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting
summary
The gist
SoccerTrack v2 introduces a novel public dataset designed to advance multi-object tracking (MOT), game state reconstruction (GSR), and ball action spotting (BAS) research in soccer analytics by
In short
SoccerTrack v2 is a new public dataset featuring 10 full-length 4K panoramic videos of university soccer matches. It provides rich annotations for multi-object tracking, game state reconstruction, and ball action spotting. This resource advances computer vision research by offering synchronized positional data and detailed event labels across multiple matches.
Key concepts
- Multi-Object Tracking (MOT)
- MOT is a technique used to follow multiple objects—like players—across video frames. In this dataset, it helps researchers track every player persistently throughout the entire match using unique track IDs and 2D pitch coordinates.
- Game State Reconstruction (GSR)
- GSR involves recreating the complete state of a game from video footage. The annotations provide crucial data points for every visible person, including their role, jersey number, and exact location on the pitch to reconstruct what is happening during play.
- Ball Action Spotting (BAS)
- BAS enables event-based analysis by labeling specific actions related to the ball. Annotations cover 12 distinct classes such as 'Pass,' 'Shot,' or 'Goal,' allowing researchers to analyze dynamic soccer events beyond simple player tracking.
Terminology used across episodes
This episode discusses
- SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting · Paper Radio
The paper
SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting · Read on arXiv
Atom Scott, Ikuma Uchida, Kento Kuroda, Yufi Kim, Keisuke Fujii
Nagoya University · University of Tsukuba
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting".
Jane: SoccerTrack v2 introduces a novel public dataset designed to advance multi-object tracking (MOT), game state reconstruction (GSR),
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re looking at a paper called "SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting." It sounds pretty technical, but the main thing is they’ve put together a dataset that combines high-resolution panoramic video with detailed game state and ball action labels.
Jane: That’s right, Tom. Essentially, they’re tackling the problem of getting a holistic view of a soccer match by giving researchers not just video clips, but all the positional data and event labels needed to understand what was happening across the entire field at once.
Lu: What makes this dataset unique is that it’s been built using BePro cameras for ten complete matches in one consistent environment <ref:2508.01802#pg0>. That means you get full pitch coverage and a lot of variety in how the game looks from different angles or under different conditions, which is pretty rare.
Meng: So, to put it simply, this dataset aims to move past just tracking individual players in short clips and instead let you reconstruct the actual tactical situation on the field frame by frame. It’s a big step toward understanding match dynamics.
Lalam: For me, it means we can train models that don't just see a player moving but actually know where they are on the pitch and what role they’re playing, which could help us build much smarter tactical AI for coaching or automated analysis.
The paper's summary: Tom: They summarize it by saying this is the first dataset to merge high-resolution panoramic video with rich per-frame GSR data and ball action spotting labels across ten complete matches in one consistent setting <ref:2508.01802#pg0>. That’s a big claim there, connecting all those pieces together.
Jane: Exactly. The summary points out that they’ve contributed 4K panoramic video data from BePro cameras covering entire matches with a complete field of view, along with GSR annotations like 2D pitch coordinates and player IDs via jersey numbers <ref:2508.01802#pg0>.
Lu: They also mention the BAS annotations for twelve action classes, including things like Pass, Drive, Shot, and Header <ref:2508.01802#pg0>. This means you can do both track players *and* spot specific ball actions simultaneously within the same video context.
Meng: The summary highlights that they address limitations in existing datasets by providing pitch-level annotations instead of just broadcast views or short clips that suffer from occlusions and incomplete visibility.
Lalam: It basically means we’re moving from just seeing *who* is where to understanding *how* the game is being played through every single movement and action captured in the video.
The paper's improvements: Tom: Now, looking at what they suggest as improvements, they focus on data integration for holistic analysis. They argue that combining tracking and event-based video analysis lets systems do both at the same time within a single match context.
Jane: That’s a key point. By having both the positional data from GSR and the event labels from BAS, you can create a system that tracks players while simultaneously detecting specific ball actions, which is much richer than either doing them separately.
Lu: They also focus on advanced game state reconstruction by using those 2D pitch coordinates, player IDs derived from jersey numbers, roles, and team affiliations to reconstruct the spatial and tactical state of a match frame-by-frame <ref:2508.01802#pg0>.
Meng: From an engineering standpoint, that positional information is crucial because it lets the AI understand the actual space on the pitch rather than just seeing a generic bounding box around a player. It gives it depth.
Lalam: And those player IDs linked to jersey numbers are super important for building persistent tracking systems; if you can keep that ID across frames, you build a much more stable model of how players move and interact throughout the whole match.
Conclusion: Tom: So, to wrap up, SoccerTrack v2 gives us a comprehensive resource combining high-resolution panoramic video with detailed per-frame positional data and event labels across multiple matches for foundational computer vision research. It’s a solid resource for anyone working on multi-object tracking, game state reconstruction, or ball action spotting.
Jane: Right. The implication is that researchers can now test more complex systems that need to understand both where players are and what actions are happening at the same time in a full match context, using this dataset as a new benchmark.
Lu: For the world of soccer analytics, this means we’re moving toward tools that can interpret entire match scenarios with much more accuracy than before because they have access to all those granular details.
Meng: It gives engineers the concrete data they need to build systems that don't just track a ball, but understand the entire flow of play based on precise positional and action information.
Lalam: Ultimately, this work supports building AI that can process complex visual information in real-world scenarios with high fidelity, which helps improve how we design these analytical tools for people who actually use them.
Tom: So there you have it for SoccerTrack v2, a dataset that really pushes the boundaries of what we can track and analyze in soccer videos. We’ll take a quick break before we look at what else is buzzing on arXiv.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck