SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting

arXiv:2508.01802 · cs.CV · Submitted 2025-08-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting".

Jane: SoccerTrack v2 introduces a novel public dataset designed to advance multi-object tracking (MOT), game state reconstruction (GSR),

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re looking at a paper called "SoccerTrack v2: A Full-Pitch Panoramic Video Dataset for Game State Reconstruction and Ball Action Spotting." It sounds pretty technical, but the main thing is they’ve put together a dataset that combines high-resolution panoramic video with detailed game state and ball action labels.

Jane: That’s right, Tom. Essentially, they’re tackling the problem of getting a holistic view of a soccer match by giving researchers not just video clips, but all the positional data and event labels needed to understand what was happening across the entire field at once.

Lu: What makes this dataset unique is that it’s been built using BePro cameras for ten complete matches in one consistent environment <ref:2508.01802#pg0>. That means you get full pitch coverage and a lot of variety in how the game looks from different angles or under different conditions, which is pretty rare.

Meng: So, to put it simply, this dataset aims to move past just tracking individual players in short clips and instead let you reconstruct the actual tactical situation on the field frame by frame. It’s a big step toward understanding match dynamics.

Lalam: For me, it means we can train models that don't just see a player moving but actually know where they are on the pitch and what role they’re playing, which could help us build much smarter tactical AI for coaching or automated analysis.

The paper's summary: Tom: They summarize it by saying this is the first dataset to merge high-resolution panoramic video with rich per-frame GSR data and ball action spotting labels across ten complete matches in one consistent setting <ref:2508.01802#pg0>. That’s a big claim there, connecting all those pieces together.

Jane: Exactly. The summary points out that they’ve contributed 4K panoramic video data from BePro cameras covering entire matches with a complete field of view, along with GSR annotations like 2D pitch coordinates and player IDs via jersey numbers <ref:2508.01802#pg0>.

Lu: They also mention the BAS annotations for twelve action classes, including things like Pass, Drive, Shot, and Header <ref:2508.01802#pg0>. This means you can do both track players *and* spot specific ball actions simultaneously within the same video context.

Meng: The summary highlights that they address limitations in existing datasets by providing pitch-level annotations instead of just broadcast views or short clips that suffer from occlusions and incomplete visibility.

Lalam: It basically means we’re moving from just seeing *who* is where to understanding *how* the game is being played through every single movement and action captured in the video.

The paper's improvements: Tom: Now, looking at what they suggest as improvements, they focus on data integration for holistic analysis. They argue that combining tracking and event-based video analysis lets systems do both at the same time within a single match context.

Jane: That’s a key point. By having both the positional data from GSR and the event labels from BAS, you can create a system that tracks players while simultaneously detecting specific ball actions, which is much richer than either doing them separately.

Lu: They also focus on advanced game state reconstruction by using those 2D pitch coordinates, player IDs derived from jersey numbers, roles, and team affiliations to reconstruct the spatial and tactical state of a match frame-by-frame <ref:2508.01802#pg0>.

Meng: From an engineering standpoint, that positional information is crucial because it lets the AI understand the actual space on the pitch rather than just seeing a generic bounding box around a player. It gives it depth.

Lalam: And those player IDs linked to jersey numbers are super important for building persistent tracking systems; if you can keep that ID across frames, you build a much more stable model of how players move and interact throughout the whole match.

Conclusion: Tom: So, to wrap up, SoccerTrack v2 gives us a comprehensive resource combining high-resolution panoramic video with detailed per-frame positional data and event labels across multiple matches for foundational computer vision research. It’s a solid resource for anyone working on multi-object tracking, game state reconstruction, or ball action spotting.

Jane: Right. The implication is that researchers can now test more complex systems that need to understand both where players are and what actions are happening at the same time in a full match context, using this dataset as a new benchmark.

Lu: For the world of soccer analytics, this means we’re moving toward tools that can interpret entire match scenarios with much more accuracy than before because they have access to all those granular details.

Meng: It gives engineers the concrete data they need to build systems that don't just track a ball, but understand the entire flow of play based on precise positional and action information.

Lalam: Ultimately, this work supports building AI that can process complex visual information in real-world scenarios with high fidelity, which helps improve how we design these analytical tools for people who actually use them.

Tom: So there you have it for SoccerTrack v2, a dataset that really pushes the boundaries of what we can track and analyze in soccer videos. We’ll take a quick break before we look at what else is buzzing on arXiv.

Atom Scott, Ikuma Uchida, Kento Kuroda, Yufi Kim, Keisuke Fujii

Nagoya University · University of Tsukuba

cs.CV

Submitted: 2025-08-03

Updated: 2026-10-05

Comments: 39 pages. Extended version with game state reconstruction and ball action spotting baselines; describes dataset release v1.2. Dataset and code: https://github.com/AtomScott/SoccerTrack-v2 and https://huggingface.co/datasets/atomscott/soccertrack-v2

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: SoccerTrack v2 introduces a novel public dataset designed to advance multi-object tracking (MOT), game state reconstruction (GSR), and ball action spotting (BAS) research in soccer analytics by

Key concepts

Multi-Object Tracking (MOT)
MOT is a technique used to follow multiple objects—like players—across video frames. In this dataset, it helps researchers track every player persistently throughout the entire match using unique track IDs and 2D pitch coordinates.
Game State Reconstruction (GSR)
GSR involves recreating the complete state of a game from video footage. The annotations provide crucial data points for every visible person, including their role, jersey number, and exact location on the pitch to reconstruct what is happening during play.
Ball Action Spotting (BAS)
BAS enables event-based analysis by labeling specific actions related to the ball. Annotations cover 12 distinct classes such as 'Pass,' 'Shot,' or 'Goal,' allowing researchers to analyze dynamic soccer events beyond simple player tracking.

Terminology

Summary

SoccerTrack v2 introduces a novel public dataset designed to advance multi-object tracking (MOT), game state reconstruction (GSR), and ball action spotting (BAS) research in soccer analytics by providing 10 full-length panoramic videos with synchronized positional data, detailed GSR annotations, and BAS labels for 12 action classes. This resource addresses limitations in existing datasets that rely on broadcast views or lack comprehensive player identification across multiple matches.

The gist

SoccerTrack v2 is the first dataset to combine high-resolution panoramic video, rich per-frame GSR data, and BAS across 10 complete matches in a consistent environment<ref:2508.01802#pg2>.

Dataset Scope and Content

The dataset comprises 10 amateur matches from university-level teams, totaling approximately 900 minutes of footage<ref:2508.01802#pg4> (a). The videos are in MP4 format at 4K resolution, capturing diverse weather conditions, locations, and gameplay styles<ref:2508.01802#pg2>. Collection involved fixed panoramic camera setups; two matches used a BePro Cerberus system, while the remaining eight utilized BePro’s standard 3-camera panoramic stitching system to ensure complete pitch coverage at 4K resolution. The data includes video files and corresponding annotation files in a JSON-based format<ref:2508.01802#pg4>.

Game State Reconstruction (GSR) Annotations

The GSR annotations are designed to support both multi-object tracking (MOT) and game state reconstruction (GSR)<ref:2508.01802#pg4>. For every frame, the GSR annotations include specific data points for every visible player, goalkeeper, and referee<ref:2508.01802#pg4>:

2D pitch coordinates (in meters)

Unique track ID (persistent throughout the match)

Role (player, goalkeeper, referee, other)

The annotations also provide player identification details such as:

• Jersey number (0–99 if visible; null otherwise)<ref:2508.01802#pg4>.

Ball Action Spotting (BAS) Annotations

BAS annotations enable event-based video analysis beyond tracking<ref:2508.01802#pg4>. These annotations capture key ball-related events derived from BePro’s proprietary event logs<ref:2508.01802#pg4> and include:

• Global timestamp (aligned to video timeline)<ref:2508.01802#pg4>

• Action class (Pass, Drive, Header, High Pass, Out, Cross, Throw In, Shot, Ball Player Block, Player Successful Tackle, Free Kick, Goal)<ref:2508.01802#pg4>.

Annotation Strategy and Access

The annotation process initially planned for full bounding box annotations across all 1.62 million frames; however, due to the extreme manual effort required (5000 hours), comprehensive bounding boxes are not included in the main dataset release<ref:2508.01802#pg4>. Instead, a curated subset with bounding boxes will be released as part of the SoccerTrack Challenge (MMSports 2025)<ref:2508.01802#pg4>. The public access for the full package (videos, annotations, baseline code) will be provided via a GitHub repository and Hugging Face Spaces<ref:2508.01802#pg4> at https://huggingface.co/datasets/atomscott/soccertrack- v2 and https://github.com/AtomScott/SoccerTrack-v2<ref:2508.01802#pg4>.

Ethical and Contextual Considerations

All recordings were carried out under informed written consent agreements approved by the ethical committee at the University of Tsukuba. To maintain anonymity in public releases, Player identities are not included in the public dataset; however, jersey numbers are provided for player reidentification and tracking purposes. This approach balances providing necessary identification data for research with adhering to ethical standards<ref:2508.01802#pg4>.

SoccerTrack v2 provides a comprehensive resource that integrates high-resolution panoramic video, detailed per-frame positional data, and event labels across multiple matches for foundational computer vision research and practical applications. The work was financially supported by JST SPRING Grant Number JPMJSP2125 and JSPS KAKENHI Grant Number 23H03282.

REFERENCES

[1] Anthony Cioppa, Silvio Giancola, Adrien Deliege, Le Kang, Xin Zhou, Zhiyu Cheng, Bernard Ghanem, and Marc Van Droogenbroeck. Soccernet-tracking: Multiple object tracking dataset and benchmark in soccer videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 3491–3502, 2022. 1, 2

[2] Yutao Cui, Chenkai Zeng, Xiaoyu Zhao, Yichun Yang, Gangshan Wu, and Limin Wang. SportsMOT: A large multi-object tracking dataset in multiple sports scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9921–9931, 2023. 1, 2

[3] Adrien Deliege, Anthony Cioppa, Silvio Giancola, Meisam J Seikavandi, Jacob V Dueholm, Kamal Nasrollahi, Bernard Ghanem, Thomas B Moeslund, and Marc Van Droogenbroeck. Soccernet-v2: A dataset and benchmarks for holistic understanding of broadcast soccer videos. In 7th International Workshop on Computer Vision in Sports (CVsports) at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR’ 21), pages 4508–4519, 2021. 1, 2

[4] Silvio Giancola, Mohieddine Amine, Tarek Dghaily, and Bernard Ghanem. Soccernet: A scalable dataset for action spotting in soccer videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 1711–1721, 2018. 1, 2

[5] Atom Scott, Ikuma Uchida, Masaki Onishi, Yoshinari Kameda, Kazuhiro Fukui, and Keisuke Fujii. Soccertrack: A dataset and tracking algorithm for soccer with fish-eye and drone videos. In 8th International Workshop on Computer Vision in Sports (CVsports) at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR’ 22), pages 3569–3579, 2022. 1

[6] Atom Scott, Ikuma Uchida, Ning Ding, Rikuhei Umemoto, Rory Bunker, Ren Kobayashi, Takeshi Koyama, Masaki Onishi, Yoshinari Kameda, and Keisuke Fujii. Teamtrack: A dataset for multi-sport multi-object tracking in full-pitch videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3357–3366, 2024. 1, 2

[7] Vladimir Somers, Victor Joos, Anthony Cioppa, Silvio Giancola, Seyed Abolfazl Ghasemzadeh, Floriane Magera, Baptiste Standaert, Amir M Mansourian, Xin Zhou, Shohreh Kasaei, et al. Soccernet game state reconstruction: End-to-end athlete tracking and identification on a minimap. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3293–3305, 2024. 1, 2

--- Page 1 ---

SoccerTrack v2: A Full-Pitch Multi-View Soccer Dataset for Game State Reconstruction

Atom Scott

Nagoya University∗

Ikuma Uchida

University of Tsukuba

Kento Kuroda

University of Tsukuba

Yufi Kim

University of Tsukuba

Abstract

SoccerTrack v2 is a new public dataset for advancing multi-object tracking (MOT), game state reconstruction (GSR), and ball action spotting (BAS) in soccer analytics. Unlike prior datasets that use broadcast views or limited scenarios, SoccerTrack v2 provides 10 fulllength, panoramic 4K recordings of universitylevel matches, captured with BePro cameras for complete player visibility. Each video is annotated with GSR labels (2D pitch coordinates, jersey-based player IDs, roles, teams) and BAS labels for 12 action classes (e.g., Pass, Drive, Shot). This technical report outlines the dataset’s structure, collection pipeline, and annotation process. SoccerTrack v2 is designed to advance research in computer vision and soccer analytics<ref:2508.

Improvements for AI systems

  1. Data integration for holistic analysis: The combined data allows for integrated tracking and event-based video analysis, enabling systems to perform simultaneous object tracking (MOT) and temporal event detection (BAS) within a single match context.

  2. Advanced game state reconstruction: By leveraging 2D pitch coordinates, player IDs (via jersey numbers), roles, and team affiliations, AI systems can reconstruct the complete spatial and tactical state of a soccer match frame-by-frame, going beyond simple object tracking to understanding positional dynamics.

  3. Multi-modal action spotting: The inclusion of 12 action classes (e.g., Pass, Drive, Shot) enables the development of sophisticated models capable of not only tracking players but also reliably detecting and classifying specific ball actions across the entire panoramic view.

Abstract

Soccer analytics draws on two kinds of information: spatio-temporal data describing where players and the ball are, and event data describing what they do. Public datasets offer them apart, or together only on broadcast footage that leaves players outside the frame unobserved. SoccerTrack v2 combines continuous full-pitch video, long player trajectories and actor-linked events in one resource: ten university-level matches, 932 minutes of fixed-camera 4K panoramic video, annotated per frame with metric pitch coordinates, jersey numbers and persistent identities, roles and team sides for all players, and with ball action events in twelve classes, linked to the acting players through the same identifiers used in the trajectories. We fix a match-level split and report baselines for two tasks. For game state reconstruction, we run a full pipeline over all twenty halves and find that GS-HOTA scores degrade as sequence length increases. For ball action spotting, we train a model on the player trajectories, with and without the ball track. The data, the split and the evaluation tooling are released so that both tasks can be developed and compared at match length on the same footage.

Related papers