H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities
Julius Broermann, Oliver Müller, Michael Döring, Jochen Baumeister
Paderborn University · SG Flensburg-Handewitt
cs.LG, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 13 pages, 6 figures, 1 table. Accepted at the 13th Workshop on Machine Learning and Data Mining for Sports Analytics (MLSA 2026), co-located with ECML PKDD 2026
Code: https://github.com/JuliusBroermann/handballaction
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper presents the first comprehensive adaptation and evaluation of Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP) frameworks for professional team handball,
Terminology
Summary
This paper presents the first comprehensive adaptation and evaluation of Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP) frameworks for professional team handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga (2021/22 to 2025/26). The authors develop Handball-xT (H-xT) using a handball-native court zoning layout featuring angular boundaries originating from the goals and concentric divisions matching the 6m and 9m lines, aggregating the defensive half into a single zone. Through simulations following the methodology of van Arem et al., they demonstrate that this native layout is systematically more robust than standard rectangular grids: For any number of zones, the grid model’s 90th percentile deviation is larger than that of the handball-native model.
The handball-native layout supports a capacity of 67 zones under the stability threshold (R90 ≤ 0.03), whereas the grid model is restricted to 45 zones. The xT surfaces reveal key tactical insights: threat increases with goal angle inside the 6m crease, a low-threat trough exists directly outside the 9m line, and wing corners outside the 6m line exhibit the highest threat outside the crease.
For Handball-VAEP (H-VAEP), the authors optimize the feature space, learning algorithm, and context length. They replace exact cumulative scores with a bucketed score difference (1b), replace the polar angle with the goal angle between rays connecting the ball to the two posts (2), and add elapsed seconds since gaining possession (3). These combined feature modifications (1b+2+3) yield compounding performance improvements over the football-native baseline. XGBoost achieves the best performance, with the gain from customizing features (+0.00596 scoring ROC-AUC over the tuned CatBoost baseline) being more than double the gain from subsequent model class selection (+0.00244). The authors select context length k = 3 as a balanced compromise between predictive quality and team-identity leakage, which increases monotonically with context length from 0.608 at k = 1 to 0.664 at k = 3 and 0.698 at k = 6. The final model configuration (tuned XGBoost, 1b+2+3, k = 3) reduces the scoring Brier score to 0.11220 and the conceding Brier score to 0.02735, while improving scoring ROC-AUC to 0.79582 and conceding ROC-AUC to 0.80683 compared to the original VAEP baseline. On three held-out seasons (2023/24 to 2025/26) with rolling training, the scoring model attains a Brier score of 0.11341 at a ROC-AUC of 0.79128 and the conceding model a Brier score of 0.02654 at a ROC-AUC of 0.79454, with excellent calibration (expected calibration error 0.006 for scoring and 0.001 for conceding over decile bins).
Empirical player ratings for the 2024/25 season (players with at least 500 total minutes and 250 minutes in offense) demonstrate strong face validity: the ten players with the highest H-VAEP per 10 minutes in offense are widely recognized elite performers who collectively received three IHF World Handball Player of the Year awards, three HBL MVP titles, two EHF Champions League Final Four MVP awards, and two HBL Best Young Player awards. The ranking is not a mere reproduction of box-score statistics—the player ranked first in H-VAEP/10o ranks only 55th in goals and 62nd in goals/10o, 9th in assists, 8th in assists/10o, and 27th in HPI. Action-type decomposition reflects distinct offensive roles: back players generate the majority of their value through passing (55% for left and right backs, 63% for centre backs), wings accumulate most of their value through dribbles (55%) and shots (33%) with passing contributing almost nothing (3%), and pivots derive value from shots (51%) and dribbles (53%) with slightly negative passing contribution (−4%). Team-level differences are also captured, such as SC Magdeburg's backs generating 38% of their value through dribbles versus SG Flensburg-Handewitt's backs showing the highest share through passing (65%). Illustrative sequences show H-VAEP rates a fast break's long pass as the most valuable action (+0.365), splits the pivot's contribution between dribble (+0.220) and shot (+0.292), and credits the goalkeeper +0.141 for an outlet pass—contributions ignored by box-score statistics and the HPI.
The comprehensive validation framework following Davis et al., Franks et al., and Robberechts et al. shows that H-VAEP and H-VAEP/10o exhibit exceptional reliability (within-season r ≥ 0.98, cross-season r ≥ 0.96), outperforming traditional metrics like goals and assists (within-season r ≥ 0.92, cross-season r ≥ 0.80) and shooting accuracy (r = 0.49 within-season). H-VAEP/10o performs best on discrimination (0.994) and stability (0.999) meta-metrics, while absolute H-VAEP retains high discrimination (0.990) but lower stability (0.910). Normalizing H-xT for playing time raises its stability from 0.164 to 0.976. Both H-VAEP (0.147) and H-VAEP/10o (0.163) score low on the independence metric, correlating strongly with goals (r = 0.72, r = 0.66) and goals/10o (r = 0.75, r = 0.77), while H-xT correlates strongly with assists (r = 0.78). Coaches confirmed the handball-specific zoning layout is markedly more intuitive than rectangular grids, and workshop discussions showed they intuitively grasped how H-VAEP decomposes and attributes value across multi-player build-up sequences. The authors release their complete code repository, including API connectors to data providers, to help professional clubs deploy these models. Future work includes extracting defensive events from raw tracking data, expanding to off-ball behavior valuation, and addressing event-detection noise for possession-based prediction targets.
Improvements for AI systems
Improvements to AI systems:
-
Domain-Adaptive Spatial Partitioning for Spatiotemporal Models – Replace generic rectangular grids with sport/domain-native zoning (e.g., angular boundaries from key points, concentric divisions matching rule lines). The improved AI system automatically learns or selects the most robust spatial discretization for any given domain, reducing prediction variance and enabling finer granularity (67 zones vs. 45) without overfitting.
-
Context-Length Optimization with Identity-Leakage Control – Implement a tunable context window (k) that balances predictive accuracy against over-reliance on team/player identity. The improved AI system can dynamically select k to maximize performance while minimizing leakage, making it applicable to other team sports or multi-agent systems where historical context can bias evaluations.
-
Feature Engineering via Domain-Specific Geometric Transformations – Replace raw polar angles with goal-angle (angle between two reference points) and bucket continuous score differences. The improved AI system can automatically derive such geometric and discretized features for any domain, leading to compounding gains (e.g., +0.00596 ROC-AUC) that exceed gains from algorithm selection.
-
Action-Type Decomposition for Role-Aware Valuation – Decompose player/agent value by action type (pass, dribble, shot) to reveal distinct functional roles. The improved AI system can generate role-specific profiles, enabling better team composition analysis, opponent scouting, and individualized training recommendations.
-
Normalization for Playing-Time Stability – Apply per-minute normalization to raw value metrics to boost reliability (stability from 0.164 to 0.976). The improved AI system can automatically detect and correct for time-dependent biases in any performance metric, making evaluations fairer across varying exposure.
-
Rolling-Training with Calibration Monitoring – Use rolling training windows and track expected calibration error (0.006 scoring, 0.001 conceding). The improved AI system can self-monitor calibration drift over time and trigger retraining or feature adjustments, ensuring sustained real-world accuracy.
-
Interpretable Sequence-Level Attribution – Provide action-by-action value decomposition for multi-agent sequences (e.g., fast break, pivot play). The improved AI system can explain contributions of each agent in a chain, enabling coaches and analysts to identify undervalued actions (e.g., goalkeeper outlet pass +0.141) that box-score metrics miss.
-
Meta-Metric Validation Framework – Automatically compute reliability, discrimination, stability, and independence metrics for any new model. The improved AI system can rank candidate models on these meta-metrics, guiding selection toward models that are both accurate and robust across seasons.
-
API-Integrated Deployment with Data Connectors – Package the AI system with modular API connectors to live data providers. The improved AI system can be deployed directly into professional clubs’ workflows, providing real-time player ratings and tactical insights without manual data handling.
-
Cross-Domain Transfer of Valuation Frameworks – The improved AI system can adapt the xT/VAEP framework to other sports (e.g., basketball, hockey) by redefining zones, features, and context lengths based on each sport’s rules, accelerating adoption of advanced analytics across domains.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks