SportD: How do VLMs physically strategize?

arXiv:2607.14616 · cs.AI, cs.CV · Submitted 2026-07-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SportD: How do VLMs physically strategize?".

Jane: The paper was written by Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Jinglin Xiao et al. from Princeton University and Rice University and University of California, Irvine and New York University and University of California, Santa Barbara.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We’re starting today with a really interesting paper titled "SportD: How do VLMs physically strategize?" which comes from a large group of researchers at Princeton, Rice, UC Irvine, NYU, and UCSB.

Jane: It's such a catchy title because it moves the conversation from what these models can see to how they actually behave in a high-stakes situation.

Tom: Exactly, Jane, because for a long time we've been testing if they can recognize a soccer ball or describe a stadium, but this is about making decisions.

Jane: It’s like the difference between reading a playbook and actually stepping onto the pitch to play the game.

Lu: I think that distinction is where things get really exciting for the future of robotics and autonomous systems.

Tom: Do you see it as a direct path to smarter machines, Lu?

Lu: Definitely, because if we can teach a model to understand tactical positioning in soccer, we're essentially teaching it to navigate complex social and physical spaces.

Meng: I wonder if that's practical for real-world deployment though.

Jane: What are you thinking, Meng?

Meng: Well, a soccer field is a very controlled environment compared to a busy street or a manufacturing plant, so I'm curious if these strategic lessons actually scale.

Lalam: I see it as more of a cultural shift in how we interact with technology.

Tom: How so?

Lalam: We are moving away from using AI as just a search engine or a text generator and toward treating it as an entity that can share our physical space and act alongside us.

Jane: That’s a big leap to take, but it's clearly where the research is heading.

Tom: Let's look at how these models actually performed when they were put to the test in these soccer scenarios.

Summary: Tom: Building on that idea of agency, "SportD: How do VLMs physically strategize?" shows us that current models are struggling significantly with actual decision-making.

Jane: They used one thousand four hundred fifteen real scenarios from the Men's and Women's World Cups to see if the models could pick the best next move, like a pass or a shot.

Tom: And the results were pretty humbling because even the best model only chose the optimal action about thirty-four point three percent of the time.

Jane: It’s actually quite startling that they're even less accurate than some human players making mistakes!

Meng: Is it just a matter of them being bad at seeing the players, then?

Tom: That's what I wondered too, Meng, but the researchers found something much deeper than just poor vision.

Jane: They found that these models are incredibly risk-averse, meaning they constantly choose "safe" options that don't actually help the team score.

Lu: The math there is really fascinating because they found a massive mismatch in how the models think about probability and value.

Tom: Can you break that down for us, Lu?

Lu: Sure, the models are great at predicting which action is likely to succeed, but they completely fail to realize which action is actually worth the risk.

Jane: So they think because a pass is easy to complete, it must be the best move?

Lu: Precisely, they treat "easy" as "valuable," even when a harder pass would have been much more effective for the team.

Lalam: That's a pattern that could be quite problematic if we apply it to society.

Tom: What do you mean by that?

Lalam: If an AI starts prioritizing the easiest path because it's "safe," it might avoid the complex, difficult tasks that are actually necessary for human progress.

Jane: It’s a bit of a wake-up call for how we train these systems.

Tom: But the authors didn't just point out the problem; they actually found a way to nudge these models toward better behavior.

Improvements: Tom: Following up on that mismatch, "SportD: How do VLMs physically strategize?" explores how we can use simple text to steer these models away from their cautious habits.

Jane: They used something called "risk-steering" prompts, where they basically told the model to look for the most ambitious play rather than just the safest one.

Tom: And it actually worked, Jane, because when they used a "SEEK" prompt, the skill levels of these frontier models jumped up significantly.

Jane: For example, GPT five point six Sol saw its skill score rise from zero point two seven to zero point three eight just by being told to favor penetrating passes and shots!

Meng: I have to play devil's advocate here, though.

Tom: Go ahead, Meng.

Meng: Using a prompt to fix a reasoning error feels like putting a bandage on a broken bone; it doesn't change the underlying architecture of the model.

Lu: I see it as more than just a bandage, Meng.

Jane: How so?

Lu: These prompts are acting like a key that unlocks capability that is already sitting there, dormant within the model's weights.

Tom: So you're saying the intelligence is present, but the training has just pushed them into this cautious corner?

Lu: Exactly, and finding ways to bake that ambition directly into their training could be our next big breakthrough.

Lalam: If we can master that, we'll move from having AI assistants to having true partners.

Jane: Partners that aren't afraid to take a calculated risk when it matters most.

Tom: It really brings us to the end of our discussion on this paper.

Conclusion: Tom: We’ve spent our time today looking at "SportD: How do VLMs physically strategize?" and seeing how much work is left to do in the realm of physical agency.

Jane: It's such a powerful reminder that being able to describe a scene is worlds apart from knowing how to act within it.

Tom: We've seen that models can be incredibly good at seeing, but they often lack the strategic "gut feeling" to value risk correctly.

Lu: I’m still thinking about the potential for these models to eventually master the physical world once we solve this valuation gap!

Meng: I'll be looking for the next paper that shows how engineers can actually implement these complex reward functions into real-world hardware.

Lalam: And I'll be watching how this shapes our culture as we learn to trust machines with more than just our information, but with our physical intentions too.

Jane: It’s been such a blast talking through this with all of you!

Tom: Thanks for tuning in, and we'll see you next time for another deep dive!

Princeton University · Rice University · University of California, Irvine · New York University · University of California, Santa Barbara

cs.AI, cs.CV

Submitted: 2026-07-16

Updated: 2026-09-14

Code: https://github.com/statsbomb/open-data

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 84/100

The gist: This paper introduces SportD, a benchmark designed to evaluate whether vision-language models (VLMs) can "turn visual understanding into good strategic actions" in physical environments.

Key concepts

Vision-Language Models (VLMs)
These AI models can process both visual and textual information. While they are proficient at describing what they see in a scene, this research highlights their difficulty in making strategic decisions or understanding how to act effectively within a physical environment.
Risk-aversion
This refers to the tendency of models to prioritize "safe" actions that are easy to complete over more difficult but valuable ones. The models struggle with a valuation gap, confusing the probability of an action succeeding with its actual strategic worth.
Risk-steering prompts
These are text-based instructions designed to steer models away from cautious habits and toward more ambitious behavior. By prompting models to favor penetrating passes or shots, researchers were able to significantly increase their strategic skill scores.

Terminology

Summary

This paper introduces SportD, a benchmark designed to evaluate whether vision-language models (VLMs) can turn visual understanding into good strategic actions in physical environments. By using professional soccer as an objective testbed with quantifiably-valued actions, the study investigates if models can move beyond mere scene description to make sound strategic decisions. This research is critical because current benchmarks focus on recognition and captioning rather than the reasoning required for agents to operate effectively in the real world.

The SportD Benchmark

SportD provides a comprehensive evaluation consisting of 1415 on-ball situations sourced from the 2022 men’s and 2023 women’s World Cups. To make a decision, a VLM is presented with:

  • Images from the previous 5 seconds, sampled at 0.5-second intervals.

  • An image of the moment immediately preceding the decision.

  • A bird’s-eye view map showing all visible player positions.

The model must select exactly one action from a set containing either a shot or a pass to a specific teammate. Every candidate action is assigned a quantitative value using the VAEP (Valuing Actions by Estimating Probabilities) model, which measures how much an action changes the acting team’s chances of scoring or conceding.

Systematic Risk Aversion

The findings indicate that VLMs consistently underperform professional players in strategic decision-making. The best-performing model selects the optimal action only 34.3% of the time, and every VLM lands between RANDOM and the real player on all primary metrics. Most significantly, models exhibit a clear preference for safer actions, choosing lower-variance, lower-value choices on 62–69% of events. This conservatism is reflected in the physical progression of play:

  • The optimal pass advances the ball by approximately 9 meters.

  • Real players advance the ball by roughly 2 meters.

  • VLMs' passes barely move the ball forwards, often resulting in near-zero or even backward progression.

Miscalibration of Value and Likelihood

The study reveals that VLMs' errors are not merely due to poor perception but stem from a miscalibration of value. While frontier models are proficient at estimating likelihood—placing the highest-success-probability action among their top choices in 83–92% of cases—they struggle to identify which actions are actually valuable, with expected-value recall falling to 56–66%. This occurs because models systematically conflate likelihood with value. While the ground truth shows a slight negative correlation between success probability and expected value (rho = -0.08), every VLM assigns them a positive relationship (rho = +0.30 to +0.52). Essentially, the models behave as though what is easier to execute [is] what is more valuable.

Steerability of Policy Bias

To test if these failures represent a capability ceiling or a policy bias, researchers used risk-steering prompts to alter the model's deliberation. By replacing the standard prompt with one that instructs the model to either seek upside (SEEK) or avoid risk (AVOID), they found that frontier models are highly steerable. A single risk-seeking instruction then lifts the frontier models towards the real players’ skills, improving performance even on subsets where shooting is an incorrect option. Conversely, steering toward retention lowers skill, confirming that the models' default policies lie on the conservative side of optimal.

Improvements for AI systems

1. Dual-Stream Probabilistic and Value Reasoning Architecture

  • Improvement: Replace end-to-end action prediction with a decoupled, multi-stage reasoning pipeline. Implement two distinct specialized heads or prompting modules: one dedicated to estimating the Success Probability (P success) of an action and a second dedicated to estimating the Change in State Value (V) resulting from that action.

  • Improved Capability: The AI will stop conflating ease of execution with strategic worth. It will be able to identify high-risk, high-reward maneuvers (e.g., a penetrating pass with low completion probability but massive goal-scoring potential) rather than defaulting to low-variance, low-value actions (e.g., a safe sideways pass).

2. Expected Utility Optimization via Value-Based Reinforcement Learning (RLPF)

  • Improvement: Transition from standard RLHF (Reinforcement Learning from Human Feedback) to RLPF (Reinforcement Learning from Physical/Strategic Feedback). The reward function must be explicitly tied to the mathematical change in expected value (V) and goal-ward progression, rather than just human-like imitation or task completion.

  • Improved Capability: The system will exhibit calibrated risk tolerance. Instead of a standardized, uniform risk aversion across all scenarios, the AI will dynamically adjust its aggression based on whether the expected payoff justifies the variance of the action, effectively minimizing Regret (R).

3. Counterfactual Strategic Deliberation (Chain-of-Thought Refinement)

  • Improvement: Integrate a mandatory Counterfactual Search step within the model's deliberation process. The model must be prompted (or trained via fine-tuning) to explicitly simulate and compare the outcomes of the top N candidate actions, specifically asking: Does this high-probability action yield significantly less value than a lower-probability alternative?

  • Improved Capability: This prevents defaulting to the most likely outcome. The AI will be able to perform a mental dry run of multiple strategic paths, allowing it to select actions that maximize long-term utility even when those actions are visually or statistically more complex.

4. Spatiotemporal Vector Integration for Physical Progression

  • Improvement: Augment visual tokens with explicit spatiotemporal vector data (e.g., goal-ward gain vectors and velocity tensors). This integrates the directionality of an action directly into the latent representation of the scene.

  • Improved Capability: The AI will prioritize progressive actions. It will avoid the current VLM failure mode of selecting actions that move the agent laterally or backward, ensuring that every strategic decision contributes to physical progress toward a defined objective or goal.

Abstract

Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next. Models on average select the optimal action around 27% of the time, less often than the professional players, and capture markedly less of the value at stake. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 72-85% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed (ρ=+0.30 to +0.52), despite no such relationship in the ground truth (ρ=-0.08). Modifying the deliberation instructions to encourage risk-taking brings the frontier models closer to the players' skill levels. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.

Sources

Related papers