Do Vision Language Models Understand Human Engagement in Games?

arXiv:2603.18480 · cs.CV, cs.AI, cs.HC · Submitted 2026-03-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Do Vision Language Models Understand Human Engagement in Games?".

Tom: Inferring human engagement from gameplay video is important for game design and playerexperience research,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: To get into the main idea, the authors are testing if vision-language models can predict human engagement from gameplay video, and they’re checking different prompting strategies to see what helps them succeed. This whole paper is about putting these complex AI systems through a real test of their ability to interpret player feelings in a game context.

Jane: Exactly, Tom; it moves beyond just looking at the visual elements of the game and tries to connect those visuals to how a person is actually feeling while playing. They are essentially asking if an AI can read the mood of a match from a video clip.

Lu: What really stands out is that they aren't just testing one approach; they’re comparing zero-shot predictions against using specific prompts based on theories like Flow or Self-Determination Theory to see if those frameworks actually help the models reason better about engagement.

Meng: I noticed they are looking at both pointwise prediction, where the model guesses the feeling for a single frame, and pairwise prediction, which tries to predict how engagement changes between consecutive frames. That suggests they are trying to capture dynamic shifts in player experience over time.

Lalam: The paper’s summary really highlights that raw zero-shot predictions are pretty weak; they only hit about fifty-seven percent accuracy across all the models tested when compared to just looking at the majority class of game labels.

The paper's summary: Tom: That low starting point is a key finding in this paper, showing that without extra help, these vision-language models struggle significantly to predict engagement accurately from gameplay footage. They found that they can’t just guess the label well on their own when they are asked to infer those deeper psychological states.

Jane: And what's particularly telling is how the authors found that few-shot demonstrations, giving the model examples, actually provide a substantial lift in accuracy, sometimes adding eighteen percentage points to their performance. That’s a pretty significant gain for improving these predictions.

Lu: I think they also pointed out that just using theory-guided prompts on its own doesn't always work as expected; it often ends up amplifying surface-level shortcuts instead of genuinely boosting the model's reasoning capabilities, which is a nuanced point.

Meng: From a practical standpoint, this means we can’t rely on theory alone to fix these models; they need concrete examples or other structured methods to really get the performance up.

Lalam: And another major summary point is that there's a definite gap between what the models perceive and what they actually understand; for instance, they tend to confuse things like visual intensity with engagement, which isn't what humans are labeling in the data.

The paper's improvements: Tom: Now we look at how the researchers suggest fixing these issues because their zero-shot results weren't very strong. They propose a few directions for future work, focusing on building a more robust pipeline rather than just tweaking the prompt slightly.

Jane: They suggest using a hybrid pipeline where the VLM features aren't just fed into a classifier but go through several steps of reasoning before they make the final judgment about engagement. This is about forcing intermediate thinking to happen.

Lu: That idea of decomposing the reasoning—identifying game state first, then assessing challenge—seems like a very promising way to tackle that perception-understanding gap they identified in their analysis.

Meng: I see the suggestion for temporal smoothing as something very practical for real-world use; if the predictions are inconsistent frame by frame, applying a technique like an exponential moving average could enforce coherence over time.

Lalam: The authors also suggested using game-specific calibration techniques, like Platt scaling, to tune the model's confidence based on how it performs on that particular game before making a final call.

Conclusion: Tom: So wrapping up this discussion of "Do Vision Language Models Understand Human Engagement in Games?", the main message is that while these models can perceive what’s happening visually, they currently lack the ability to truly infer the underlying human engagement state without some kind of structured help.

Jane: It seems like we need to move away from relying on simple visual input alone and instead incorporate multiple layers of reasoning and calibration to get reliable results for player experience research. This paper really lays out exactly where the current technology is succeeding and where it needs more work.

Lu: The implication for the broader field is that we need to combine vision with structured psychological theory to build models that actually capture complex human experiences, not just visual patterns.

Meng: For us in development, this points toward building systems that are aware of context—knowing whether a frame is part of an active match or just a static menu screen—before trying to guess the feeling.

Lalam: I think the biggest impact here is showing that for research into player experience, we need models that understand the difference between what's visually intense and what actually makes someone feel engaged.

Tom: That’s a solid summary of "Do Vision Language Models Understand Human Engagement in Games?". Thanks to Lu, Meng, and Lalam for weighing in on all these points. We’ll be right back after the break.

University of Maryland, College Park · University of Southern California · Arizona State University

cs.CV, cs.AI, cs.HC

Submitted: 2026-03-19

Updated: 2026-10-01

Importance score: 70/100

The gist: Inferring human engagement from gameplay video is important for game design and playerexperience research, yet it remains unclear whether vision–language models (VLMs) can infer such latent

Key concepts

Player Engagement
This is a subjective quality of the user experience that changes over time in a game. It cannot be directly observed; instead, researchers try to infer it by looking at visual cues like close matches or intense action, requiring complex reasoning.
Few-Shot Demonstrations
Giving the VLM a few examples of correct predictions helps it learn better than asking it to predict from scratch. This shows that providing concrete examples can substantially boost performance, suggesting VLMs need guided learning to improve their inference.
Perception–Understanding Gap
This refers to the difference between what a model sees (perception) and what it actually knows about the game's meaning (understanding). The study found models are good at seeing surface details but struggle with complex, deep interpretations required for accurate engagement prediction.

Terminology

Summary

Inferring human engagement from gameplay video is important for game design and playerexperience research, yet it remains unclear whether vision–language models (VLMs) can infer such latent psychological states from visual cues alone. The gist: raw zero-shot predictions are unreliable (∼57%, below the majority-class baseline), but few-shot demonstrations yield substantial pointwise gains (up to 75.0% for Qwen), while theory-guided prompting alone can amplify surface-level shortcuts rather than improve reasoning.

Background and Research Questions

Player engagement is a multi-dimensional, subjective, and temporally dynamic quality of the user experience that must be inferred rather than observed. Games provide an ideal testbed because they decouple visual intensity from engagement; for instance, a scoreboard showing a close match requires multi-step reasoning to infer tension. The study investigates three research questions: Q1: Can VLMs predict engagement from gameplay frames, and do few-shot demonstrations or retrieval strategies improve prediction? Q2: Do theory-aligned probes (Flow, GameFlow, SDT, MDA) help VLMs capture engagement-related dimensions? Q3: Are VLM predictions consistent across different games and across different models? The work is the first to combine all seven dimensions: visual input, VLM evaluation, game domain, engagement prediction, theory-grounded probes, cross-game transfer, and systematic failure analysis.

Experimental Setup and Methodology

The researchers use the GameVibe FewShot (GVFS) dataset across nine first-person shooter games. The task involves predicting a binary engagement label (High/Low) from 16 frames sampled from a 1-second gameplay window, evaluated in two settings: pointwise prediction and pairwise prediction of engagement change between consecutive windows. To test the capability of VLMs, three architecturally distinct models are evaluated: InternVL3.5-8B-Instruct, Qwen3-VL-8B-Instruct (open-source), and GPT-4o (proprietary). Six prompting strategies are tested, including zero-shot prediction (S1), theory-guided prompts grounded in Flow, GameFlow, Self-Determination Theory (SDT), and MDA frameworks (S2), and retrieval augmented prompting using both VLM embeddings and CLIP embeddings for S3 through S6.

Key Findings on Performance

Overall accuracy is well below the majority-class baseline, with averaged zero-shot prediction achieving ∼57% across all three models. Performance varies dramatically by game, ranging from 83.1% on Borderlands 3 (vs. 81.4% majority) to 30.5% on CS 1.6 (vs. 78.0% majority). Theory-guided prompting alone does not consistently improve over plain zero-shot prompting (S2 vs S1), but when combined with few-shot demonstrations (S5), theory guidance often yields additional improvements, suggesting it is most effective when paired with concrete examples. Retrieval-augmented prompting shows model-dependent effects; for instance, Qwen3-VL benefits substantially from CLIP-based retrieval. Pairwise prediction remains consistently difficult across all strategies.

Systematic Failure Modes

The analysis identifies five systematic failure modes exposing a perception–understanding gap:

  1. Visual Intensity Bias: VLMs systematically conflate visual intensity (color saturation, edge density, etc.) with engagement, correlating significantly with this score on CS:GO18 while showing no correlation with human labels.

  2. Surface Feature Shortcuts: Models rely on shortcuts such as Relatedness bias (teammates visible) and Sensation bias (explosions/blood splatter), which map directly to theory dimensions but in a shallow, stimulus-response manner.

  3. Post-Match Context Blindness: VLMs cannot distinguish active gameplay from post-match screens, often classifying scoreboards showing close victories as static, no active gameplay and predicting Low engagement.

  4. Temporal Inconsistency: Predictions lack temporal coherence; the flip rate for VLM predictions is dramatically higher than the human label flip rate, indicating that models achieve near chance-level temporal consistency.

  5. Confident but Wrong Reasoning: There is an inverse confidence–accuracy relationship; high-confidence predictions achieve only 38.5% accuracy, confirming that confident VLM reasoning is systematically misleading.

Implications and Recommendations

The findings suggest a capability hierarchy where VLMs succeed at perception and description but fail at interpretation and inference. The study recommends five concrete directions for future work: (1) Hybrid pipeline using VLM features as input to a supervised classifier; (2) Temporal smoothing of per-window predictions; (3) Decomposed reasoning via chain-of-thought to identify game state before assessing challenge or feedback; (4) Game-specific calibration techniques like Platt scaling; and (5) Structured output with verification to process extracted game states by deterministic logic.

Improvements for AI systems

Based on the findings in this paper, here are specific, actionable improvements for AI systems designed for game analytics and player experience research:


) 1. Implement a Multi-Stage Reasoning Pipeline (Chain-of-Thought):

Instead of a single VLM prediction, build a pipeline that forces intermediate reasoning steps. The system should be decomposed into sequential modules:

(i) Game State Identification: Determine if the current frame is in an active gameplay state, menu/pause screen, or post-match screen (addressing Failure Mode 3).

(ii) Engagement Dimension Assessment: Analyze the frame against four theoretical frameworks (Flow, GameFlow, SDT, MDA) to quantify specific dimensions like challenge level or social presence.

(iii) Contextual Synthesis: Combine the scores from steps (i) and (ii), conditioned on whether the state is active gameplay or not, to output a final engagement label.

[Improved System Capability]: This system moves beyond simple perception by enforcing multi-step reasoning, allowing it to distinguish between intense but routine moments and genuinely engaging ones, as demonstrated by the failure of S1 and S2 alone.

) 2. Develop Domain-Specific Calibration Layers:

Integrate a game-specific calibration step based on per-game majority class performance metrics (as suggested in Section 6.3). Before final output, apply Platt scaling or temperature scaling to the VLM's logits using a small set of labeled windows from that specific game.

[Improved System Capability]: This addresses the model-specific biases observed across different games (e.g., CS:GO18 vs. Borderlands 3), ensuring that the model’s confidence and prediction calibration are tuned to the unique visual language and engagement patterns of that specific title, mitigating confident but wrong reasoning (Failure Mode 5).

) 3. Prioritize Concrete Demonstrations Over Elaborate Prompts:

When using retrieval-augmented prompting (S3-S6), shift the focus from complex theory-grounded prompts to a strategy where demonstration examples are explicitly balanced for label distribution. The system must incorporate a mechanism to detect and correct memory imbalances (as shown in Figure 3).

[Improved System Capability]: This ensures that retrieval augmentation actually corrects weaknesses instead of amplifying them. By dynamically balancing the positive/negative labels retrieved from memory, the system prevents systematic over-prediction of negative engagement states when few positive examples are available, leading to more reliable few-shot performance gains.

) 4. Integrate Temporal Smoothing and Coherence Enforcement:

Implement a post-processing layer (e.g., an Exponential Moving Average or lightweight LSTM) that operates on the sequence of per-window predictions to enforce temporal smoothness based on high autocorrelation metrics (as shown in Table 7).

[Improved System Capability]: This directly combats the Temporal Inconsistency (Failure Mode 4), ensuring that engagement predictions are temporally coherent, which is crucial for modeling player experience over time rather than treating every frame as an independent event.

) 5. De-emphasize Visual Intensity Reliance:

Explicitly train or fine-tune the VLM to decouple visual intensity features (color saturation, edge density) from engagement labels during pre-training or prompt design, as these features are shown to be highly correlated with incorrect predictions (Failure Mode 1).

[Improved System Capability]: The system will learn to ignore spectacle cues and focus on the latent psychological state rather than simply classifying high-arousal visual stimuli. It will avoid conflating a visually intense explosion with a low-engagement moment if it can identify the underlying context (e.g., post-match screen).

Sources

Related papers