Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents

arXiv:2608.07490 · cs.HC, cs.AI · Submitted 2026-06-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents".

Jane: The paper was written by Yingying Guo, Zhuoxuan Ju, Ruibo Ming, Ruicheng Feng and Jinjin Gu from The Chinese University of Hong Kong, Shenzhen and Georgetown University and INSAIT, Sofia University St. Kliment Ohridski and Tencent.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the channel, everyone. I'm Tom, and sitting across from me is the wonderful Jane. Today we're digging into a fresh arXiv paper that's got a title that really makes you stop and think: "Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents."

Jane: Tom, I have to say, when I first read that title, I was hooked. It's not just about whether AI can win games. It's about whether AI actually learns from playing them, the way people do. That's a whole different question.

Tom: Exactly. And the authors—Yingying Guo, Zhuoxuan Ju, Ruibo Ming, Ruicheng Feng, and Jinjin Gu—they're coming at this from a really interesting angle. They're not just measuring final scores. They're watching how decisions change over time.

Jane: Right. So imagine you're learning Othello. Your first instinct might be to flip as many discs as possible because it looks like you're winning. But after a few games, you realize that's a trap. You learn to think about mobility, about positioning, about what comes next.

Tom: And that's what they call "experience-sensitive learning." The game itself isn't changing, but the player's internal model of the game is. And that's what the paper is really about—making that invisible change visible.

Jane: For humans, that's natural. We all learn from experience. But for language agents, the big language models that play these games, the question is whether they're actually learning in that same way or just getting better at pattern-matching within a single episode.

Tom: And spoiler alert, Jane—the paper finds that current agents are pretty bad at this. They don't show the same stable, interpretable shifts that humans do. Their improvements are noisy, transient, and often disappear.

Jane: Which is a huge deal, because if we want AI to be truly adaptive, to learn from doing, then this is the gap we need to close. This paper gives us a framework to actually measure that gap.

Tom: And that framework is what we're going to dig into next. Stick around, because we're about to get into the nitty-gritty of how they built these games and metrics to catch learning in the act.

Summary: Jane: So Tom, we've set the stage with the title. Now let's talk about what this paper actually does. The summary lays it out pretty clearly: they're introducing a whole framework for studying how gameplay experience changes decision-making behavior.

Tom: And they do it with a suite of three games. We've got Four in a Row, Othello6, and CircleCat. Each one has this built-in tension between a greedy, short-sighted move and a smarter, more global strategy.

Jane: Right. In Four in a Row, the greedy move is just extending or blocking the most obvious line. But strong play requires thinking about forced threats and multi-step setups. In Othello6, the greedy move is flipping the most discs right now, but that can wreck your future mobility.

Tom: And CircleCat is this spatial game where you're trying to trap a cat by placing walls. The greedy move is blocking the cell closest to the cat, but smart play means thinking about the whole escape structure, closing off bottlenecks and exits.

Jane: So they designed these games to be experience-sensitive. That means repeated play can actually change how you evaluate a position. And they built metrics to track that change—not just win rates, but behavioral shifts.

Tom: The two big cross-game metrics are Greedy Value Difference and Greedy Trap Avoidance. GVD measures how much better your moves are than the greedy baseline. GTA measures how often you avoid falling into a greedy trap when a better move exists.

Jane: And then they have game-specific metrics too. Things like solver alignment in Four in a Row, mobility control in Othello6, and escape-path lengthening in CircleCat. These are the instruments that make learning visible.

Tom: Then they collected human gameplay data from thirty-two participants, over seven hundred games total. And they ran four different self-evolving language agents through the same games, using the same metrics.

Jane: And the results? Humans show this beautiful, interpretable shift from greedy to global thinking over just a few games. The agents? Not so much. They show noisy, transient gains that don't stick.

Tom: It's like the agents are learning the vocabulary of strategy without actually learning to speak the language fluently. They can say the right things, but they don't consistently do the right things.

Jane: And that's the core finding. We'll get into the details of those human trajectories and agent failures in the next segment, so stay with us.

Improvements: Tom: Jane, we've talked about what the paper found. But what does it suggest we should do about it? What are the improvements it's pointing toward?

Jane: Well, the paper doesn't just say "agents are bad at this." It actually diagnoses why. They identify a four-level hierarchy of bottlenecks that prevent agents from turning experience into stable strategic improvement.

Tom: Let's break those down. Level one is the simulation bottleneck. That's when an agent can't reliably compute what the board looks like after a move. In Four in a Row, that means misjudging gravity or line completion even when it's talking about checking for wins.

Jane: Level two is the evaluation bottleneck. The agent recognizes that future consequences matter, but it can't actually compare candidate moves by their downstream value. It might say "I need to think about mobility" but then make a move that destroys its own mobility.

Tom: Level three is the proceduralization bottleneck. This is when an agent stores good strategic principles in its memory, but those principles never become executable rules. It's like reading a chess book and then making the same blunder anyway.

Jane: And level four is the experience-integration bottleneck. Even after accumulating lots of memory and reflection text, the agent repeats structurally similar errors. It doesn't consolidate what it's learned into persistent behavioral change.

Tom: So the improvement the paper is really calling for is a shift in how we build self-evolving agents. It's not enough to just store reflections. Agents need to convert those reflections into verifiable, executable decision procedures.

Jane: And that's a concrete design goal. Instead of just appending text to a memory file, the agent should be building something like a decision rule: "When I see this pattern, I must check this specific thing before acting."

Tom: The paper's framework gives us the metrics to test whether that's actually happening. You can't just claim your agent learned from experience. You can run it through these games and see if the Greedy Trap Avoidance curve actually goes up and stays up.

Jane: That's the practical value. It turns "learning from experience" from a vague aspiration into something measurable. And that's a big deal for anyone building agents that need to adapt to new environments.

Tom: We'll see how that plays out in the actual experiments next, when we look at the first page of the paper and the concrete numbers behind all this.

First Page: Tom: So Jane, let's get into the actual first page of "Experience-Sensitive Game Learning." The abstract really sets the tone. It says most benchmarks emphasize final outcomes rather than how players learn from repeated interaction.

Jane: And that's the core critique. Win rates and final scores are static snapshots. They don't tell you anything about the journey. This paper wants to study the journey itself.

Tom: The abstract also introduces that key term, "experience-sensitive game learning." And it defines it as settings where repeated gameplay can change a player's future decisions by inducing reusable strategic knowledge.

Jane: I love that phrase, "reusable strategic knowledge." It's not just getting better at one specific board state. It's learning principles that transfer across many states within the game.

Tom: And they use Othello as their motivating example. A novice flips the most discs because it looks like progress. But experienced players learn that greedy flips can reduce mobility and expose unstable discs. Same board, different interpretation.

Jane: That's the heart of it. What changes isn't just the outcome. It's the heuristic used to evaluate actions. And that's what the paper's metrics are designed to capture.

Tom: Now, the first page also mentions the three contributions. First, they formulate the framework. Second, they introduce the game suite and behavioral metrics. Third, they collect human data and evaluate self-evolving agents in the same metric space.

Jane: And that third point is crucial. They're not using humans as a performance ceiling. They're using humans as learning subjects whose behavioral changes can be analyzed with the same tools used for agents.

Tom: So when we see that human Greedy Trap Avoidance improves and stabilizes, while agent GTA is noisy and transient, we're making an apples-to-apples comparison. That's the methodological strength here.

Jane: The first page also hints at the related work landscape. They're building on text-game platforms like TextWorld and Jericho, but they're pushing beyond simple task completion toward process-level evaluation.

Tom: And they're connecting to classic cognitive science too. Work by Chase and Simon on chess perception, Gobet and Simon on templates, Gershman on computational rationality. The idea that expertise is about reusable perceptual chunks and heuristics.

Jane: So this paper is sitting at this intersection of cognitive science and AI evaluation. It's saying, if we want agents to learn like humans, we need to measure learning like we measure human learning.

Tom: And that's a profound shift. We'll wrap up our thoughts on the whole paper in our conclusion segment right after this.

Conclusion: Jane: Alright Tom, we've covered a lot of ground on "Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents." Let's bring it all together.

Tom: The big picture is this. The paper gives us a way to see whether players—human or machine—are actually learning from repeated gameplay. Not just winning more, but changing how they make decisions.

Jane: And the human data shows that learning is real and interpretable. People shift from greedy heuristics to global strategies. They avoid traps more often. They improve on solver alignment, mobility control, and spatial containment.

Tom: The agents, on the other hand, show a different story. They accumulate memory and reflection, but they don't convert that into durable behavioral change. Their improvements are noisy and often regress.

Jane: And the four-level bottleneck hierarchy explains why. Simulation failures, evaluation failures, proceduralization failures, and experience-integration failures. Each one blocks a different stage of learning.

Tom: The impact here is significant. For anyone building self-evolving agents, this paper is a wake-up call. Storing reflections isn't enough. You need executable, verifiable decision procedures.

Jane: And for the broader AI community, it's a reminder that final outcomes can hide important differences in how systems learn. Process-level evaluation matters.

Tom: So as we say goodbye to this paper, we're left with a clear challenge. Can we build agents that not only play well but learn well? Can we make their learning curves look like human learning curves?

Jane: That's the next frontier. And this paper gives us the tools to measure progress toward it.

Tom: Thanks for joining us on the channel. We'll be back soon with another paper to break down. Until then, keep learning from experience.

Jane: And keep asking whether your AI is actually doing the same. See you next time.

Yingying Guo, Zhuoxuan Ju, Ruibo Ming, Ruicheng Feng, Jinjin Gu

The Chinese University of Hong Kong, Shenzhen · Georgetown University · INSAIT, Sofia University St. Kliment Ohridski · Tencent

cs.HC, cs.AI

Submitted: 2026-06-15

Updated: 2026-08-11

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 57/100

The gist: The paper "Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents" studies how repeated gameplay changes the decision-making behavior of both humans and large language

Key concepts

Experience-Sensitive Game Learning
This refers to settings where repeated gameplay can change a player's future decisions by inducing reusable strategic knowledge. The paper studies how this learning manifests in both humans and language agents.
Greedy Value Difference (GVD)
This metric measures how much better an agent's moves are compared to a greedy baseline. It helps quantify the improvement an agent makes beyond simply making the most obvious move at any given moment.
Four-level bottleneck hierarchy
This describes four stages preventing agents from turning experience into stable strategic improvement: simulation, evaluation, proceduralization, and experience-integration bottlenecks.

Terminology

Summary

The paper Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents studies how repeated gameplay changes the decision-making behavior of both humans and large language model (LLM) agents. The authors introduce a framework called experience-sensitive game learning, which they define as settings in which repeated gameplay can change a player's future decisions by inducing reusable strategic knowledge. They argue that "in an experience-sensitive game, effective play is not determined only by reasoning from the rules. It also depends on whether the player can discover patterns, revise misleading heuristics, learn and reuse lessons from prior episodes."

The paper's central motivation is that "existing studies on game-based LLM agent evaluation have largely emphasized within-episode competence, treating games as tests of reasoning, planning, or task completion within individual episodes. Much less attention has been paid to the second form: whether repeated interaction changes the way an agent evaluates states, selects actions, and reuses experience across episodes."

The authors' contributions are threefold: "(1) We formulate experience-sensitive game learning as a framework for studying how humans and language agents change their behavior through repeated gameplay; (2) We introduce a suite of experience-sensitive games and game-specific behavioral metrics that make experience-driven behavioral change observable beyond final performance; and (3) we collect human gameplay trajectories and evaluate recent self-evolving language agents in the same metric space, enabling us to analyze human learning, agent self-improvement, and human-agent behavioral differences."

The paper introduces a suite of three games selected according to four criteria: "1. success should require reusable strategic principles rather than only local pattern matching; 2. each game should contain a natural but potentially misleading greedy heuristic; 3. each game should admit interpretable process-level metrics computable from action traces; 4. each game should be suitable for both human data collection and agent evaluation under a shared interface."

The three games are: Four in a Row, Othello6, and CircleCat. Four in a Row is described as our solver-aligned adversarial search environment where strong play requires more than extending visible lines or blocking immediate threats. Players must learn to recognize immediate wins, forced threats, multi-step threats, and the strategic value of central columns. Othello6 is a 6x6 version of Othello that serves as our environment for studying greedy-trap avoidance where the strategic learning target is therefore a shift from immediate flip maximization toward mobility-aware and positionally robust evaluation. CircleCat is a spatial containment game on a hex-neighbor board where effective play requires reasoning about the global escape structure of the board.

The authors define two cross-game metrics. Greedy Value Difference (GVD) measures the average improvement of a player's actions over a game-specific greedy baseline, computed as a softmax-normalized value difference. Greedy Trap Avoidance (GTA) focuses on states where that greedy baseline leads to a strategically poor action, measuring the fraction of times the player avoids the greedy trap by choosing an action that is better than the greedy baseline.

They also define game-specific behavioral metrics. For Four in a Row: EAR (solver alignment), CPQ (critical-state decision quality), and OCI (opening center control). For Othello6: MCI (mobility control) and CPQ (search-critical decision quality). For CircleCat: EPD (escape-path lengthening), EPC (escape-option reduction), and CAR (reachable-region compression).

For human data collection, the authors collected human gameplay trajectories for the three environments from 32 participants, including master and doctoral students, resulting in a total of 709 games. After filtering, the retained data consist of 12 participants for Four in a Row, 11 for Othello6, and 9 for CircleCat. Participants were required to report no recent experience with the corresponding games and completed at least 11 full games. Trajectories were divided into Early, Mid, and Final stages, corresponding to the first three games, the middle five games, and the final three games.

The human results show that across the three environments, the median human trajectory improves within the first two game groups, suggesting that participants begin to move away from locally greedy heuristics after limited gameplay experience. The game-specific metrics show broad improvement along the strategic dimensions targeted by each environment: in Othello6, CPQ improves for 10/11 participants and MCI improves for 7/11; in CircleCat, CAR improves for 8/9, while EPC and EPD improve for 9/9; in Four in a Row, CPQ improves for 11/12 participants, while EAR and OCI improve for 12/12.

The authors then evaluate four representative agent designs: CEL, EVOTEST, EVOLVER, and REASONINGBANK, running 128 gameplay episodes per agent per game. The results show that "current self-evolving agents do not exhibit the stable behavioral improvement observed in human trajectories. Although individual methods sometimes improve over short stretches, the gains are often noisy, method-specific, and not consistently maintained across later episodes. The authors conclude that the issue is not simply that agents achieve lower final performance than humans, but that their behavioral changes are less durable across repeated gameplay."

The paper proposes a four-level hierarchy of bottlenecks explaining agent failures. Level 1: the simulation bottleneck concerns whether the agent can reliably compute the next state after an action. Level 2: the evaluation bottleneck concerns whether the agent can compare candidate actions by downstream value. Level 3: the proceduralization bottleneck concerns whether useful strategic language becomes executable decision-time constraints. Level 4: the experience-integration bottleneck concerns whether repeated interaction produces stable behavioral change.

The authors conclude: "We introduced experience-sensitive game learning as a framework for measuring how humans and language agents change their decision-making through repeated gameplay. Our human trajectories show clear and interpretable behavioral changes: players increasingly avoid greedy traps and improve along strategic dimensions such as solver alignment, mobility control, and spatial containment. In contrast, recent self-evolving language agents show noisier and less durable gains."

The paper lists limitations including Limited number of environments, Variability and cost of human data, and Comparison with trainable game agents. Regarding the latter, the authors state: "Our goal is not to test whether a trainable policy can eventually master a game, but to examine whether language agents can use limited repeated interaction to revise their decision-making in a way that is behaviorally comparable to human learning."

Improvements for AI systems

Based on the paper's findings, here are the specific improvements I can implement in an AI system, along with what the improved system can do:

Improvements to the AI System

  1. Implement a hierarchical failure-detection module that monitors the agent's decision-making at four levels: state simulation, action evaluation, strategy proceduralization, and experience integration. The system will log and flag failures at each level in real-time, such as when the agent miscomputes a board transition (Level 1), fails to compare future consequences (Level 2), stores a strategy without executing it (Level 3), or repeats a structurally similar error despite accumulated memory (Level 4).

  2. Add a greedy-trap avoidance mechanism that, before each action, computes a game-specific greedy baseline (e.g., maximum-flip in Othello, nearest-cell block in CircleCat, local line completion in Four in a Row) and compares the candidate action's value against that baseline. If the candidate action is worse than the greedy baseline in a state where the baseline is known to be a trap (as defined by a threshold), the system will flag the action and force a re-evaluation.

  3. Introduce a cross-episode behavioral consistency check that tracks the agent's Greedy Value Difference (GVD) and Greedy Trap Avoidance (GTA) scores over sliding windows of five episodes. The system will detect when gains are noisy or transient (e.g., improvement in one window followed by regression in the next) and trigger a memory-compression step that distills the most recent successful strategies into compact, executable rules rather than accumulating verbose text.

  4. Implement a state-simulation verifier that, after the agent selects an action, independently simulates the resulting state using a deterministic game engine and compares it against the agent's claimed next state. If a mismatch is detected, the system will log the error, correct the state, and add a corrective example to the agent's memory to prevent future simulation failures.

  5. Add a proceduralization enforcement layer that converts high-level strategic principles (e.g., preserve mobility, seal shared gateways) into concrete, state-triggered action constraints. For example, in Othello, the system will compute the number of legal moves for both players after each candidate action and reject actions that lead to a forced pass unless an immediate tactical win justifies it. In CircleCat, the system will require the agent to place walls that reduce the cat's reachable boundary cells by at least one before the cat enters the near-boundary zone.

  6. Implement a long-horizon credit-assignment module that, after a loss, traces the terminal failure (e.g., forced pass, boundary escape) back to the earliest decision that contributed to it. The system will store the causal chain (e.g., move at step 5 reduced mobility, leading to forced pass at step 31) and use this to prioritize which past decisions to revise in the next episode, rather than only reflecting on the terminal symptom.

What the Improved AI System Can Do

  • Reliably simulate game states: The system will no longer miscompute gravity, landing positions, or line completions in Four in a Row, and will correctly predict future legal moves in Othello, preventing forced passes caused by simulation errors.

  • Avoid greedy traps consistently: In Othello, the system will reject maximum-flip moves that reduce future mobility; in CircleCat, it will place walls that close bottlenecks and reduce reachable exits rather than merely blocking cells near the cat; in Four in a Row, it will choose solver-aligned moves over locally visible line completions.

  • Convert verbal strategies into executable behavior: The system will not just store principles like preserve mobility but will enforce them as hard constraints at decision time, leading to stable improvements in mobility control (MCI) and critical-state decision quality (CPQ) across episodes.

  • Exhibit durable, experience-sensitive learning: The system will show monotonic or near-monotonic improvement in GVD and GTA across repeated gameplay, similar to human trajectories, rather than noisy and transient gains. It will consolidate successful strategies into compact rules, avoiding the memory bloat that currently dilutes useful lessons.

  • Learn from causal mistakes: After a loss, the system will identify and revise the earliest contributing decision, preventing the recurrence of structurally similar failures (e.g., repeated endgame mobility collapse in Othello or near-boundary escapes in CircleCat) across consecutive episodes.

  • Provide interpretable behavioral diagnostics: The system will output process-level metrics (EAR, CPQ, OCI, MCI, EPD, EPC, CAR) alongside final outcomes, allowing researchers to verify that improvement is due to strategic shifts (e.g., solver alignment, mobility control, containment) rather than random variation or interface familiarity.

Sources

Related papers