LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning

arXiv:2608.12626 · cs.CL, cs.AI, cs.MA · Submitted 2026-08-12 · Read on arXiv

Yi Wu, Zhimin Hu

University of Chicago · University of Wisconsin-Madison

cs.CL, cs.AI, cs.MA

Submitted: 2026-08-12

Updated: 2026-08-14

Journal ref: Published at Reasoning and Planning for LLMs at ICLR 2025

Code: https://github.com/ethanyiwu/EpicStar

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: This paper introduces EpicStar, a framework for enhancing strategic reasoning in Large Language Models (LLMs) within long-horizon environments, using StarCraft II as the testbed.

Terminology

Summary

This paper introduces EpicStar, a framework for enhancing strategic reasoning in Large Language Models (LLMs) within long-horizon environments, using StarCraft II as the testbed. The authors argue that LLMs suffer from strategic drift in long-horizon tasks—where finite attention resources prevent the model from maintaining strategic coherence over thousands of steps, causing localized decisions [to] fail to sustain a coherent trajectory across reasoning.

The paper identifies that while LLMs have already achieved human-level reasoning in short-horizon settings, their performance often degrades in long-range sequential reasoning tasks. The authors argue that the failures are mainly due to the fact that the agent tends to progressively overfit to local observations and consequently loses sight of its global objectives. Existing approaches either rely on more prompting loops to compress information or introduce more rules to stabilize the reasoning, but the authors contend that robust strategic reasoning requires a transition from reactive prompting to structured reuse of past experience.

EpicStar integrates episodic memory with working memory to enable agents to learn memory as policy. The framework consists of several key components:

The agent continuously stores successful gameplay episodes in a structured memory bank. During inference, it retrieves relevant episodes, acting as a strategic heuristic to guide its reasoning. The retrieval mechanism uses binary search around the current game time and computes differences between current and past observations using two metrics: the number of items that changed (Ditem) and the number of values that changed (Dvalue). The top n memories are selected based on a weighted combination (α = 0.5, β = 0.5) with n = 3.

The agent maintains a working memory for tracking environmental changes through an observation queue that captures recent kmax observations with a frame interval of L (kmax = 4, L = 24). A dynamic gating mechanism balances reusing past actions directly with new reasoning for situational adaptation. This exploration-exploitation process (Algorithm 1) includes an exploration action queue that stores LLM-proposed actions for future interpolation into retrieved action sequences.

Retrieved episodes are fused via a context-aligned mechanism to provide high-level context and structure for ongoing reasoning. This involves bidirectional modulation: 1) modulation from working memory to episodic memory through instructions encouraging retrieved actions to be executable in the current scenario, and 2) modulation from episodic memory to working memory by prompting the LLM to generate a high-level strategy description from retrieved episodes.

The evaluation used TextStarCraft II as the interface, with the Chain of Summarization (CoS) method as the primary baseline. Experiments were conducted at difficulty levels 5 and 6 (corresponding to entry-level and above-average human performance). Four OpenAI models were tested: gpt-3.5-turbo, gpt-4-turbo, gpt-4o-mini, and gpt-4o. The agent played as Protoss against built-in Zerg opponents across various strategies (timing, rush, power, macro, air) on two maps (Abyssal Reef LE and Ever Dream LE). Episodic memory was bootstrapped from a rule-based agent collecting successful trajectories from 20 rounds against Levels 6 and 7 opponents, retaining episodes from five winning games, yielding a total of 4,592 episodes.

At Level 5, EpicStar achieved win rates of 67.5% (gpt-4o-mini) and 75.0% (gpt-4-turbo), compared to CoS's 60.0% (gpt-4-turbo). At Level 6, EpicStar achieved 30.0% (gpt-4o-mini) while CoS only reached 8.3% (gpt-3.5-Turbo)—EpicStar nearly doubled the win rate to 15% with the same model.

A critical finding is that our token consumption is only 14.5% of CoS when using the same models. Specifically, CoS consumed an average of 518,485.4 tokens per game, while EpicStar used only 75,039.7 tokens per game—nearly an order of magnitude lower.

EpicStar showed superior Average Population Utilization (APU)—0.7991 (gpt-3.5-turbo) and 0.8449 (gpt-4-turbo) versus CoS's 0.7608 and 0.7194—indicating more effective macro management.

The ablation study revealed that both components contribute to performance:

  • Without exploration: Win rate dropped from 67.5% to 60.0% at Level 5, and from 30.0% to 17.5% at Level 6

  • Without contextual fusion: Win rate dropped to 65.0% at Level 5, and to 12.5% at Level 6

The paper notes that strategy coherence becomes increasingly important at higher difficulty levels, where misaligned exploration significantly impairs performance. At Level 6, the gap widened substantially particularly in the Power, Rush, and Timing strategies, where the ablated agents' win rates drop close to zero while EpicStar retains a modest but non-trivial win rate.

A case study comparing EpicStar with and without exploration showed that EpicStar wins the game faster and achieves faster expansion, indicating it reasons about its current strategic situation in the game and utilizes the limited resources more efficiently.

The paper's three main contributions are:

  1. EpicStar framework: an LLM-based agentic framework that incorporates structured episodic memory to support long-horizon strategic reasoning, mitigating strategic drift by retrieving and adapting past trajectories to maintain coherence over time

  2. Situational modulation mechanism: a dynamic gating module that integrates episodic recall with working memory, and a contextual fusion step that translates retrieved episodes into strategic guidance

  3. Empirical evidence: demonstrating that both episodic and working memory are essential for strategic performance, with even a small number of high-quality episodic memories leading to marked improvements in win rate and adaptability

The authors argue that even a small, curated bank of past trajectories can act as an effective substitute for costly exploration at inference time, reframing episodic memory not merely as a storage mechanism but as an implicit, non-parametric policy that complements the LLM's own reasoning.

Limitations acknowledged include: limited empirical evidence on the extent to which this performance ceiling depends on the base model's capabilities, the episodic memory being bootstrapped from a relatively small set of victories, and uncharacterized performance scaling with larger or noisier memory banks. The authors also note the safety consideration of the risk of overfitting to opponent styles seen during memory collection.

The paper concludes that by integrating episodic and working memory systems, we enable agents to maintain coherent strategic trajectories and adapt dynamically to evolving game scenarios, achieving much lower token consumption, indicating that episodic memory can serve as an implicit mechanism for policy optimization. The research underscores the importance of human cognitive mechanisms in developing AI that can navigate and excel in complex strategic environments, bridging the gap between artificial intelligence and cognitive science.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement episodic memory banks as non-parametric policy layers — Store successful decision trajectories (not just raw data) and retrieve them via similarity metrics (e.g., weighted item/value changes) to guide current actions. This reduces token consumption by 85% compared to pure prompting loops while maintaining or improving task performance.

  2. Add dynamic gating between working and episodic memory — Use a gating mechanism that decides when to reuse past action sequences verbatim versus when to generate new reasoning. This balances exploitation of learned strategies with exploration of novel situations, preventing overfitting to local observations.

  3. Enable bidirectional contextual modulation — When retrieving past episodes, (a) inject instructions that retrieved actions must be executable in the current state, and (b) force the LLM to generate a high-level strategy summary from retrieved episodes before acting. This ensures retrieved memories are adapted, not blindly copied.

  4. Maintain a short-term observation queue with time-indexed retrieval — Track recent environmental states (e.g., last 4 frames with interval 24) and use binary search around the current time step to find relevant past episodes. This provides temporal locality and reduces retrieval noise in long-horizon tasks.

  5. Bootstrap episodic memory from rule-based agents — Pre-populate the memory bank with successful trajectories from a simpler, non-LLM agent (e.g., 20 games, retaining 4,500 episodes). This provides high-quality seeds without expensive LLM exploration, enabling immediate strategic coherence.

  6. Use token efficiency as a design constraint — Structure the reasoning pipeline to minimize per-step token usage (e.g., retrieve 3 memories, use compact summaries) rather than scaling context windows. This makes the system deployable in cost-sensitive or latency-constrained environments.

What the improved AI system can do:

  • Maintain strategic coherence over thousands of decision steps without performance degradation (e.g., win rate increases from 8.3% to 30% in hard tasks).

  • Operate with 7x lower token consumption than baseline methods while achieving higher task success.

  • Adapt to increasing difficulty levels by leveraging retrieved strategies rather than relying on more prompting or rules.

  • Balance between following proven action sequences and generating novel responses when the situation deviates from stored experiences.

  • Scale to long-horizon domains beyond games (e.g., robotics task planning, multi-step negotiation, autonomous driving) where maintaining global objectives is critical.

Sources

Related papers