Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction.
Terminology
Abstract
Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction. However, prior work demonstrated that LLMs struggle to use representations learned in context on a graph tracking task, where the model needs to construct a representation of the graph governing data generation process and use it for subsequent predictions. In this paper, we show that extending this to few-shot settings, where each demonstration is generated from a different world with either the same or different graph topologies, enhances its prediction on 6 models from 4 model families. To understand this improvement, we linearly probe a low-dimensional world representation that encodes graph information in the hidden states. Notably, we find that few-shot demonstrations relocate the world representation and increase its predictive use. Specifically, for each model, these world representations shift in directions nearly orthogonal to their original subspace, and interventions on these representations selectively impair performance more than interventions on other subspaces. Consistent with this insight, we show that few-shot demonstrations with observations from different worlds improve performance on ARC-AGI-1&2, web agent tasks, and Othello. Our findings elucidate the role and internal mechanisms of few-shot demonstrations in in-context world modeling. More broadly, our work advances our understanding of how LLM agents learn from in-context observations and provides implications for their further improvement.
Sources
- Understanding intermediate layers using linear classifier probes
- LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
- ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
- On the Measure of Intelligence
- Mind2Web: Towards a Generalist Agent for the Web
- Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals
- AI Must Embrace Specialization via Superhuman Adaptable Intelligence
- The Llama 3 Herd of Models
- World Models
- ARC Is a Vision Problem!
- Language Models (Mostly) Know What They Know
- Emergent World Models and Latent Variable Estimation in Chess-Playing Language Models
- When can transformers compositionally generalize in-context?
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- What Does it Mean for a Neural Network to Learn a "World Model"?
- Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
- Ministral 3
- TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning
- Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
- The Hydra Effect: Emergent Self-repair in Language Model Computations
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection