Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri
JPMorganChase
cs.AI
Submitted: 2026-08-20
Updated: 2026-08-24
Comments: Preprint for conference full paper at 20th ACM Conference on Recommender Systems (RecSys '26), Minneapolis, MN, USA
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 72/100
The gist: The paper proposes an Inverse Theory of Mind (IToM) pipeline for content recommendation that reasons backward from observed user interactions to infer the beliefs, preferences, and decision-making
Terminology
Summary
The paper proposes an Inverse Theory of Mind (IToM) pipeline for content recommendation that reasons backward from observed user interactions to infer the beliefs, preferences, and decision-making traits that explain behavior, producing structured natural-language personas without requiring qualitative data such as interviews or surveys.
The pipeline consists of five stages: (1) Summarize—converting raw interaction events into natural language sentences; (2) Perceive—reconstructing the decision context by analyzing page content alongside action summaries, producing a structured percept with main choices, sub choices, other choices, user choice, and perception description; (3) Infer—performing backward belief inference at both action-level and session-level, generating multiple candidate belief statements grounded in the reconstructed percept; (4) Optimize—deduplicating and selecting a compact, diverse belief set via Maximal Marginal Relevance (MMR); and (5) Synthesize—generating multiple diverse persona hypotheses and aggregating them with confidence-weighted averaging.
The authors evaluate on the OPeRA dataset, which pairs fine-grained Amazon shopping interaction traces with ground-truth personality assessments, attitudinal surveys, and interview transcripts. They validate across four tasks: next action prediction, shopping category prediction, Big Five personality inference, and shopping preference survey alignment.
Key results include: In next action prediction on Claude Opus 4.6, the inferred persona substantially outperforms the ground-truth persona on action generation accuracy (25.69% vs. 17.06%, +50.6% relative), action type weighted F1 (83.90% vs. 78.20%), and click type weighted F1 (49.78% vs. 43.16%). For shopping category prediction, the multi-hypothesis persona achieves NDCG@5 of.701 (+6.5% over the best baseline) with Hit@3 of 100%. For Big Five personality prediction, the method achieves an overall MAE of 0.762, outperforming all non-oracle baselines including population norms (0.836, −8.8% relative improvement) and closing 88% of the gap from random guessing to the oracle mean. For shopping preference survey alignment, the method achieves 76.6% normalized accuracy (MAE = 0.937 on a 1–5 Likert scale).
The ablation study reveals what the authors term persona bias
: single-interpretation methods (direct inference and questionnaire administration) produce negative average Pearson correlations with ground truth (−0.103 for questionnaire, −0.008 for direct inference), meaning their predictions are anti-correlated with actual personality. Multi-hypothesis reasoning reverses the average correlation from negative to +0.103 and cuts MAE from 0.945 to 0.762.
The authors also present a case study demonstrating cross-modal transferability with a persona-driven spatial banking application on VisionOS, where personas inferred from 2D browsing data inform 3D layout and content selection. The paper concludes that inferred personas match or exceed ground-truth personas, multi-hypothesis reasoning is essential for accurate personality prediction, and the natural-language persona representation transfers directly to generative UI engines or spatial XR interfaces without retraining.
Improvements for AI systems
Improvements to AI Systems:
-
Add Inverse Theory of Mind (IToM) as a core reasoning module – Instead of relying solely on forward prediction (input → output), the AI system will first reconstruct the user's decision context (perceived choices, constraints, alternatives) from raw interaction logs, then perform backward inference to generate multiple candidate belief statements about the user's preferences, goals, and cognitive biases. This turns the AI from a pattern matcher into a causal explainer of user behavior.
-
Implement multi-hypothesis persona generation with confidence-weighted aggregation – The system will no longer output a single user profile. It will generate 5–10 diverse persona hypotheses (each a structured natural-language description of beliefs, traits, and decision rules), deduplicate them via MMR, and then aggregate predictions (e.g., next action, category preference) by weighting each hypothesis by its self-reported confidence. This eliminates
persona bias
(anti-correlation with ground truth) seen in single-interpretation methods. -
Enable cross-modal persona transfer without retraining – Because personas are expressed as structured natural language (not embeddings or task-specific features), the same inferred persona can be directly fed into any downstream generative UI, recommender, or spatial interface (e.g., VisionOS app). The AI system will use the persona text as a prompt to condition layout, content selection, and interaction style—no fine-tuning or feature engineering needed per modality.
-
Add a
perception reconstruction
preprocessing layer – Before any inference, the AI will reconstruct the user's decision context from raw events (e.g., page content + click sequence) into a structured percept: main choices, sub-choices, other choices, user choice, and a perception description. This forces the AI to model what the user saw and considered, not just what they did, improving accuracy on next-action and category prediction. -
Replace questionnaire-based personality assessment with behavioral inference – The AI will infer Big Five traits (openness, conscientiousness, etc.) directly from interaction traces using the IToM pipeline, achieving MAE of 0.762 (vs. 0.836 for population norms) and reversing negative correlations to positive (+0.103). This removes the need for intrusive surveys and enables real-time, passive personality modeling.
-
Implement action-level and session-level dual inference – The AI will reason backward at two granularities: (a) per-action beliefs (why did the user click this item?) and (b) session-level beliefs (what overall goal or preference pattern explains the sequence?). This dual-layer reasoning improves robustness and yields more grounded, less hallucinated persona statements.
What the Improved AI System Can Do:
-
Explain user behavior in natural language – Given a user's browsing/click history, it outputs a coherent persona like:
This user prioritizes product durability over price, tends to compare 3–4 alternatives before deciding, and is highly conscientious but low in openness to new brands.
This is directly usable by human analysts or downstream systems. -
Predict next actions with 50.6% higher accuracy than using ground-truth personas – The system outperforms even the
true
personality profile for action generation (25.69% vs. 17.06% accuracy), meaning its inferred persona is more predictive of future behavior than a survey-derived one. -
Predict shopping categories with 100% Hit@3 and NDCG@5 of 0.701 – It can rank the most likely product categories a user will explore next, even when the user's preferences are implicit.
-
Infer Big Five personality traits without any self-report – Achieves MAE of 0.762 on a 1–5 scale, closing 88% of the gap from random guessing to the oracle (survey-based) mean, and reverses the anti-correlation problem of single-shot methods.
-
Align with user survey preferences at 76.6% normalized accuracy – It can predict how a user would answer Likert-scale preference questions (e.g.,
I prefer quality over price
) purely from their interaction data. -
Transfer the same persona to a 3D spatial interface – For example, in a VisionOS banking app, the inferred persona from 2D web browsing automatically determines which financial products to show, how to arrange them spatially (e.g., conservative users get a calm, single-column layout; exploratory users get a dynamic, multi-panel grid), and what language tone to use—without any retraining.
-
Avoid persona bias in real-time – By always maintaining multiple hypotheses and weighting them, the system never collapses into a single, potentially wrong interpretation, ensuring its predictions remain positively correlated with actual user traits even in noisy, sparse data.
Abstract
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.
Sources
- Turning large language models into cognitive models
- Generative Inverse Deep Reinforcement Learning for Online Recommendation
- Large Language Models for User Interest Journeys
- Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models
- Kernel Ridge Regression for Efficient Learning of High-Capacity Hopfield Networks
- Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
- LLM-Rec: Personalized Recommendation via Prompting Large Language Models
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- User-LLM: Efficient LLM Contextualization with User Embeddings
- LLMs achieve adult human performance on higher-order theory of mind tasks
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
- Multiagent Inverse Reinforcement Learning via Theory of Mind Reasoning
- Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection