Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations
cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns.
Terminology
Abstract
As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.
Sources
- Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces
- Learning to Simulate Human Dialogue
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Dialogue
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
- WhatIf: Interactive Exploration of LLM-Powered Social Simulations for Policy Reasoning
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants
- Simulating Human-like Daily Activities with Desire-driven Autonomy
- Human vs. Agent in Task-Oriented Conversations
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- ReAct: Synergizing Reasoning and Acting in Language Models
- Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection