HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning
cs.AI, cs.CL, cs.CY
Submitted: 2025-10-16
Updated: 2026-08-31
Terminology
Sources
- Relational inductive biases, deep learning, and graph networks
- ToMBench: Benchmarking Theory of Mind in Large Language Models
- Towards Measuring the Representation of Subjective Global Opinions in Language Models
- DreamCoder: Growing generalizable, interpretable knowledge with wake-sleep Bayesian program learning
- Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
- Ground(less) Truth: A Causal Framework for Proxy Labels in Human-Algorithm Decision-Making
- HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models
- WikiWhy: Answering and Explaining Cause-and-Effect Questions
- Language Models, Agent Models, and World Models: The LAW for Machine Reasoning and Planning
- CommunityLM: Probing Partisan Worldviews from Language Models
- MMToM-QA: Multimodal Theory of Mind Question Answering
- Lyfe Agents: Generative agents for low-cost real-time social interactions
- FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
- Simulating Society Requires Simulating Thought
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Machine Theory of Mind
- ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning
- A Roadmap to Pluralistic Alignment
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection