Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
cs.AI, cs.MA, cs.SI
Submitted: 2026-09-21
Updated: 2026-09-22
Comments: Accepted for publication in the Journal of Artificial Societies and Social Simulation (JASSS). 43 pages (34 main text, 9 supplementary information), 4 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check.
Terminology
Abstract
Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.
Sources
- Out of One, Many: Using Language Models to Simulate Human Samples
- Calibrating Agent-based Models to Microdata with Graph Neural Networks
- Bayesian Workflow
- Automatic Posterior Transformation for Likelihood-Free Inference
- A Trust Crisis In Simulation-Based Inference? Your Posterior Approximations Can Be Unfaithful
- G-Sim: Generative Simulations with Large Language Models and Gradient-Free Calibration
- Agent-Based Model Calibration using Machine Learning Surrogates
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- Sampling-Based Accuracy Testing of Posterior Estimators for General Inference
- Fast $\epsilon$-free Inference of Simulation Models with Bayesian Conditional Density Estimation
- Generative Agents: Interactive Simulacra of Human Behavior
- BayesFlow: Learning complex stochastic models with invertible neural networks
- Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
- Whose Opinions Do Language Models Reflect?
- Validating Bayesian Inference Algorithms with Simulation-Based Calibration
- Robust Neural Posterior Estimation and Statistical Model Criticism
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection