Augmented Hypothesis Testing with Persona-Based LLM Simulations
cs.LG, cs.AI, stat.AP
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: Work accepted at COLM Workshop on Agent Behavior
License: http://creativecommons.org/licenses/by/4.0/
The gist: A/B testing requires large sample sizes, long timelines, and significant costs.
Terminology
Abstract
A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.
Sources
- GPT-4 Technical Report
- PPI++: Efficient Prediction-Powered Inference
- Non-clairvoyant Scheduling with Partial Predictions
- On Tradeoffs in Learning-Augmented Algorithms
- SimGym: Traffic-Grounded Browser Agents for Offline A/B Testing in E-Commerce
- Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- LLM Generated Persona is a Promise with a Catch
- SimAB: Simulating A/B Tests with Persona-Conditioned AI Agents for Rapid Design Evaluation
- Prediction-Powered Semi-Supervised Learning with Online Power Tuning
- Llama 2: Open Foundation and Fine-Tuned Chat Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks