GenAI-Powered Inference

arXiv:2507.03897 · cs.LG, stat.ME, stat.ML · Submitted 2026-08-16 · Read on arXiv

Kosuke Imai, Kentaro Nakamura

Harvard University

cs.LG, stat.ME, stat.ML

Submitted: 2026-08-16

Updated: 2026-08-18

Project page: https://gpi-pack.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 80/100

The gist: The paper introduces GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images.

Terminology

Summary

The paper introduces GenAI-Powered Inference (GPI), a statistical framework for both causal and predictive inference using unstructured data, including text and images. GPI leverages open-source Generative Artificial Intelligence (GenAI) models—such as large language models and diffusion models—not only to generate unstructured data at scale but also to extract low-dimensional representations that are guaranteed to capture their underlying structure. Applying machine learning to these representations, GPI enables estimation of causal effects while quantifying associated estimation uncertainty. Unlike existing approaches to representation learning, GPI does not require fine-tuning of generative models, making it computationally efficient and broadly accessible.

The paper illustrates the versatility of the GPI framework through three applications: (1) estimating the effects of Chinese social media censorship while adjusting for textual confounders, (2) isolating the impact of specific image features from that of other correlated features in the same image, and (3) assessing the persuasiveness of political rhetoric. An open-source software package is available for implementing GPI.

The key methodological insight is that GPI leverages GenAI models to generate unstructured data X and extract the corresponding internal representations R. For existing texts or images, GenAI models are prompted to reproduce the original content, enabling the recovery of consistent internal representations. Because GenAI models are configured to produce X as a deterministic function of R, the latter is guaranteed to contain all information necessary to generate the former. As a result, GPI eliminates the need to finetune R, unlike standard approaches that rely on generic pretrained embeddings such as BERT.

The use of R offers several important advantages over directly modeling the raw data X. First, R provides a substantially lower-dimensional representation of the original high-dimensional input. Second, it captures rich and relevant information distilled from the vast data used to train the GenAI model. Third, by relying on open-source GenAI models that operate deterministically, GPI promotes scientific replicability. Finally, GPI can further improve statistical efficiency by combining the results based on multiple GenAI models.

GPI applies machine learning to the internal representation R to estimate a deconfounder f(R), which serves as a low-dimensional approximation of the latent confounding features U. This step is critical because directly adjusting for the high-dimensional unstructured data X can violate the positivity assumption and lead to biased inference. While the deconfounder is not necessarily unique, the paper formally shows that any valid f(R) enables the nonparametric identification of causal effects when adjusted for together with the observed confounders Z.

GPI learns a low-dimensional deconfounder f(R) that is predictive of the outcome within each treatment group while remaining shared across treatment groups. By discarding components of R that predict treatment assignment alone, f(R) also makes the overlap assumption more likely to hold than standard approaches in the literature. The paper emphasizes the importance of carefully tuning the neural network architecture used in GPI: the network must be sufficiently expressive to preserve outcome-relevant information, yet parsimonious enough to remove irrelevant variation.

In the first application (Text as Confounder), the paper reanalyzes data from 75,324 Weibo posts collected by the Weiboscope project to investigate whether users whose posts are censored are more likely to face future censorship and whether they respond by engaging in self-censorship. The paper uses two open-source large language models, LLaMA3 with eight billion parameters and Gemma3 with one billion parameters, to reproduce each focal post and extract its internal representation. The paper estimates the average treatment effect on the treated (ATT) using double machine learning with two-fold cross-fitting, with standard errors clustered at the user level. In contrast to the original analysis, GPI finds that prior censorship significantly reduces users' subsequent posting activity, providing evidence of self-censorship. Moreover, GPI reveals a much greater effect of prior censorship on the likelihood of future censorship. GPI's full-sample estimates are substantially more efficient than those based on the matched sample, highlighting GPI's advantages in statistical efficiency.

In the second application (Image as Treatment), the paper uses a dataset of more than 9,000 protest images to estimate the causal effect of nighttime scenes on perceived violence. The paper reproduces each protest photo using two different versions of Stable Diffusion model (versions 1.5 and 2.1) and extracts their internal representations. The paper finds that the estimated effect of nighttime scenes on perceived violence is reduced by more than half once other image features are adjusted for through the GPI methodology. The estimates are similar across the two different versions of Stable Diffusion model. The internal representation of Stable Diffusion model is of substantially lower dimension (16,384 = 64 × 64 × 4) than the original representation (786,432 = 512 × 512 × 3).

In the third application (Structural Model of Texts), the paper reanalyzes a forced-choice conjoint experiment to evaluate the persuasiveness of different types of political rhetoric. The paper estimates a semiparametric version of the original structural model, which allows for greater flexibility in capturing complex relationships. The paper finds that appeal to authority is the most persuasive, while ad hominem attacks significantly reduce persuasiveness. These findings are consistent across GenAI models and each model's point estimates are highly correlated. The GPI estimates are substantially more precise than the original analysis results.

The paper concludes by noting that this work opens several promising avenues for future research, including developing methods to interpret the estimated deconfounder, extending GPI to discover effective treatment features from unstructured data, and naturally extending the GPI framework to other types of unstructured data, including audio and video.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Instead of using generic pre-trained embeddings (e.g., BERT, Sentence-BERT) that may discard outcome-relevant information, I will extract internal representations from generative models by prompting them to reproduce the original unstructured data (text/images). This guarantees the representation contains all information necessary to reconstruct the input.

Capability: The AI system can now perform causal inference with unstructured data without the risk of losing confounding information during representation learning. This is particularly valuable when the downstream task requires adjusting for latent confounders embedded in text or images.

Improvement: I will implement a neural network architecture (TarNet extension) that simultaneously estimates a low-dimensional deconfounder and outcome model by minimizing only the outcome prediction loss (Equation S5). Unlike DragonNet, which adds a treatment-prediction head that can exacerbate overlap violations, this approach discards treatment-only information.

Improvement: I will implement the optimal linear combination of estimates from multiple GenAI models (Proposition 2). By computing influence functions for each model's estimator and using the optimal weight formula, I can combine results to minimize asymptotic variance.

Improvement: I will extend parametric structural models (e.g., Bradley-Terry for pairwise comparisons) to semiparametric versions where the confounding function h(R) is estimated by a neural network rather than assumed to be linear or additive.

Improvement: I will implement a sensitivity analysis procedure that measures how much causal estimates change when additional text-based confounders are included. Using the delta method for GPI estimators (as opposed to bootstrap for matching methods), I can quantify this robustness with proper standard errors.

Improvement: I will apply the GPI framework to image data by using Stable Diffusion models' internal representations (64×64×4 latent space) rather than raw pixels (512×512×3). The convolutional neural network encoder further reduces dimensionality while preserving spatial structure.

Improvement: I will implement automated hyperparameter tuning (using Optuna) specifically for causal inference architectures, with the key principle that the deconfounder must be expressive enough to preserve outcome-relevant information but parsimonious enough to remove treatment-only variation.

Sources

Related papers