When Can We Work in Embedding Space? What Text Embeddings Preserve
econ.EM, cs.CL, stat.ML
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so.
Terminology
Abstract
When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.
Sources
- Adventures in Demand Analysis Using AI
- Inference for Regression with Variables Generated by AI or Machine Learning
- From Unstructured Data to Demand Counterfactuals: Theory and Practice
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- CAREER: A Foundation Model for Labor Sequence Data
Related papers
- SLIM: Stochastic Learning and Inference in Overidentified Models
- High-dimensional censored MIDAS logistic regression for corporate survival forecasting
- Cross-Fitting-Free Debiased Machine Learning with Multiway Dependence
- Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities
- Causal Inference in Possibly Nonlinear Factor Models
- Mining Causality: AI-Assisted Search for Instrumental Variables