Total Variation Distance Estimation in Autoregressive Models
Eric Price, Kevin Tian, Zhiyang Xun, Yusong Zhu
cs.LG, cs.DS, stat.ME, stat.ML
Submitted: 2026-07-21
Comments: 39 pages, 11 figures, code is available at https://github.com/XunZhiyang/llm-tv-estimation
Code: https://github.com/XunZhiyang/llm-tv-estimation
License: http://creativecommons.org/licenses/by/4.0/
The gist: Modern LLM deployments use a number of implementation choices and inference optimizations (e.g., batching, custom kernels, and quantization) on top of fixed weights, so two engines serving "the same
Terminology
Abstract
Modern LLM deployments use a number of implementation choices and inference optimizations (e.g., batching, custom kernels, and quantization) on top of fixed weights, so two engines serving "the same model" can produce meaningfully different distributions. We study the problem of estimating the total variation (TV) distance between two length- n autoregressive distributions to additive error epsilon, under three access models. (1) Under sample access, we use (n squared K/epsilon 2) queries, where K is the maximum support of the next-token distribution. This improves upon the (n cubed m/epsilon 5) -query estimator of Meel et al. (2025), where m at least K is the total size of the token alphabet. (2) Under logit access, we use O(n/epsilon 2) queries, and this is tight. (3) Under noisy logit access, we smoothly interpolate between the above two guarantees: if probability values are given to relative error sigma, we use ((n+n 2 sigma 2)/epsilon 2) queries. We complement our theoretical results with an empirical evaluation of our algorithms, for example measuring the distance between SGLang and vLLM serving identical weights. Our experiments highlight the robustness and practicality of estimating the total variation distance, which remains estimable where the KL divergence is infinite. Our code is available at https://github.com/XunZhiyang/llm-tv-estimation.
Sources
- A Chasm Between Identity and Equivalence Testing with Conditional Queries
- Tight simulation of a distribution using conditional samples
- Improved Bounds for High-Dimensional Equivalence and Product Testing using Subcube Queries
- Better Estimation of the Kullback--Leibler Divergence Between Language Models
- A short note on an inequality between KL and TV
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- On Distribution Testing in the Conditional Sampling Model
- Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks