Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

arXiv:2609.13201 · cs.LG, cs.IT, math.IT, math.PR · Submitted 2026-08-14 · Read on arXiv

cs.LG, cs.IT, math.IT, math.PR

Submitted: 2026-08-14

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

The gist: Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs.

Terminology

Abstract

Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the main data points when the number of anomalies is small and is ``taken" over by the anomalies above a certain threshold. We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.

Related papers