Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification

arXiv:2608.12340 · cs.CL · Submitted 2026-06-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification".

Jane: The paper was written by Keito Inoshita from Kansai University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Jane, have you seen this paper title? "Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification." That's a mouthful, but it's basically throwing down a gauntlet.

Jane: It really is, Tom. And honestly, the title tells you exactly where the authors land. They're saying that all this effort we put into making augmented text sound new and diverse? It might be completely missing the point.

Tom: Right, because when you're dealing with imbalanced data—like a dataset where one emotion shows up five hundred times more than another—you need to beef up those rare classes. And the obvious instinct is to use a big language model to generate fresh, varied examples.

Jane: Exactly. And the paper's title is the spoiler: that instinct might be wrong. The authors benchmarked eleven different augmentation methods across seven datasets, and the ones that just retrieved and reused existing text in a smart way beat the fancy LLM generation methods.

Tom: So the "class-structure preservation" part means keeping the shape of how each class actually looks in the data. And "diversity" means generating superficially new sentences. And the paper says preservation wins.

Jane: Which is a pretty big deal, because a lot of recent work has been riding the wave of "just ask the LLM to make more data." This paper is a reality check on that whole approach.

Tom: And it's not some tiny experiment either. They ran five seeds per method-dataset pair, used proper statistical tests, and even tested two different LLM families to make sure it wasn't just one model being weak.

Jane: The authors are from Kansai University, and they've made all the code public. So anyone can check their work. That's the kind of rigor that makes a benchmark like this actually useful to the field.

Tom: I love that they didn't just say "LLMs are bad." They said "LLMs are unnecessary here, and here's the mechanism why." That's the difference between a headline and a contribution.

Jane: And the mechanism is fascinating. It's not about the text being unique. It's about whether the generated text actually stays inside the region where that class's real examples live. If it drifts outside, you're just adding noise.

Tom: So the title is almost a thesis statement for the whole paper. Class-structure preservation beats diversity. We'll dig into how they actually proved that next.

Jane: And whether this means we should all just abandon LLM augmentation entirely, or whether there's a smarter way to combine the two. That's the question I want answered.

Abstract: Tom: So we're back with "Class-Structure Preservation Beats Diversity," and Jane, the abstract really sets the stage. They're claiming that every LLM-based method they tested was either statistically equivalent to or worse than a classical method called EmbSMOTE.

Jane: EmbSMOTE, for anyone just tuning in, is basically a text version of a classic oversampling trick. You take a rare class example, find its nearest neighbors in the embedding space, and interpolate between them to create a synthetic point. Then you grab the actual real text closest to that point.

Tom: So you're not generating new sentences at all. You're retrieving real ones that sit in the right neighborhood. And that approach beat six different LLM-based methods, including some pretty sophisticated ones.

Jane: The gap got bigger as the imbalance got worse. On GoEmotions-twenty-eight which has twenty-eight emotion classes and an imbalance ratio over five hundred, the gap reached about zero point zero six three in macro F1. That's a real difference in a hard task.

Tom: And they didn't just measure performance. They measured five different properties of the augmented text—uniqueness, vocabulary diversity, text length, label distribution—and correlated them with performance. And here's the kicker.

Jane: Uniqueness had essentially zero correlation with how well the classifier did. You can generate a thousand unique sentences and it doesn't help. But vocabulary diversity did correlate positively. So it's not "diversity is bad." It's "surface-level uniqueness is irrelevant."

Tom: That's such a clean result. It means the field has been optimizing for the wrong thing. People have been chasing "make the text sound different" when they should have been asking "does this text stay inside the class's actual distribution?"

Jane: And they also found that LLM-generated text tends to be longer than the original training data, and that length inflation actually hurt performance. Same with forcing the label distribution to be uniform—that also hurt.

Tom: Which is ironic, because a lot of augmentation methods explicitly try to balance the classes by generating equal amounts for each. And this paper says that's actively counterproductive.

Jane: The abstract frames it as "class-conditional structural fidelity." Which is a fancy way of saying: the augmented samples need to look like they could have been drawn from the same distribution as the real examples of that class.

Tom: And that's the lens we should use to evaluate any augmentation method going forward. Not "how creative is the output" but "how faithful is it to the class geometry."

Jane: They even introduced their own method, VoidGen, which tries to target sparse regions in the embedding space before generating. And it still didn't beat the simple retrieval approach. That's a strong signal that the problem isn't where you generate from—it's the generation itself.

Tom: We'll get into VoidGen and what it means for future work in a bit. But first, let's talk about what's actually on page one, because there's a lot of context there about why this benchmark was even necessary.

Page 1: Tom: We're still on "Class-Structure Preservation Beats Diversity," and page one does a great job of setting up the problem. Jane, why is imbalanced text classification such a persistent pain point?

Jane: Because most real-world text data is imbalanced. Think about customer support tickets—most are routine, a tiny fraction are urgent escalations. Or medical notes—most are normal, a few are critical. If you train a classifier on that, it just learns to predict the majority class and ignores the rare ones.

Tom: And the paper points out that on something like GoEmotions-twenty-eight an augmentation-free baseline can fall below zero point two zero macro F1 even with five thousand training examples. That's barely better than guessing for a twenty-eight-class problem.

Jane: Right. And the standard fix has been data augmentation. The classical methods—like EDA, which does synonym replacement and random word swaps—have been around for years. But then LLMs came along and everyone assumed they'd be strictly better because they can generate genuinely new text.

Tom: And the paper's whole point is that this assumption was never properly tested. They say, and I'm paraphrasing here, that no existing benchmark included an embedding-space SMOTE-style retrieval method as a reference. So people were comparing LLM methods against weak baselines and declaring victory.

Jane: That's the "controlled benchmark" part of the title. They wanted to create a fair playing field where the classical retrieval methods got their best shot, and then see if LLMs could actually beat them.

Tom: They also mention that some contemporaneous studies had already found LLM augmentation only helps when seed samples are extremely scarce. So the writing was on the wall, but nobody had done the systematic comparison.

Jane: And the paper's contribution is exactly that systematic comparison. Eleven methods, seven datasets, five seeds each, proper statistical testing. That's the kind of benchmark that lets you make claims with confidence.

Tom: They also introduce VoidGen on page one as a "methodological probe." It's not meant to be the best method—it's meant to test a hypothesis about whether targeting sparse regions in the embedding space before generation helps.

Jane: And spoiler alert: it doesn't really help. It performs comparably to EmbSMOTE on most datasets but trails on the most imbalanced one. So even a clever Locate-then-Decode approach can't overcome the fundamental issue.

Tom: Which brings us back to their core thesis. The issue isn't where you generate from. It's that generation itself doesn't guarantee class membership. Retrieval does, by construction.

Jane: Because when you retrieve a real text, you know it belongs to that class. When you generate text, you're hoping the LLM stays on-topic. And hope is not a great engineering strategy.

Tom: We'll get into the actual experimental setup and results in the next segment. But page one really sets the stage for why this benchmark was needed and what they found.

Jane: And I think the most exciting implication is that this could save people a lot of compute. LLM-based augmentation is expensive. If a simple retrieval method works as well or better, that's a huge practical win.

Results: Tom: Welcome back to our discussion of "Class-Structure Preservation Beats Diversity." We've set the stage, and now Jane, we need to talk about the actual numbers. What did they find?

Jane: The headline result is in Table four. On six of the seven datasets, the top performers were classical methods. EmbSMOTE specifically was best or tied-for-best on the imbalanced multi-class datasets like Emo, GoEmotions-thirteen and GoEmotions-twenty-eight.

Tom: And the LLM methods just couldn't keep up. LLM-Paraphrase, AugGPT, CoTAM, LLM2LLM, CIEGAD—none of them statistically matched the classical methods on those imbalanced datasets.

Jane: The statistical testing is important here. They used Welch's t-tests with five seeds per cell. So when they say a method is worse, it's not a fluke of one random seed. It's a consistent difference.

Tom: And they broke it down by class size, which is where it gets really interesting. The gap between LLM and classical methods isn't in the tail classes—the ones with almost no training data. And it's not in the head classes, where everyone does fine.

Jane: It's in the middle classes. The ones with between twenty and two hundred fifty examples. That's where classical methods like EmbSMOTE steadily accumulate signal by preserving the class geometry, while LLM methods inject noise and lose their edge.

Tom: That's such a precise finding. It tells you exactly where augmentation matters and where it doesn't. If a class has three examples, no method can save it. If a class has five hundred, you don't need augmentation. The sweet spot is in between.

Jane: And they also did the diversity correlation analysis we mentioned earlier. Uniqueness rate had essentially no correlation with macro F1. But vocabulary diversity—measured by type-token ratio—had a moderate positive correlation.

Tom: So it's not that diversity doesn't matter at all. It's that the kind of diversity that matters is vocabulary breadth, not sentence-level uniqueness. You want the augmented text to use a wide range of words, but you don't need each sentence to be novel.

Jane: They also found that LLM-generated text tends to be longer, and that length inflation hurt performance. Same with forcing label distributions to be uniform. Both of those are common LLM augmentation behaviors, and both are net-negative.

Tom: And then they did the LLM family sensitivity analysis. They replaced Llama-three point one-8B with Qwen3-8B and reran CIEGAD. And the results were... mostly the same.

Jane: On the saturated datasets, no difference. On the imbalanced ones, Qwen3 was a bit better—up to zero point zero five two macro F1 on GoEmotions-twenty-eight. But even with the stronger LLM, CIEGAD still couldn't beat EmbSMOTE.

Tom: That's the key result for me. Upgrading the LLM narrowed the gap but didn't close it. Which means the bottleneck isn't LLM quality. It's the generate-then-verify paradigm itself.

Jane: And they confirmed this with the training set size sweep. At ntrain of five hundred or one thousand CIEGAD actually beat EmbSMOTE. But at two thousand and five thousand the order flipped. So LLM augmentation has a niche—very low-resource settings—but it's not the universal solution people hoped for.

Tom: So the practical takeaway is pretty clear. If you have a few thousand examples per class, use retrieval-based oversampling. It's cheaper, faster, and works better.

Jane: And if you're in a truly low-resource setting, LLM generation might help. But you should verify that it's actually staying on-class before you trust it.

Tom: We'll wrap this up in the conclusion, but I think this paper is going to change how people approach augmentation. It's a much-needed reality check.

Conclusion: Tom: And we've reached the end of our discussion on "Class-Structure Preservation Beats Diversity: A Comprehensive Benchmark of Text Augmentation Methods for Imbalanced Text Classification." Jane, what's the one thing you want listeners to remember?

Jane: That the goal of augmentation isn't to create novel text. It's to create text that faithfully represents the class. And retrieval-based methods like EmbSMOTE do that by construction, while LLM generation can only hope to do it.

Tom: The paper showed that across seven datasets and eleven methods, classical retrieval consistently matched or beat LLM-based generation, especially as imbalance increased. And the gap was driven by middle-class dynamics, not tail classes.

Jane: And the diversity analysis proved that uniqueness doesn't predict performance. What matters is vocabulary breadth and staying inside the class distribution. That's a finding that should reshape how we evaluate augmentation methods.

Tom: They also introduced VoidGen as a probe, and even that couldn't beat simple retrieval on the hardest dataset. Which reinforces their central claim: class-structure preservation beats diversity.

Jane: The practical implication is huge. LLM-based augmentation is expensive—hours of GPU time per dataset. EmbSMOTE is nearly free. If it works as well or better, why wouldn't you use it?

Tom: And the authors were careful to note the limitations. They only tested 8B-class LLMs. A 70B model might behave differently. And they only looked at English text classification. So there's room for future work.

Jane: But as a benchmark, this is exactly what the field needed. A controlled, statistically rigorous comparison that lets us make informed decisions instead of chasing hype.

Tom: So we'll say goodbye to this paper with a sense of clarity. The next time someone tells you "just use an LLM to generate more data," you can ask them one question.

Jane: "Does the generated text stay inside the class distribution?" And if they can't answer that, you know what to do.

Tom: Thanks for joining us, and we'll see you next time on the show.

Keito Inoshita

Kansai University

cs.CL

Submitted: 2026-06-03

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

The gist: The paper presents a controlled empirical benchmark of 11 text augmentation methods plus a no-augmentation baseline, evaluated on seven public text classification datasets, to compare classical

Key concepts

Class-Structure Preservation
This concept refers to ensuring that augmented text remains within the original distribution or 'region' where real examples of a specific class exist. The paper argues that generated text must stay faithful to the class geometry rather than just being unique.
Text Augmentation
This is the process of creating synthetic data points to address imbalanced datasets in classification tasks. Methods range from simple techniques like synonym replacement (EDA) to complex methods like using large language models (LLMs) to generate entirely new sentences.
EmbSMOTE
This is a classical oversampling technique used for text. Instead of generating new sentences, it finds the nearest neighbors of a rare class example in embedding space and interpolates between them to create synthetic points, retrieving actual real text closest to that point.
Imbalanced Text Classification
This is the problem where one class (e.g., routine messages) appears far more frequently than another (e.g., urgent escalations). Training a classifier on such data often leads it to ignore the rare, but important, classes.

Terminology

Summary

The paper presents a controlled empirical benchmark of 11 text augmentation methods plus a no-augmentation baseline, evaluated on seven public text classification datasets, to compare classical perturbation, embedding-space retrieval, and LLM-based generation approaches for imbalanced text classification.

Core Research Question and Motivation: The authors note that no empirical NLP benchmark exists that includes embedding-space SMOTE-style retrieval (EmbSMOTE) as a reference method and that "a systematic NLP benchmark under controlled imbalance conditions, comparing classical perturbation, embedding-space retrieval, and modern LLM-based generation under a unified experimental protocol with sufficient seeds and statistical corrections to render negative findings credible, is notably lacking."

Experimental Design: The benchmark covers 11 methods × 7 public datasets × 5 seeds with datasets spanning SST-2, AG News, Emo, TREC, GoEmotions-13, DBpedia, and GoEmotions-28, spanning a spectrum from K = 2 to K = 28 classes and from imbalance ratio (IR) = 1.12 to IR = 527.67. All experiments use a single NVIDIA H100 NVL GPU, DistilBERT as the classifier, and an augmentation ratio of r=1.0 with ntrain=5,000. The total computational budget is approximately ∼400 GPU-hours on a single H100, of which more than 80% is consumed by LLM-based augmentation methods.

Main Findings:

  1. LLM-based methods are statistically equivalent or inferior to EmbSMOTE: All LLM-based augmentation methods are shown to be statistically equivalent or inferior to EmbSMOTE, with the gap widening as class imbalance increases and reaching ∆F1macro ≈ 0.063 on GoEmotions-28. None of the six LLM-based methods (LLM-Paraphrase, LCG, AugGPT, CoTAM, LLM2LLM, and CIEGAD) statistically matches the classical methods on these datasets.

  2. Three-tier performance structure: The first tier comprises the four classical methods (EDA, AEDA, BackTrans, and EmbSMOTE), which remain within the statistical equivalence region of EmbSMOTE and suffer at most one loss each. The second tier consists of LLM-based methods (LLM-Paraphrase, LCG, AugGPT, CoTAM, LLM2LLM) which collectively record at least four losses each and only two wins in total. The third tier contains CIEGAD, which reaches only 0/4/3 and thus approaches but does not surpass the classical tier without a single win.

  3. Class-structure preservation, not diversity, is the operative variable: The effective variable is not surface-level diversity but class-conditional structural fidelity, namely the degree to which augmented samples preserve the class-conditioned geometry of the training distribution. The uniqueness rate exhibits at most a weakly negative correlation with macro F1 (r= − 0.166), while EmbSMOTE shows the lowest uniqueness rate across datasets, between 58% and 61%, yet attains the best or tied-best performance on 6 of 7 datasets.

  4. LLM-specific artifacts are harmful: "Mean character length and forced label uniformity both yield negative correlations with macro F1 (r= − 0.335 and r= − 0.170, respectively), and both are characteristic artifacts of LLM-based augmentation: chat-aligned LLMs produce longer texts than the original training corpus, while methods such as LCG and CIEGAD explicitly steer the label distribution toward uniformity."

  5. Per-class analysis reveals the gap is concentrated in middle classes: "The gap between LLM and classical methods is concentrated not in tail classes nor in head classes but in the middle-class region (20–250 examples), where training signal exists yet is sparse. In that region, classical methods stably accumulate incremental signal by faithfully preserving within-class distributional geometry, whereas LLM-based methods contaminate gradient updates with off-boundary generations and lose their competitive edge."

  6. LLM family sensitivity: Replicating CIEGAD with Qwen3-8B instead of Llama-3.1-8B shows that upgrading the LLM does not close the gap to retrieval-based EmbSMOTE. Even at its best, CIEGAD-Qwen still falls below EmbSMOTE on all imbalanced multi-class datasets. The authors conclude the bottleneck of LLM-based augmentation is not LLM quality at the 8B scale, but rather the absence of class-structure preservation in the generate-then-verify formulation.

  7. Training set size sensitivity: "The method ranking is not constant along the curve: it reverses at approximately ntrain ≈2,000. For ntrain ≤1,000, CIEGAD achieves the best performance, slightly outperforming EmbSMOTE. By contrast, for ntrain ≥2,000, the order reverses: EmbSMOTE and AEDA take the lead, and the gap with CIEGAD widens as ntrain grows, reaching ∆=+0.076 macro F1 at ntrain =5,000."

VoidGen Methodological Probe: The authors introduce "a so-called Locate-then-Decode formulation, VoidGen, in which the conventional 'generate-then-verify' order followed by existing LLM-based augmentation is inverted: sparse regions in the sentence embedding space are identified prior to generation, and LLM decoding is conditioned on those targets via a learned projector. VoidGen performs comparably to EmbSMOTE on most datasets yet trails on the most imbalanced GoEmotions-28, reinforcing the conclusion that pre-generation void targeting alone is insufficient to guarantee structural fidelity. VoidGen also failed with a ZeroDivisionError on 4 of 5 seeds for GoEmotions-28, owing to numerical instability in the void detection procedure."

Practical Recommendations: In saturated settings where the unaugmented macro F1 exceeds 0.85 on binary or low-class-count tasks, the choice of augmentation method is empirically irrelevant, and the cheapest classical option suffices. For fine-grained imbalanced multi-class classification, EmbSMOTE should be adopted as the default, and LLM-based augmentation should be required to satisfy a high evidentiary standard before deployment.

Limitations: The authors note the LLMs evaluated are limited to two models of the 8B class (Llama-3.1-8B and Qwen3-8B), and conclusions may change with 70B-class or GPT-4-class models, all datasets are English text classification tasks, the classifier is fixed to DistilBERT, and the augmentation ratio is fixed at r=1.0.

Improvements for AI systems

Based on this paper, here are the specific improvements I can implement in AI systems:

Improvement: Implement EmbSMOTE (embedding-space SMOTE with nearest-neighbor retrieval) as the default augmentation strategy instead of LLM-based generation.

What the improved system can do:

  • Achieve up to +0.063 macro F1 improvement over LLM-based methods on highly imbalanced datasets (e.g., GoEmotions-28 with IR≈528)

  • Guarantee class membership by construction—every augmented sample is a real training example, eliminating label noise from generation drift

  • Operate with zero parameter updates and no LLM inference cost, reducing augmentation runtime from 1 hour to seconds per dataset

Sources

Related papers