DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations

arXiv:2405.07780 · cs.LG, cs.AI, cs.CV · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Variations".

Jane: The paper was written by Zhiyong Yang, Qianqian Xu, Sicong Li, Zitai Wang, Xiaochun Cao et al. from University of Chinese Academy of Sciences and Institute of Computing Technology, Chinese Academy of Sciences and Institute of Information Engineering, Chinese Academy of Sciences and Sun Yat-sen University and Peng Cheng Laboratory.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a fresh arXiv paper that’s got a real mouthful of a title: "DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations." Jane, I’m going to need you to break that down for me, because I’m already lost at "Test Agnostic."

Jane: Ha, happy to, Tom. So, imagine you train a model to recognize cats and dogs, but most of your pictures are of cats. That’s a long-tail problem—a few common classes and lots of rare ones. Usually, we test the model on a balanced set, but this paper says, what if the test data is also skewed? What if you suddenly get way more dogs than cats? That’s the "test agnostic" part—we don’t know what the test distribution will look like.

Tom: Okay, so it’s not just about training on imbalanced data, but being ready for any imbalance at test time. That sounds like a much harder problem.

Jane: Exactly. And the "Dir" in DirMixE stands for Dirichlet, which is a fancy way to describe randomness in probabilities. The authors use it to model all the possible test distributions we might encounter.

Tom: So they’re not just picking a few fixed scenarios, they’re sampling from a whole universe of possibilities?

Jane: Precisely. They break the variations down into "global" and "local." Global is the big swing—like going from mostly cats to mostly dogs. Local is the small jitter around those big swings. And their method, DirMixE, uses a team of experts, each one trained to handle a specific slice of that universe.

Tom: A team of experts, each with their own specialty. So instead of one model trying to do everything, you’ve got a specialist for the "mostly cats" world, one for the "balanced" world, and one for the "mostly dogs" world.

Jane: You got it. And the clever part is how they combine those experts at test time, using a self-supervised trick to figure out which expert to trust for the unknown distribution they’re seeing.

Tom: That’s a really elegant way to think about it. It’s not just about making one model robust; it’s about building a system that can adapt on the fly. I’m curious to see how they actually pull that off.

Jane: Me too. The paper’s got a lot of math, but the core idea is super intuitive. It’s about being prepared for the unexpected, which is a great mindset for any real-world application.

Tom: Alright, we’ve got the gist. Let’s take a quick break, and when we come back, we’ll dig into the summary and see how they actually built this thing.

Summary: Jane: Welcome back. We’re still on "DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations." Tom, we talked about the problem, but the paper’s summary really lays out the solution’s architecture.

Tom: Right, and the first thing that jumps out is that they’re not just using one model. They’re using a Mixture of Experts, which we touched on. But the key is how they assign those experts.

Jane: Instead of assigning an expert to a single, fixed distribution, they assign each expert to a *range* of distributions. They do this by defining a "meta-distribution," which is like a distribution of distributions. Each expert is trained to handle a specific part of that meta-distribution.

Tom: So it’s like having a specialist for "slightly skewed towards cats," not just "all cats." That way, when the test data comes in with some random skew, there’s an expert that’s a good fit.

Jane: Exactly. And they don’t just average the experts’ outputs. They use a self-supervised method to learn weights for each expert based on the test data itself. So if the test data looks like a uniform distribution, the "uniform" expert gets more weight.

Tom: That’s the "test-agnostic" part in action. The model is figuring out the test distribution on its own, without being told.

Jane: And the results back this up. On standard benchmarks like CIFAR-one hundred and ImageNet, their method, DirMixE, consistently beats other state-of-the-art approaches, especially when the test distribution is heavily skewed in an unexpected direction.

Tom: So it’s not just a theoretical idea; it’s actually performing better in practice. That’s always good to hear.

Jane: Definitely. They also show that their method is more stable. The performance doesn’t swing wildly between different test distributions, which is a huge plus for real-world reliability.

Tom: Stability is key. If a model is great on one test set and terrible on another, you can’t trust it. This paper seems to be tackling that head-on.

Jane: They are. And they’ve also extended this to work with foundation models, which is a big deal. They’re not just training from scratch; they’re fine-tuning large pre-trained models efficiently.

Tom: Oh, that’s interesting. So they’re making this work with the big models everyone’s using now. How do they manage that without retraining the whole thing?

Jane: That’s the next part of the paper. They have a clever method called Latent Skill Finetuning, or LSF, that we should get into.

Tom: Sounds like a perfect segue. Let’s talk about that after the break.

Improvements: Tom: We’re back with "DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations." Jane, you mentioned Latent Skill Finetuning, or LSF. What’s the big idea there?

Jane: So, foundation models are huge. Retraining them for a new task is expensive. LSF is a way to fine-tune them using only a tiny fraction of the parameters. Think of it like having a shared toolbox of skills, and each expert just picks the tools it needs.

Tom: A shared toolbox. So instead of each expert having its own completely separate set of updates, they all pull from a common pool of "latent skills."

Jane: Exactly. They implement this using popular methods like LoRA and Adapters. These are ways to add small, trainable matrices to the frozen model. The trick here is that these small matrices for different experts are built from a shared set of base skills.

Tom: So you’re getting the diversity of multiple experts, but you’re not paying the full cost of training them all from scratch. That’s a huge efficiency win.

Jane: It is. And they also have a clever initialization scheme for these skills, based on the singular values of the original model. It helps the training start off on the right foot, avoiding some of the pitfalls of starting from zero.

Tom: So they’re not just randomly initializing these new parameters; they’re using the structure of the pre-trained model to give them a head start.

Jane: Right. And the paper also has a lot of theory to back this up. They prove that their method has good generalization guarantees, which means it should work well on data it hasn’t seen before.

Tom: That’s important. It’s one thing to show it works on benchmarks, but having a theoretical foundation gives you more confidence it’ll hold up in the wild.

Jane: And their theory is pretty sharp. They show that the generalization error depends on the number of trainable parameters, not the total size of the model. That’s a strong statement.

Tom: So the efficiency isn’t just a practical trick; it’s theoretically sound. That’s a nice combination.

Jane: It is. They’ve really thought this through, from the high-level concept down to the implementation details. I’m impressed by the depth of the work.

Tom: Me too. It feels like they’ve covered all the bases. Let’s bring in the rest of the team to get their take on the impact of this.

Conclusion: Tom: Alright, we’re wrapping up our discussion on "DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Vartiations." Jane, what’s the final takeaway for our listeners?

Jane: The big takeaway is that this paper gives us a robust way to handle the messy, unpredictable nature of real-world data. It’s not just about training on a skewed dataset; it’s about being ready for any skew you might see later. And they do it efficiently, even with huge foundation models.

Tom: And it’s not just a clever idea. They’ve shown it works on major benchmarks and backed it up with solid theory. That’s a complete package.

Jane: Absolutely. This could be really impactful for any application where the data distribution might shift over time—like medical diagnosis, where disease prevalence changes, or fraud detection, where patterns evolve.

Tom: That’s a great point. It’s about building systems that are resilient, not just accurate on a static test set. This paper is a step towards that.

Jane: We’ll be keeping an eye on this line of research. Thanks for joining us, and we’ll see you next time.

Zhiyong Yang, Qianqian Xu, Sicong Li, Zitai Wang, Xiaochun Cao, Qingming Huang

University of Chinese Academy of Sciences · Institute of Computing Technology, Chinese Academy of Sciences · Institute of Information Engineering, Chinese Academy of Sciences · Sun Yat-sen University · Peng Cheng Laboratory

cs.LG, cs.AI, cs.CV

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: Conference version: Zhiyong Yang, Qianqian Xu, Zitai Wang, Sicong Li, Boyu Han, Shilong Bao, Xiaochun Cao, and Qingming Huang. Harnessing Hierarchical Label Distribution Variations in Test Agnostic Long-tail Recognition. ICML, 56624-56664, 2024

DOI: 10.1109/TPAMI.2025.3647124

Project page: https://xcurveopt.github.io

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 56/100

The gist: This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced.

Key concepts

Test Agnostic
This refers to the challenge where the distribution of data used for testing is unknown or heavily skewed. Instead of assuming a balanced test set, DirMixE is designed to handle any possible imbalance in real-world scenarios.
DirMixE
The proposed method uses a team of experts, each trained to handle specific slices of the entire universe of possible data distributions. It employs a self-supervised trick to determine which expert is most appropriate for the unknown test distribution.
Mixture of Experts (MoE)
This architecture involves multiple specialized models or 'experts.' Instead of one model doing everything, it uses a system where different experts are assigned to specific ranges of data distributions, allowing the system to adapt on the fly.

Terminology

Summary

This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced. The authors argue that the variation in these distributions can be broken down hierarchically into global and local levels. Global variations reflect a broad range of diversity, while local variations typically arise from milder changes, often focused on a particular neighbor.

Traditional methods predominantly use a Mixture-of-Expert (MoE) approach, targeting a few fixed test label distributions that exhibit substantial global variations. However, the local variations are left unconsidered. To address this issue, the authors propose a new MoE strategy, DirMixE, which assigns experts to different Dirichlet meta-distributions of the label distribution, each targeting a specific aspect of local variations. Additionally, the diversity among these Dirichlet meta-distributions inherently captures global variations. This dual-level approach also leads to a more stable objective function, allowing better sampling of different test distributions to quantify the mean and variance of performance outcomes.

Building on this idea, the authors develop a general Latent Skill Finetuning (LSF) framework for parameter-efficient finetuning of foundation models, providing implementations based on LoRA and Adapter. Theoretically, they derive upper bounds on the generalization error for both standard learning and PEFT. Under mild assumptions, they show that variance-based regularization helps tighten these bounds. Furthermore, they prove that the covering number of the PEFT hypothesis class scales with the number of trainable parameters. Finally, extensive experiments on CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist validate the effectiveness of DirMixE.

The paper extends the authors' ICML 2024 paper, with major improvements including: a new method for foundation models (LSF), new theoretical results with tighter generalization bounds, and new experiments evaluating LSF on multiple datasets.

The paper is structured as follows: Section 2 reviews related work on loss modification, experts ensembling, and test-agnostic long-tail recognition. Section 3 formulates the problem, introducing the hierarchical sampling process where test label distributions are sampled from a meta-distribution. Section 4 details DirMixE for traditional deep learning models, including the construction of the meta-distribution as a mixture of Dirichlet distributions, the mixture-of-experts strategy, and the empirical approximation of the objective function using Monte Carlo methods with semi-variance regularization. Section 5 introduces the LSF framework for fine-tuning foundation models with LoRA and Adapter implementations, including an improved initialization regime. Section 6 provides theoretical analysis, including upper bounds on the semi-variance/variance ratio, generalization bounds for the induced subclass, and generalization analysis for DirMixE-LSF. Section 7 presents extensive experiments, and Section 8 concludes the paper.

Key theoretical contributions include: Theorem 1 provides lower bounds on the ratio of variance to semi-variance for exponential, Gamma, and Pareto distributions; Theorem 2 extends this to mixture distributions; Theorem 3 provides a generalization bound for the induced subclass; Theorem 4 provides a generalization bound for the LSF scheme; and Theorem 5 analyzes model averaging error at test time.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do.

1. Hierarchical Test-Distribution Modeling for Robustness

  • Improvement: Replace the common assumption of a fixed or balanced test label distribution with a hierarchical probabilistic model. The system will sample test label distributions from a Dirichlet Mixture Model (with forward, uniform, and backward components), capturing both global (inter-component) and local (intra-component) variations.

  • What the improved AI system can do: It can maintain high accuracy not just on balanced or long-tailed test sets, but also on unseen, arbitrarily imbalanced distributions (e.g., inverse long-tail where rare classes become common). This is critical for deployment in dynamic environments where class priors shift over time.

2. Mixture-of-Experts (MoE) with Variance-Aware Training

  • Improvement: Implement a MoE architecture where each expert is trained on a specific Dirichlet component (e.g., one for head-heavy, one for uniform, one for tail-heavy distributions). The training objective is not just to minimize average loss, but to minimize the mean and semi-variance of the loss across sampled test distributions. This prevents over-optimization on easy distributions and stabilizes performance across the entire spectrum.

  • What the improved AI system can do: It will exhibit lower variance in performance across different test-time label distributions. The system will be more reliable, avoiding catastrophic failures on tail classes even when the test distribution is heavily skewed towards them.

3. Parameter-Efficient Fine-Tuning (PEFT) with Latent Skill Sharing (LSF)

  • Improvement: For foundation models (e.g., CLIP), implement a novel PEFT framework called Latent Skill Finetuning (LSF). Instead of training each expert independently, all experts share a common set of latent skills (low-rank matrices or adapters). Each expert's update is a learned linear combination of these shared skills, promoting knowledge transfer and reducing the number of trainable parameters.

  • What the improved AI system can do: It can be fine-tuned for long-tail recognition with significantly fewer trainable parameters (e.g., via LoRA or Adapters) while achieving better or comparable performance to full fine-tuning. This makes it feasible to adapt large models on limited hardware and with smaller datasets.

4. Improved Initialization for Stable Training (SVD-based)

  • Improvement: For the LoRA-based LSF, replace the standard Kaiming/zero initialization with a principled SVD-based initialization. The pretrained weight matrix is decomposed, and its singular values/vectors are reallocated to the shared latent skills and a residual matrix. This ensures that the initial gradient flow is non-zero for all parameters, preventing slow or stalled training in early epochs.

  • What the improved AI system can do: It will converge faster and more reliably during fine-tuning. The system avoids the dead neuron problem in early training, leading to better final performance, especially when the number of training epochs is limited.

5. Test-Time Self-Supervised Expert Aggregation

  • Improvement: Implement a test-time aggregation mechanism that learns the optimal weights for combining expert outputs without needing labeled test data. This is done via self-supervision, which maximizes mutual information between predictions and the unknown test distribution.

  • What the improved AI system can do: It can dynamically adapt its ensemble weights at inference time. For a given test batch, it will automatically assign higher weight to the expert best suited for the current (unknown) label distribution, leading to superior performance without any manual tuning or access to test labels.

6. Sharp Generalization Guarantees via Semi-Variance Regularization

  • Improvement: Theoretically, the system's generalization error is bounded by a function that includes the empirical semi-variance of the loss. By actively minimizing this term during training, the system is directly optimizing an upper bound on its true generalization error.

  • What the improved AI system can do: It provides a provable guarantee that its performance on unseen test distributions will be close to its training performance, with a tighter bound than standard methods. This offers a higher level of trust and predictability for safety-critical applications.

Abstract

This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and arbitrarily imbalanced. We argue that the variation in these distributions can be broken down hierarchically into global and local levels. The global ones reflect a broad range of diversity, while the local ones typically arise from milder changes, often focused on a particular neighbor. Traditional methods predominantly use a Mixture-of-Expert (MoE) approach, targeting a few fixed test label distributions that exhibit substantial global variations. However, the local variations are left unconsidered. To address this issue, we propose a new MoE strategy, DirMixE, which assigns experts to different Dirichlet meta-distributions of the label distribution, each targeting a specific aspect of local variations. Additionally, the diversity among these Dirichlet meta-distributions inherently captures global variations. This dual-level approach also leads to a more stable objective function, allowing us to sample different test distributions better to quantify the mean and variance of performance outcomes. Building on this idea, we develop a general Latent Skill Finetuning (LSF) framework for parameter-efficient finetuning of foundation models. We provide implementations based on LoRA and Adapter. Theoretically, we derive upper bounds on the generalization error for both standard learning and PEFT. Under mild assumptions, we show that the variance-based regularization helps tighten these bounds. Furthermore, we prove that the covering number of the PEFT hypothesis class scales with the number of trainable parameters. Finally, extensive experiments on CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist validate the effectiveness of DirMixE.

Related papers