UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval

summary

Video file (mp4)

The gist

The paper addresses the challenge of Unsupervised Domain Adaptation (UDA) for neural retrievers.

In short

The episode discusses 'UnIte,' a paper from Seoul National University about uncertainty-based iterative document sampling for domain adaptation in information retrieval. The method filters out noisy documents and actively seeks samples where the model is most uncertain about the new domain's knowledge. This approach leads to significant retrieval quality gains while being more efficient.

Key concepts

Aleatoric uncertainty
This refers to noise within the data itself, such as a document that is off-topic or just unusual. The method filters these out because they do not teach the model anything useful.
Epistemic uncertainty
This represents what the model does not yet know. The authors use this type of uncertainty to actively seek out documents that are on-topic but confusing to the model, treating them as valuable learning opportunities.
Iterative Document Sampling
The process involves repeatedly training the model and re-sampling documents based on its changing uncertainty. This allows the model to constantly check its own knowledge gaps and learn more effectively over time.

Terminology used across episodes

This episode discusses

The paper

UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval · Read on arXiv

Jongyoon Kim, Minseong Hwang, Seung-won Hwang

Seoul National University

Unsupervised domain adaptation generalizes neural retrievers to an unseen domain by generating pseudo queries on target domain documents. The quality and efficiency of this adaptation critically depend on which documents are selected for pseudo query generation. The existing document sampling method focuses on diversity but fails to capture model uncertainty. In contrast, we propose **Un**certainty-based **Ite**rative Document Sampling (UnIte) addressing these limitations by (1) filtering documents with high aleatoric uncertainty and (2) prioritizing those with high epistemic uncertainty, maximizing the learning utility of the current model. We conducted extensive experiments on a large corpus of BEIR with small and large models, showing significant gains of +2.45 and +3.49 nDCG@10 with a smaller training sample size, 4k on average.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval".

Jane: The paper was written by Jongyoon Kim, Minseong Hwang and Seung-won Hwang from Seoul National University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we are digging into a paper that has a pretty dense title, but a really practical problem at its core. It’s called “UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval.” Jane, I’m going to lean on you to unpack that title for our listeners.

Jane: Happy to, Tom. So imagine you’ve trained a search engine on one set of documents, like general web pages. Now you want it to work on a completely different set, say medical research papers. That’s domain adaptation. The trick is, you don’t have labeled queries for the new domain, so you have to generate fake ones. But you can’t generate them for all hundred thousand documents, so you have to pick which ones to train on.

Tom: And that picking part is where things get interesting. The old way was basically random, or just picking diverse documents. But this paper from Seoul National University says that’s not enough. They argue you need to think about uncertainty.

Lu: Exactly, Tom. And I love how they frame it. They split uncertainty into two types. Aleatoric uncertainty is the noise in the data itself, like a document that’s off-topic or just weird. Epistemic uncertainty is what the model doesn’t know yet. The authors argue you want to filter out the first kind and actively seek out the second.

Jane: Right, so if a document is just an outlier, it’s not going to teach the model anything useful. But if the model is unsure about a document that’s clearly on-topic, that’s a learning opportunity. That’s the sweet spot they’re aiming for.

Tom: And they’re not just doing this once. They’re doing it iteratively. The model learns, its uncertainty changes, and then they re-sample. It’s like the model is constantly checking its own blind spots.

Meng: I’m already thinking about the engineering side. The paper reports some pretty solid gains, like over three points of nDCG@ten on average with the larger model. That’s a real jump in retrieval quality, and they’re doing it with fewer training samples than the baseline.

Tom: So it’s not just smarter, it’s more efficient. That’s a win-win. We’ll get into the nitty-gritty of how they measure that uncertainty in the next segment, because that’s where the real cleverness is.

Summary: Tom: So we’ve established that “UnIte” is about picking the right documents to train a search model on a new domain. Jane, can you walk us through the core idea of how they actually decide what’s a good document?

Jane: Sure. They use two filters. First, they get rid of the noisy outliers, the high aleatoric uncertainty stuff. They do this with a simple lexical check, like seeing if a document shares enough words with its neighbors. If it’s too far away, it’s probably off-topic and gets tossed.

Lu: And that’s a smart move because it’s model-free. They’re not using the retriever’s own embeddings to judge the data. That would be circular. They’re using a pure lexical distance, so they’re not accidentally throwing away documents that the model just hasn’t learned to represent yet.

Jane: Exactly. Then, for the documents that survive that filter, they measure epistemic uncertainty. This is the clever part. They don’t just look at the model’s confidence. They compare the model’s understanding of a document against the actual statistics of the target domain.

Tom: So they’re not asking “is the model sure?” They’re asking “does the model know what’s important in this new field?”

Jane: Precisely. They look at how well the model predicts the high-IDF terms, the rare, domain-specific words. If the model can’t predict those, that document is a high-value training sample.

Meng: And the iterative part is crucial. The paper shows that if you just sample once, you waste your budget on documents that become trivial after the first round of training. Their loop re-evaluates the model’s knowledge every time, so it’s always chasing the actual knowledge gaps.

Lu: The results back that up. On TREC-COVID, they get a four-point boost over the diversity-based baseline with a small model. And on Robust04, it’s even more dramatic, a five-point jump. It’s a clear signal that uncertainty-aware selection is far more effective than just picking for variety.

Tom: So it’s not just about finding hard examples, it’s about finding the right hard examples that are actually relevant to the new domain. That’s a much more sophisticated approach. Let’s talk about how they keep this from being a computational nightmare in the next segment.

Improvements: Tom: We’ve talked about the core idea of “UnIte,” but I want to get into the practical improvements it brings. Meng, you were looking at the efficiency angle. What stands out?

Meng: The biggest thing is the early stopping criterion. They monitor the average epistemic uncertainty across the whole domain. It drops as the model learns, and then it starts to rise again when the model is just seeing redundant samples. That’s the plateau point, and they stop right there.

Tom: So they’re not just guessing when to stop. The model’s own uncertainty is telling them when it’s had enough.

Meng: Exactly. And that saves a lot of compute. The paper shows they often stop at three thousand to four thousand samples instead of the full five thousand budget. That’s a direct reduction in training time and pseudo-query generation cost.

Lu: I also appreciate the resampling penalty. The naive approach would be to keep sampling from the biggest clusters, but they dynamically shift the budget toward clusters that haven’t been explored much. That prevents the model from overfitting to the dominant topics and ensures it gets exposure to the long tail of the domain.

Jane: And that’s a subtle but important improvement. It’s not just about picking the most uncertain documents. It’s about balancing that with diversity so you don’t end up with a model that’s great at one subtopic but useless for the rest.

Tom: The ablation study really shows this. Removing the epistemic sampling drops performance by over two points, and removing the aleatoric filter drops it even further. Each piece is doing real work.

Meng: And the gains scale with the model. With the four-billion parameter Qwen model, they see a three point four nine point average improvement over the baseline. That suggests the smarter sampling is even more valuable when you have a more capable model that can actually absorb the information.

Lu: It makes sense. A bigger model has a larger capacity to learn, so giving it the right data matters even more. This method is essentially a smarter data selection strategy that gets more valuable as your models get bigger.

Tom: So it’s not just a small tweak. It’s a fundamental improvement to the data pipeline that pays off more as the models get more powerful. That’s a great note to end on. Let’s wrap this up.

Conclusion: Tom: Alright, let’s bring it home. We’ve been talking about “UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval,” and it’s been a fascinating look at how to make search models smarter with less data.

Jane: To recap, the paper from Seoul National University tackles the problem of adapting a search engine to a new domain without any labeled queries. Instead of randomly picking documents to generate training data, they use a two-pronged uncertainty approach.

Tom: They filter out the noisy outliers, the high aleatoric uncertainty, and then they actively seek out the documents the model is most confused about, the high epistemic uncertainty. And they do this iteratively, so the model is always learning from its own blind spots.

Lu: And the results are convincing. They see consistent gains across multiple datasets and multiple model sizes, from a small DPR model all the way up to a four-billion parameter model. The efficiency gains are just as important, with early stopping saving a significant chunk of the compute budget.

Meng: From an engineering standpoint, this is a drop-in improvement to the data pipeline. You don’t need to change the retriever architecture or the training objective. You just change which documents you feed it. That makes it a very practical contribution.

Jane: The implications are pretty broad. Any system that relies on synthetic data for fine-tuning, not just search, could benefit from this uncertainty-aware selection. It’s a smarter way to spend your data budget.

Tom: So, a big thank you to the authors for this work. It’s a clear and well-executed idea that makes domain adaptation more effective and more efficient. We’ll be keeping an eye on where this line of research goes next.

Jane: And that’s a wrap on “UnIte.” Thanks for listening, everyone. We’ll see you on the next one.

More episodes

← Home