UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval

arXiv:2604.25142 · cs.IR, cs.AI · Submitted 2026-04-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval".

Jane: The paper was written by Jongyoon Kim, Minseong Hwang and Seung-won Hwang from Seoul National University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we are digging into a paper that has a pretty dense title, but a really practical problem at its core. It’s called “UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval.” Jane, I’m going to lean on you to unpack that title for our listeners.

Jane: Happy to, Tom. So imagine you’ve trained a search engine on one set of documents, like general web pages. Now you want it to work on a completely different set, say medical research papers. That’s domain adaptation. The trick is, you don’t have labeled queries for the new domain, so you have to generate fake ones. But you can’t generate them for all hundred thousand documents, so you have to pick which ones to train on.

Tom: And that picking part is where things get interesting. The old way was basically random, or just picking diverse documents. But this paper from Seoul National University says that’s not enough. They argue you need to think about uncertainty.

Lu: Exactly, Tom. And I love how they frame it. They split uncertainty into two types. Aleatoric uncertainty is the noise in the data itself, like a document that’s off-topic or just weird. Epistemic uncertainty is what the model doesn’t know yet. The authors argue you want to filter out the first kind and actively seek out the second.

Jane: Right, so if a document is just an outlier, it’s not going to teach the model anything useful. But if the model is unsure about a document that’s clearly on-topic, that’s a learning opportunity. That’s the sweet spot they’re aiming for.

Tom: And they’re not just doing this once. They’re doing it iteratively. The model learns, its uncertainty changes, and then they re-sample. It’s like the model is constantly checking its own blind spots.

Meng: I’m already thinking about the engineering side. The paper reports some pretty solid gains, like over three points of nDCG@ten on average with the larger model. That’s a real jump in retrieval quality, and they’re doing it with fewer training samples than the baseline.

Tom: So it’s not just smarter, it’s more efficient. That’s a win-win. We’ll get into the nitty-gritty of how they measure that uncertainty in the next segment, because that’s where the real cleverness is.

Summary: Tom: So we’ve established that “UnIte” is about picking the right documents to train a search model on a new domain. Jane, can you walk us through the core idea of how they actually decide what’s a good document?

Jane: Sure. They use two filters. First, they get rid of the noisy outliers, the high aleatoric uncertainty stuff. They do this with a simple lexical check, like seeing if a document shares enough words with its neighbors. If it’s too far away, it’s probably off-topic and gets tossed.

Lu: And that’s a smart move because it’s model-free. They’re not using the retriever’s own embeddings to judge the data. That would be circular. They’re using a pure lexical distance, so they’re not accidentally throwing away documents that the model just hasn’t learned to represent yet.

Jane: Exactly. Then, for the documents that survive that filter, they measure epistemic uncertainty. This is the clever part. They don’t just look at the model’s confidence. They compare the model’s understanding of a document against the actual statistics of the target domain.

Tom: So they’re not asking “is the model sure?” They’re asking “does the model know what’s important in this new field?”

Jane: Precisely. They look at how well the model predicts the high-IDF terms, the rare, domain-specific words. If the model can’t predict those, that document is a high-value training sample.

Meng: And the iterative part is crucial. The paper shows that if you just sample once, you waste your budget on documents that become trivial after the first round of training. Their loop re-evaluates the model’s knowledge every time, so it’s always chasing the actual knowledge gaps.

Lu: The results back that up. On TREC-COVID, they get a four-point boost over the diversity-based baseline with a small model. And on Robust04, it’s even more dramatic, a five-point jump. It’s a clear signal that uncertainty-aware selection is far more effective than just picking for variety.

Tom: So it’s not just about finding hard examples, it’s about finding the right hard examples that are actually relevant to the new domain. That’s a much more sophisticated approach. Let’s talk about how they keep this from being a computational nightmare in the next segment.

Improvements: Tom: We’ve talked about the core idea of “UnIte,” but I want to get into the practical improvements it brings. Meng, you were looking at the efficiency angle. What stands out?

Meng: The biggest thing is the early stopping criterion. They monitor the average epistemic uncertainty across the whole domain. It drops as the model learns, and then it starts to rise again when the model is just seeing redundant samples. That’s the plateau point, and they stop right there.

Tom: So they’re not just guessing when to stop. The model’s own uncertainty is telling them when it’s had enough.

Meng: Exactly. And that saves a lot of compute. The paper shows they often stop at three thousand to four thousand samples instead of the full five thousand budget. That’s a direct reduction in training time and pseudo-query generation cost.

Lu: I also appreciate the resampling penalty. The naive approach would be to keep sampling from the biggest clusters, but they dynamically shift the budget toward clusters that haven’t been explored much. That prevents the model from overfitting to the dominant topics and ensures it gets exposure to the long tail of the domain.

Jane: And that’s a subtle but important improvement. It’s not just about picking the most uncertain documents. It’s about balancing that with diversity so you don’t end up with a model that’s great at one subtopic but useless for the rest.

Tom: The ablation study really shows this. Removing the epistemic sampling drops performance by over two points, and removing the aleatoric filter drops it even further. Each piece is doing real work.

Meng: And the gains scale with the model. With the four-billion parameter Qwen model, they see a three point four nine point average improvement over the baseline. That suggests the smarter sampling is even more valuable when you have a more capable model that can actually absorb the information.

Lu: It makes sense. A bigger model has a larger capacity to learn, so giving it the right data matters even more. This method is essentially a smarter data selection strategy that gets more valuable as your models get bigger.

Tom: So it’s not just a small tweak. It’s a fundamental improvement to the data pipeline that pays off more as the models get more powerful. That’s a great note to end on. Let’s wrap this up.

Conclusion: Tom: Alright, let’s bring it home. We’ve been talking about “UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval,” and it’s been a fascinating look at how to make search models smarter with less data.

Jane: To recap, the paper from Seoul National University tackles the problem of adapting a search engine to a new domain without any labeled queries. Instead of randomly picking documents to generate training data, they use a two-pronged uncertainty approach.

Tom: They filter out the noisy outliers, the high aleatoric uncertainty, and then they actively seek out the documents the model is most confused about, the high epistemic uncertainty. And they do this iteratively, so the model is always learning from its own blind spots.

Lu: And the results are convincing. They see consistent gains across multiple datasets and multiple model sizes, from a small DPR model all the way up to a four-billion parameter model. The efficiency gains are just as important, with early stopping saving a significant chunk of the compute budget.

Meng: From an engineering standpoint, this is a drop-in improvement to the data pipeline. You don’t need to change the retriever architecture or the training objective. You just change which documents you feed it. That makes it a very practical contribution.

Jane: The implications are pretty broad. Any system that relies on synthetic data for fine-tuning, not just search, could benefit from this uncertainty-aware selection. It’s a smarter way to spend your data budget.

Tom: So, a big thank you to the authors for this work. It’s a clear and well-executed idea that makes domain adaptation more effective and more efficient. We’ll be keeping an eye on where this line of research goes next.

Jane: And that’s a wrap on “UnIte.” Thanks for listening, everyone. We’ll see you on the next one.

Jongyoon Kim, Minseong Hwang, Seung-won Hwang

Seoul National University

cs.IR, cs.AI

Submitted: 2026-04-28

Updated: 2026-08-18

Comments: ACL 2026 (Findings)

Journal ref: The 64th Annual Meeting of the Association for Computational Linguistics, 2026

Code: https://github.com/ldilab/UnIte

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: The paper addresses the challenge of Unsupervised Domain Adaptation (UDA) for neural retrievers.

Key concepts

Aleatoric uncertainty
This refers to noise within the data itself, such as a document that is off-topic or just unusual. The method filters these out because they do not teach the model anything useful.
Epistemic uncertainty
This represents what the model does not yet know. The authors use this type of uncertainty to actively seek out documents that are on-topic but confusing to the model, treating them as valuable learning opportunities.
Iterative Document Sampling
The process involves repeatedly training the model and re-sampling documents based on its changing uncertainty. This allows the model to constantly check its own knowledge gaps and learn more effectively over time.

Terminology

Summary

The paper addresses the challenge of Unsupervised Domain Adaptation (UDA) for neural retrievers. While neural retrievers pre-trained on large datasets like MS-MARCO perform well in the seen domain, they are limited to unseen domains. UDA via pseudo-query generation has emerged as a solution, where the retriever is fine-tuned on target-domain documents paired with generated queries. However, for corpora exceeding 100k documents, the number of generator calls scales with corpus size and is often infeasible under typical budgets, making document sampling the central bottleneck.

The authors identify critical flaws in existing sampling methods. Prior works use random sampling, while DUQGen employs a diversity-based approach using an external embedding model (Contriever). The authors observe that DUQGen's sampling tends to select documents from two problematic regions: (i) low-density regions containing atypical or outlier documents, and (ii) high-confidence regions where the model's predictive uncertainty is low (already-learned areas). Training exclusively on each region drops performance by 24.15 and 19.32 nDCG@10 respectively relative to DUQGen (62.56), demonstrating limited learning signal.

The authors propose Uncertainty-based Iterative Document Sampling (UnIte) which addresses these limitations through two complementary mechanisms:

AU reflects inherent ambiguity or noisy data. The authors use a density proxy based on lexical distance to the k-th nearest neighbor using BM25 scores. The distance is defined as:


Dk(d) = 1 / (ε + BM25(d, nk))

where ε = 1e-6 prevents zero division. Distances are normalized using modified z-scores, and documents with z(d) > z thr are filtered out as high-AU outliers. This is model-free, ensuring clean separation from epistemic uncertainty. The authors set k=3 and z thr=1.5, filtering approximately 5% of documents.

EU reflects the model's knowledge gaps. The authors measure EU by contrasting the model's representation with target domain statistics. Specifically:

  • Pre-compute token-level IDF for the target domain

  • Project document embeddings to vocabulary space via the model's MLM head to obtain token probabilities p(te d; θ t)

  • Compute EU score:


Uk(d; θ t) = Σ [log IDF(t) - p(te d; θ t)] for t ∈ Tk(d; θ t)

with k=1000 tokens. Documents with high EU are those where the model fails to predict high-IDF domain terms.

The method employs an iterative loop that:

  1. Recomputes EU each round as the model adapts

  2. Uses early stopping based on domain-averaged EU with Exponential Moving Average (EMA, α=0.4), stopping when uncertainty plateaus

  3. Uses Maximal Marginal Relevance (MMR) within clusters to balance uncertainty and diversity: score(d; θ t) = λ·Ûk(d; θ t) + (1-λ)·Ψ(d; θ t) with λ=0.5

  4. Introduces a resampling penalty that dynamically redistributes sampling weight toward underrepresented clusters: wi = Ci / (Pi + ε)

The authors evaluate on five large-scale BEIR datasets (>100k documents): TREC-COVID, Robust04, TREC-NEWS, Quora, and HotpotQA. They test four retrievers: DPR, coCondenser, COCO-DR, and Qwen3-Embedding-4B. Baselines include Random, GPL, Quality, and DUQGen. Pseudo-queries are generated using Llama3-8B-Instruct with temperature 0.8 and top-p 0.9.

UnIte consistently improves nDCG@10 across retrievers:

  • DPR: +2.45 average improvement over DUQGen (49.06 vs 46.61)

  • coCondenser: +0.75 improvement (55.69 vs 54.94)

  • COCO-DR: +0.26 improvement (62.27 vs 62.01)

  • Qwen3-Embedding-4B: +3.49 improvement (72.80 vs 69.31)

Statistically significant improvements (p<0.05) are marked with asterisks, with notable gains on TREC-COVID (+4.04 for DPR) and TREC-NEWS (+5.08 for DPR).

  • Removing EU sampling drops performance by approximately 4 nDCG@10 points on Robust04

  • Removing both AU and EU causes substantial gaps of 5 and 9 nDCG@10 points on TREC-COVID and Robust04 respectively

  • Removing resampling penalty degrades performance by 1.5 to 8.5 nDCG@10 across datasets

UnIte's domain-distribution-aware EU estimation outperforms MC-Dropout and Entropy by an average of 2.53 nDCG@10, demonstrating that effective document selection for domain adaptation requires incorporating the target domain distribution into EU estimation.

The authors show that the local minimum in average EU coincides with peak retrieval performance, validating EU as a reliable early stopping criterion. UnIte achieves 0.94 nDCG@10 improvement per 1k samples with DPR compared to DUQGen, typically converging at 3-5k samples versus DUQGen's fixed 5k.

  • Training time: 8 minutes with early stopping versus 10 minutes for baseline 5k

  • AU filtering: 120 seconds once

  • EU estimation: 150 seconds per iteration

  • Net time savings of 880 seconds with early stopping

Hyperparameters were tuned on FiQA (held-out development set) to avoid information leakage. Key findings:

  • Z thr = 1.5 achieves optimal performance (29.27 nDCG@10) while filtering 5% of documents

  • λ = 0.5 yields optimal balance between uncertainty and diversity

  • α = 0.4 balances noise reduction and trend preservation

The authors acknowledge: (1) preliminary treatment of multi-vector and re-ranking architectures (ColBERT, MonoT5) with approximations that depart from the model's native training objectives, (2) reliance on IDF statistics for domain distribution estimation, and (3) the resampling penalty does not explicitly account for rare minority topics that might be relevant to the target domain.

The paper demonstrates that uncertainty-aware sampling as a promising direction for budget-efficient domain adaptation, with UnIte yielding consistent gains across five BEIR corpora and four retrievers with fewer pseudo-queries.

Improvements for AI systems

Based on the scientific paper UnIte: Uncertainty-based Iterative Document Sampling for Domain Adaptation in Information Retrieval, here are the specific improvements I can make to an AI system and what the improved system can do:

  • Implementation: Add a dual-uncertainty estimation layer (Aleatoric + Epistemic) to the retrieval system's data pipeline.

  • Aleatoric Filter: Use BM25-based k-NN distance (k=3) to identify and filter out low-density/outlier documents (z-score > 1.5) before any training or query generation.

  • Epistemic Estimator: Compute per-document EU scores by comparing the model's token predictions (via MLM head) against domain IDF statistics, focusing on top-1000 tokens.

  • Dynamic Re-sampling: Replace one-shot sampling with a loop that:

  • Recomputes EU scores after each training iteration (500 docs/iteration)

  • Uses MMR (λ=0.5) to balance uncertainty and diversity within clusters

  • Applies a resampling penalty (inverse frequency weighting) to prevent over-sampling dominant clusters

  • Stops early when domain-averaged EU plateaus (EMA with α=0.4)

  • Targeted Pseudo-Query Creation: Generate queries only for documents that pass both uncertainty filters, using LLM prompting with domain-specific in-context examples (e.g., Llama-3-8B with temperature 0.8).

  1. Achieve higher retrieval accuracy with fewer training samples: +2.45 to +3.49 nDCG@10 improvement over diversity-only baselines (DUQGen) across BEIR benchmarks (TREC-COVID, Robust04, Quora, TREC-NEWS, HotpotQA).

  2. Reduce computational cost: Early stopping typically converges at 3-5k samples instead of 5k, saving 20% training time and 880 seconds of dataset construction time.

  3. Adapt to new domains more effectively: The system identifies and prioritizes documents that address the model's current knowledge gaps, rather than randomly sampling or relying on static diversity.

  4. Improve few-shot learning: The uncertainty-based sampling can be applied to any supervised learning task (e.g., classification, NER) to select the most informative training examples, improving accuracy with limited labels.

  5. Enhance active learning: The iterative EU estimation and resampling penalty can be generalized to any active learning scenario, ensuring the model continuously learns from diverse, high-uncertainty regions without redundant sampling.

  6. Better domain adaptation for NLP models: The AU filtering (removing noisy/outlier data) combined with EU prioritization (focusing on knowledge gaps) can be applied to fine-tune any pretrained language model for a new domain, improving downstream task performance.

  • For a question-answering system: It can quickly adapt to a new corpus (e.g., biomedical literature) by selecting the most informative documents to generate training queries, improving answer retrieval accuracy by 5-8% nDCG@10 within 4k samples.

  • For a document classification system: It can filter out irrelevant/outlier documents and prioritize those that the model is most uncertain about, leading to faster convergence and higher F1 scores with 50% fewer labeled examples.

  • For a recommendation system: It can identify which user-item interactions are most informative for adapting to a new user base, improving recommendation quality while minimizing data collection costs.

Abstract

Unsupervised domain adaptation generalizes neural retrievers to an unseen domain by generating pseudo queries on target domain documents. The quality and efficiency of this adaptation critically depend on which documents are selected for pseudo query generation. The existing document sampling method focuses on diversity but fails to capture model uncertainty. In contrast, we propose **Un**certainty-based **Ite**rative Document Sampling (UnIte) addressing these limitations by (1) filtering documents with high aleatoric uncertainty and (2) prioritizing those with high epistemic uncertainty, maximizing the learning utility of the current model. We conducted extensive experiments on a large corpus of BEIR with small and large models, showing significant gains of +2.45 and +3.49 nDCG@10 with a smaller training sample size, 4k on average.

Sources

Related papers