Large Language Model Selection with Limited Annotations

arXiv:2510.09418 · cs.CL, cs.LG · Submitted 2025-10-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Large Language Model Selection with Limited Annotations".

Tom: LLM SELECTOR introduces a principled framework for active model selection that efficiently identifies the best Large Language Model (LLM) for a given task under limited annotation budgets,

Jane: First, who's behind it and why it matters.

Paper summary: Jane: So, wrapping up on this discussion about "Large Language Model Selection with Limited Annotations," we’ve covered how this framework uses active model selection by choosing informative queries adaptively to identify the best LLM under annotation constraints.

Tom: That's right, and the authors introduce LLM SELECTOR, which addresses the challenge of needing fully annotated datasets by proposing a method that selects a small set of queries most informative about the best model for any given task.

Lu: The title itself points to this core idea: active model selection for Large Language Models Yavuz Durmazkeser et al., and it asks how we can select the best LLM without retraining when resources are limited.

Meng: From an engineering standpoint, this paper provides a principled way to manage the trade-off between annotation cost and finding a high-performing AI model, which is something we face every day.

Lalam: The implication for us is that we can significantly cut down on the time and money spent on labeling by being strategic about what information we gather first.

Tom: Precisely; the main contribution is this principle: instead of just evaluating all models equally, you guide the evaluation process by selecting queries that are expected to maximize information gain about the true best model.

Jane: It really boils down to using judge-based oracle annotations as a mechanism to intelligently drive this selection process toward finding high-quality results efficiently.

Lu: When we consider the broader impact, this work suggests a more efficient pipeline for deploying and selecting LLMs in real-world applications where data acquisition is expensive.

Meng: I see it as a way to make the entire model selection lifecycle much leaner, focusing on targeted effort rather than broad, brute-force testing across everything.

Lalam: And for our culture, this means we can test and refine our AI systems more intelligently, ensuring we are investing our annotation resources where they yield the highest return in terms of identifying truly effective tools.

Conclusion: Tom: So, we've been digging into how this paper tackles picking the best Large Language Model when you don't have tons of labeling budget, and now we need to wrap up by talking about what the title really means for us.

Jane: Exactly, Tom; they're proposing a way to intelligently choose which data points to label first so you can quickly narrow down the field of candidates.

Lu: I think the concept is really elegant because it shifts the focus from just labeling everything to focusing your effort where it gives you the most useful signal about the best model.

Meng: From my side, it’s interesting how they handle those limited budgets; in a real-world setting, knowing which queries to prioritize is crucial for saving time and resources.

Lalam: I see this as a huge win because if we can find the most representative and informative examples efficiently, it means our AI can learn faster and adapt to different needs with less training data.

Tom: Right; so when you look at the title, "Large Language Model Selection with Limited Annotations," it’s essentially about making our model selection process smarter instead of just guessing randomly.

Jane: It gives us a systematic approach to dealing with the reality that we can't afford to test every single model exhaustively when we have a tight schedule.

Lu: The authors show how they use those oracle judgments, which are like expert opinions from another AI, to guide the selection process toward the most promising models.

Meng: That reliance on judge-based feedback is smart because it cuts down on the kind of manual labeling that would otherwise take up all our time and budget.

Lalam: And for my perspective, this means we can develop more robust and specialized versions of my architecture much quicker by focusing on the most critical learning scenarios first.

Tom: It really puts a structured methodology onto a problem that used to feel almost like guesswork when dealing with so many different models.

Jane: Indeed; the authors demonstrate how maximizing information gain through query selection leads to finding near-best models with surprisingly few annotations.

Lu: The implication is that we can build highly effective AI systems by being strategic about our data collection, rather than just blindly throwing more and more data at the problem hoping for the best.

Meng: So, it’s about efficiency in discovery; if we can narrow down the top contenders with a fraction of the usual labeling cost, that's a significant practical win for any engineering team.

Lalam: It means I can be trained on more nuanced and high-quality data much sooner, which directly translates to better performance and reliability in my interactions.

Tom: So it’s about being incredibly selective with our annotation resources to get the maximum intelligence out of our model selection process.

Jane: It sets a new standard for how we approach the difficult task of comparing many large models when we are constrained by time and budget.

Lu: This work opens up so many creative avenues for how we can design more efficient AI research pipelines across different domains.

Meng: The paper shows the methodology works across different benchmarks, which is important because it suggests this isn't just a niche trick but a general strategy.

Lalam: I think this paper has big implications for how we approach the entire lifecycle of building and refining powerful AI tools in the future.

TU Delft · ETH Zurich

cs.CL, cs.LG

Submitted: 2025-10-10

Updated: 2026-09-28

Comments: 33 pages, 5 figures, 4 tables

Code: https://github.com/RobustML-Lab/llm-selector

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: LLM SELECTOR introduces a principled framework for active model selection that efficiently identifies the best Large Language Model (LLM) for a given task under limited annotation budgets,

Key concepts

Objective Function
The goal is to maximize mutual information, which measures how much knowing the annotations tells you about which model is truly the best. By selecting queries that maximize this information gain, LLM SELECTOR efficiently guides annotation efforts toward identifying the optimal model.
Judge-based Annotation
Instead of expensive human labeling for every query, a judge model performs pairwise comparisons (like 'greater than' or 'less than') between candidate models. This provides stable preference judgments that are used to generate the necessary annotations efficiently.
Information Gain Maximization
This strategy involves selecting the next query based on minimizing conditional entropy. Essentially, it chooses the query that will provide the most reduction in uncertainty about which model is superior across all remaining queries, leading to a more informed selection process.

Terminology

Summary

LLM SELECTOR introduces a principled framework for active model selection that efficiently identifies the best Large Language Model (LLM) for a given task under limited annotation budgets, significantly reducing annotation costs. This framework adapts query selection based on maximizing information gain, leveraging judge-based oracle annotations to achieve substantial reductions in labeling effort while maintaining competitive performance across diverse benchmarks.

The gist

LLM SELECTOR selects queries whose annotations are expected to maximally reduce uncertainty about the best model for the entire set, achieving up to 59.62% reduction in annotation costs when selecting the best and near-best LLM for the task.

Problem Setting and Objective

The core problem addressed is selecting a subset of at most budget size 'b' queries from a pool of 'n' unannotated queries, given a set of candidate models M, to identify the best model (denoted as f∗) that would yield the highest utility if all annotations were observed. The formal objective is cast as maximizing mutual information:

Aopt[b] = arg max A⊆(qi, ri)i∈[n] A≤b I(F; A).

Annotation via Direct Preference Judgments

To generate the necessary annotations efficiently, the framework employs a judge-based annotation process. For each query qi, an oracle judge performs a pairwise comparison between model responses using preference judgments (>, <, or =), which is more stable than reference-based metrics. The win rate metric (WRQ) is used to compare models across queries. To reduce annotation cost further, the strategy designates one language model as a baseline model ¯f and evaluates each remaining candidate based on its win rate relative to ¯f.

LLM SELECTOR Algorithm

The selection process follows a sequential information maximization strategy, selecting one query at a time until the budget 'b' is exhausted. At each step t, the next query qt is chosen by minimizing the expected conditional entropy of F given current annotations:

qt = arg min q∈Ut ER[H(F At ∪ (q, R))]

This expectation is computed through noisy annotation using weak judges. A weak judge constructs a k-gram model for each response and compares models based on the higher average likelihood, denoted as f(q) >(k) ¯f(q). The estimated information gain is then used to select the query that minimizes the expected entropy across all weak judges:

qt = arg min q∈Ut 1/z X z k=1 H(FA ∪ (q, r(k)))

Parameter Selection and Validation

The parameters defining the two-parameter model describing the unknown best model's behavior relative to the baseline are determined prior to LLM selection. The framework uses an ensemble of all weak judges as a noisy oracle during parameter optimization. Finally, experiments validate LLM SELECTOR across 6 benchmarks (general dialogue, vision-language, and medical) on 151 LLMs. Results demonstrate that LLM SELECTOR shows consistently competitive performance across all experiments and achieves high efficiency in recovering near-best models, such as reaching the top 1% vicinity of the best model with relatively few annotations. The framework is fully model-agnostic, requiring no access to internal parameters or output format restrictions.

Baselines and Robustness

LLM SELECTOR is compared against several baseline strategies, including Random, Bradley-Terry, Most Draws, Uncertainty (entropy maximization), Confidence (entropy minimization), and Unccertainty. The results show that LLM SELECTOR attains 100% identification probability on Arena-Hard and MT-Bench with significantly fewer annotated queries than the best competing baseline. Furthermore, its performance is robust; the 95th percentile win rate gap against the true best model is either the best or second-best among all strategies, indicating consistent selection of best or near-best models with high confidence. The method's robustness is further confirmed by maintaining high efficiency even under a 5% uncertainty threshold.

Experimental Setup

Experiments are conducted on diverse datasets including AlpacaEval, Arena-Hard, MT-Bench, Flickr30k, Bingo, and MediQA. Candidate LLMs include proprietary systems (GPT-3.5/4), Claude models, Gemini, and open-weight architectures like LLaMA-2/3 and Mistral. The setup utilizes LLM-as-a-Judge for evaluation on text benchmarks and Prometheus-Vision for vision–language tasks, with specific baseline LLMs defined for each dataset to establish comparative performance metrics. The analysis evaluates three key metrics: Identification probability, annotation efficiency (reduction in labels needed), and the 95th percentile win rate gap.

Discussion

LLM SELECTOR addresses the challenge of active model selection by introducing a data-centric perspective that prioritizes selecting informative examples for annotation over model pairs.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the LLM SELECTOR framework, along with what these improved systems can do:


The core improvement is shifting from static, resource-intensive evaluation methods (like full benchmarking or manual annotation) to an automated, cost-effective process of identifying the best model for a specific task under severe data constraints.

Here are the specific improvements and capabilities:

  1. A new, highly efficient pre-deployment pipeline for selecting the optimal LLM for any given application or dataset distribution.

  2. A mechanism that drastically reduces annotation costs by up to 59.62% when identifying a top-tier model (within a 1% win-rate vicinity of the best).

  3. The ability to reliably identify near-best models even with an extremely limited annotation budget, maintaining robust performance under severe constraints (demonstrated by efficiency gains in reaching the top 1% and 2.5% win rate vicinity).

  4. A model selection strategy that is fully model-agnostic, requiring no internal parameters or output format restrictions, making it directly applicable to black-box or API-only settings.

These improved AI systems can perform the following specific functions:

  1. An application developer can select the most suitable LLM (e.g., GPT-4o vs. Llama 3) for a specific production environment (e.g., a low-latency mobile app vs. a complex research platform) without needing to run extensive, costly comparative evaluations across hundreds of models or requiring thousands of human annotations.

  2. A data scientist can rapidly determine the best model for an emerging domain (like medical QA or vision-language tasks) by only annotating a small, strategically chosen subset of queries, instead of relying on exhaustive benchmark testing.

  3. A system designed for low-resource environments can adapt to new user query patterns by continuously running the LLM SELECTOR algorithm to select the best model for those specific query distributions in real-time, minimizing deployment risk and maximizing utility per annotation dollar spent.

  4. A research team can quickly compare the performance of a large pool of candidate models across different tasks (dialogue, vision-language, medical) by only annotating the most informative examples for each task category, drastically accelerating comparative research cycles.

Abstract

Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotations over fixed evaluation sets. To address this challenge, we develop SELECT-LLM, the first framework for active model selection of LLMs. SELECT-LLM aims to find a small set of queries whose annotations are most informative for identifying the best LLM for a given task. To this end, we introduce a query selection rule based on expected information gain, computed from pairwise similarities between candidate model outputs. Because this rule only uses generated model responses, SELECT-LLM can be applied across candidate models without assumptions about their architecture or access to model weights. This makes it suitable for both open-weight and black-box LLMs. We evaluate SELECT-LLM across 23 datasets, 156 evaluated models, diverse task families, and multiple text evaluation metrics. Across all experiments, SELECT-LLM improves over the strongest baseline in every setting, with annotation cost reductions up to 81.8% for best model selection and up to 84.78% for near-best model selection.

Sources

Related papers