OnceSelect: Reusable Data Selection for Efficient Multimodal Instruction Tuning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "OnceSelect: Reusable Data Selection for Efficient Multimodal Instruction Tuning".
Tom: Multimodal instruction tuning requires massive datasets, making data selection critical for training efficiency and cost.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show! We're talking about a paper that’s really making waves in how we train multimodal instruction tuning models. Today we’re looking at "OnceSelect: Reusable Data Selection for Efficient Multimodal Instruction Tuning." Jane, you ready to break down what this work is all about?
Jane: Absolutely, Tom. This paper tackles the big question of how to make those massive datasets actually useful without needing everything. The main idea behind "OnceSelect" is creating a single mechanism that we can train just once and then use it for any dataset or model we throw at it later, which really saves a ton of time and compute resources.
Lu: I find the concept fascinating because it moves the focus from just brute-forcing data quantity to understanding the underlying semantic structure of the instructions themselves; that's a very creative approach to data curation.
Meng: From an engineering standpoint, that reusability is huge because we don't have to rebuild a complex scoring network every single time we switch VLM architectures or datasets. But I’m curious how robust this selection process really is when the underlying data distribution shifts significantly between training and deployment environments.
Lalam: If we think about our culture, this suggests that instead of constantly feeding models everything, we can distill the essence of what truly matters in instruction tuning, leading to a more focused and efficient learning process for all.
Tom: That’s exactly right; it shifts the focus from volume to signal quality. So, "OnceSelect" claims it solves the problem where huge instruction datasets are often redundant or even incorrectly annotated, suggesting that diversity is what really drives performance five <ref:2605.26761#pg0>. Jane, can you elaborate on what this framework actually proposes?
Jane: Well, "OnceSelect" proposes a two-phase system. First, they encode the image-instruction pairs using a frozen CLIP model to get a joint multimodal feature vector and then partition those features into clusters using K-Means clustering eleven <ref:2605.26761#pg1>. This step helps expose the semantic structure of the data.
Lu: The clustering part is key because it lets us select samples based on their position within that semantic space rather than relying solely on how well a specific model predicts an answer. It's about grouping things that mean something similar together, which is very insightful <ref:2605.26761#pg0>.
Meng: So, we’re essentially using clustering to create pseudo-labels for supervision without needing a human annotator upfront? That sounds like a clever shortcut for getting initial training signals.
Lalam: It means the data itself starts telling us what the groups are, which is much more natural than trying to label everything manually before training begins.
Tom: Right, and then they take that core set of samples—the ones closest to the cluster centers—and train a lightweight selector on them for just a few epochs <ref:2605.26761#pg2>. This selector is then used to score every new candidate by its maximum class probability <ref:2605.26761#pg2>.
Paper summary: Jane: And here’s where the clever part comes in: they intentionally don't train this selector all the way to convergence, only for a few epochs, which they suggest keeps it uncertain about trickier samples. They then select samples with low confidence scores—the ones that "resist clean cluster assignment" <ref:2605.26761#pg2>.
Lu: That exploitation of residual uncertainty is a very strong signal for informativeness; it means we're keeping the edge cases that are hard to classify, which is where the real learning happens <ref:2605.26761#pg0>.
Meng: If we only train it for a few epochs, how do they ensure that the selector doesn't just pick random outliers instead of genuinely useful data points? That sounds like a delicate balancing act.
Lalam: The method relies on the initial cluster structure being good enough to guide the selection, and that low-confidence metric is designed specifically to filter out the common, well-behaved samples that would otherwise dominate training.
Tom: It seems like they’ve managed to build a system that filters out typical data while keeping the complex, informative tail of any distribution <ref:2605.26761#pg0>. So, we're filtering for the unusual ones? Jane, what are the actual implications of this single reusable selector?
Jane: The primary advantage is its transferability; because it only depends on a frozen joint multimodal space, that one selector can be applied to entirely new datasets or even target models with different architectures without needing any recomputation <ref:2605.26761#pg0>.
Lu: That ability to apply a single learned mechanism across different domains and model sizes is what really opens up possibilities for widespread efficient instruction tuning across the entire AI ecosystem <ref:2605.26761#pg1>.
Meng: I wonder about the practical deployment. If we have this selector, does it mean we can dramatically reduce the time needed to prepare training data for a new VLM? Or is there still too much overhead in the initial encoding step?
Tom: The cost analysis suggests it’s quite efficient; they report only twenty point zero GPU-hours on LLaVA-665K for fine-tuning, which is significantly faster than some other selection baselines we looked at <ref:2605.26761#pg0>.
Jane: So, to summarize the core idea of "OnceSelect," it’s a framework that trains a lightweight selector once on frozen CLIP features to identify informative samples based on low confidence scores, allowing for efficient training across diverse datasets and model architectures.
Lu: It really shifts the entire paradigm toward data efficiency by focusing selection on semantic novelty rather than just sheer volume <ref:2605.26761#pg1>.
Meng: I think the impact will be seen in how quickly we can iterate on instruction tuning without drowning in massive, redundant datasets.
Lalam: For us, it means we can focus our resources on training models that truly capture nuanced understanding rather than just memorizing common patterns.
Tom: It's a very elegant solution for handling the data bottleneck, and I think the future of efficient multimodal instruction tuning is definitely headed in this direction. We’ve seen how "OnceSelect" handles diverse datasets; next up, we look at what this means for practical application and the long-term trajectory of instruction tuning research.
Conclusion: Tom: So, we've been looking at how this framework works in detail, and now we need to wrap up our discussion on "OnceSelect: Reusable Data Selection for Efficient Multimodal Instruction Tuning." Jane, can you give us a quick recap of what this paper actually proposes?
Jane: Absolutely, Tom. Essentially, the core idea is a method that trains one selector once on frozen CLIP features to filter out redundant data from any new dataset without retraining anything else one. It uses clustering and then scores samples based on how confident the selector is in assigning them to a specific cluster.
Lu: I think the clever part lies in using that low-confidence signal, which helps identify those truly interesting or atypical samples that standard methods might ignore <ref:2605.26761#pg0>. It’s about finding the edges of the data distribution rather than just taking the bulk of it.
Meng: From a practical standpoint, this means we can drastically cut down on the time spent curating massive instruction datasets before we even start training our models <ref:2605.26761#pg0>. It’s about making the data preparation phase much leaner for deployment.
Lalam: For me, it's exciting because this approach suggests that we can distill the essential learning signal from a huge pile of text and images, which means our resulting models might become much more focused on high-level reasoning rather than just pattern matching <ref:2605.26761#pg1>.
Tom: That focus on the essential signal is what really gets my attention. Jane, when we look at the title and authors of "OnceSelect: Reusable Data Selection for Efficient Multimodal Instruction Tuning," what does that tell us about the scope of this work?
Jane: The title really highlights two big things: reusability and efficiency. It shows they didn't just create a tool for one specific model or dataset; they built something that can be applied across different models and any dataset you throw at it, which is a major practical win <ref:2605.26761#pg0>.
Lu: And the authors clearly focused on creating a mechanism that works well within the frozen multimodal space of CLIP, which is smart because it keeps the dependency on specific vision-language models flexible <ref:2605.26761#pg0>.
Meng: I see the implication for our teams being that we can standardize our data intake process. If this method works as well across different VLM architectures, then we don't need separate pipelines for every new model we want to test <ref:2605.26761#pg0>.
Lalam: This reusability is huge because it means the culture shifts toward building modular tools instead of monolithic solutions. We can create a single data pipeline component that serves many different AI applications, which seems like a very powerful way to scale our impact <ref:2605.26761#pg1>.
Tom: So, we're looking at a single selector that handles the whole data selection job for us across different environments. This really suggests that the bottleneck in instruction tuning might be less about having more data and more about how intelligently we select what we use <ref:2605.26761#pg0>.
Jane: Exactly, Tom. The authors show that this method can achieve strong performance even on datasets they haven't seen during the selection training phase, which speaks to its generalizability <ref:2605.26761#pg0>.
Lu: I think the future work they hint at involves making those hyperparameters—like how many clusters we use or how much we train the selector—more adaptive so it can handle even more unpredictable distribution shifts <ref:2605.26761#pg3>.
Meng: That's interesting because if the system could adapt its selection criteria on the fly, it would make deployment much smoother in real-world scenarios where data quality might vary unexpectedly <ref:2605.26761#pg3>.
Lalam: I think this capability means we can move toward a culture of truly adaptive AI systems, where the data pipeline learns to filter itself based on the specific needs of the target model, which is a big step forward for building robust applications <ref:2605.26761#pg1>.
Tom: It sounds like "OnceSelect" is positioning us to move past the era where we just dump massive amounts of data into instruction tuning and start focusing on quality filtering instead <ref:2605.26761#pg0>. Next up, we have to look at how this efficiency translates into real-world model performance gains.
Mingkang Dong, Hongyi Cai, Xiwen Lei, Jie Li, Tao Zhang, Muxin Pu
Faculty of Computer Science, Universiti Malaya · School of Mathematics and Computer Sciences, Nanchang University · Faculty of Computer Science, University of Science and Technology Beijing · School of Information Technology, Monash University Malaysia
cs.CV
Submitted: 2026-05-26
Updated: 2026-10-04
Comments: Mingkang Dong and Muxin Pu contributed equally to this work. Yuqian Fu is the corresponding author
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Multimodal instruction tuning requires massive datasets, making data selection critical for training efficiency and cost.
Key concepts
- Once-for-All (OFA) Data Filtering Framework
- This is the core process that trains a single selector once. It involves encoding image-instruction pairs with a frozen CLIP model to get features, clustering them, and then training an MLP selector on pseudo-labels derived from these clusters. This trained selector can then be applied to any new data without retraining.
- Cluster-centric Pseudo-labeling
- Instead of manual labeling, the framework uses clustering to create supervision. Samples close to each cluster centroid are grouped into a core set, and they are assigned a pseudo-label equal to their cluster index. This establishes initial training signals based on the inherent semantic structure of the data distribution.
- Low Confidence Threshold Selection
- After training, the selector scores every sample. The framework selects samples within each cluster that have the lowest confidence (lowest maximum class probability). This retains 'atypical' or hard-to-classify samples, preserving informative data while discarding typical ones.
- Model-Agnostic Selector
- The trained selector relies only on frozen CLIP features, not on any specific target Vision-Language Model (VLM). This means a single selector can be used to select data for various different VLMs across different architectures, ensuring high transferability.
Terminology
Summary
Multimodal instruction tuning requires massive datasets, making data selection critical for training efficiency and cost. The proposed framework addresses this by developing a reusable selector that trains once and applies to any dataset or model without recomputation.
The gist: A single, transferable selector provides an effective and reusable solution for efficient multimodal instruction tuning.
How it works
The Once-for-All (OFA) Data Filtering Framework operates in two main phases: training the selector once, and applying it anytime to new data. The process begins by encoding image-instruction pairs using a frozen CLIP model to create a joint multimodal feature vector, which is then normalized. These features are then partitioned into clusters using K-Means clustering on the normalized embeddings.
Clustering and Pseudo-labeling
The clustering step exposes the semantic structure of the data, allowing for selection based on distribution rather than specific model performance. To supervise the selector without manual annotation, a cluster-centric pseudo-labeling scheme is employed. Samples lying closest to each centroid are collected into a core set, and each sample in this core set is given a pseudo-label equal to its cluster index. This process establishes the supervision for training the selector by reflecting the overall diversity of the data distribution.
Selector Training
A lightweight selector, denoted as an MLP built on top of frozen CLIP features, is trained on this core set of pseudo-labeled samples. The loss function used is cross-entropy loss to predict cluster-derived pseudo-labels. Crucially, the selector is deliberately not trained to convergence,
being trained for only a few epochs (e.g., three). This under-training ensures that the selector becomes confident on typical, easily separable samples while remaining uncertain on atypical, non-trivial ones—this residual uncertainty is exploited as the signal for informativeness.
Data Selection and Transferability
Once trained, the frozen selector is used to score every candidate sample in any new dataset. The confidence of a sample is defined as its maximum class probability,
which measures how confidently the selector assigns it to a single cluster. Samples are then retained based on a per-cluster threshold, specifically keeping those with lowest confidence
(low F(xi)) within each cluster. This process retains samples that resist clean cluster assignment,
thereby preserving the informative tail of any dataset's distribution while discarding typical and redundant samples.
Transferability and Performance
The primary advantage of OFA is its reusability. Because the selector depends only on the frozen joint multimodal space and not on any target VLM, a single trained selector can be applied to (i) datasets unseen during training, by encoding and scoring their samples directly, and (ii) target models of different scales and architectures. Experiments demonstrate that OFA achieves strong performance across unseen datasets and model architectures while using only a fraction of the data. For instance, when the selector trained on LLaVA-665K is applied to Vision-Flan-186K, OFA reaches a relative performance of 110.6%, surpassing full-data training. Furthermore, the method is model-agnostic,
as a subset selected once can benefit VLMs across architectures without any reselection.
Hyperparameter Sensitivity
The framework's behavior is governed by several hyperparameters: the number of clusters K, the core radius γ, the selector training budget E, and the confidence threshold τ. The analysis shows that performance is stable across reasonable ranges of K and selection ratios (e.g., 15% selection ratio yields peak performance). However, the training budget E must remain small; increasing it beyond three epochs erodes the low-confidence signal that distinguishes informative samples from redundant ones. This sensitivity analysis confirms that while OFA is robust, an automatic or cluster-adaptive scheme for selecting these hyperparameters would enhance its robustness to distribution shift.
Cost Analysis
OFA demonstrates superior computational efficiency compared to existing selection baselines. The total cost of the OFA process, including encoding and scoring, is relatively low—only 20.0 GPU-hours on LLaVA-665K for fine-tuning. This is significantly faster than methods like Self-Filter (73.5 GPU-hr) or COINCIDE (55.5 GPU-hr). Unlike methods relying on external LLM APIs, OFA does not incur substantial monetary and time overheads, making it a cost-effective solution for efficient multimodal instruction tuning.
Conclusion
OFA successfully addresses the bottleneck in data selection by proposing a framework that trains a lightweight selector once on frozen CLIP representations to identify informative samples based on low confidence. This single, reusable mechanism allows for efficient training across diverse datasets and model architectures, confirming its efficacy as a once-for-all
solution.
References
[1] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023.[Online]. Available:
Improvements for AI systems
Here are specific improvements to AI systems based on the Once-for-All (OFA) Data Filtering Framework, along with what these improved systems can achieve:
The core improvement lies in creating a highly efficient, reusable, and data-agnostic instruction tuning pipeline that maximizes learning from diverse data distributions while minimizing computational cost.
-
Improve the efficiency of Multimodal Instruction Tuning by achieving near full-data performance using only a small fraction of the original training data (e.g., 15% selection ratio).
-
Enable cross-dataset and cross-architecture fine-tuning without requiring any retraining or recomputation of the selection mechanism when switching between new instruction tuning corpora or target VLM architectures.
-
Mitigate model overfitting and hallucination in downstream tasks by specifically selecting
low-confidence
samples, which are identified as atypical, boundary cases, or ambiguous inputs that resist simple classification.
Specific capabilities of the improved AI system:
-
A VLM fine-tuned via OFA can achieve state-of-the-art performance on a wide range of benchmarks (VQAv2, GQA, MMBench) by selectively retaining samples that challenge its current knowledge base rather than simply memorizing common patterns.
-
The system can be deployed across different instruction tuning datasets (e.g., moving from LLaVA-665K to Vision-Flan-186K) with a single, pre-trained selector, instantly adapting to the new data distribution without needing to recompute complex selection metrics or retrain the selection model.
-
The resulting models will exhibit superior performance on reasoning tasks (e.g., SQA-I) and anomaly detection tasks because they are explicitly trained on the
informative tail
of the data distribution, which includes cluttered scenes, multi-turn dialogues requiring inference beyond direct perception, and atypical scene compositions that simpler models often fail to categorize correctly. -
The system will be significantly more cost-effective in terms of training time and computational resources compared to existing gradient-based or scoring-based selection methods, as the selector training is performed only once on a frozen representation space, making subsequent data filtering extremely fast (linear in dataset size).
-
The VLM will be robust against
data shift
and distribution changes because the selection signal is derived from the intrinsic semantic structure of the data (via frozen CLIP embeddings and K-means clustering) rather than being tied to a specific model's internal gradients or loss function, ensuring transferability across different visual domains and text modalities.
Abstract
Multimodal instruction tuning is widely used to adapt multimodal large language models (MLLMs), yet the large-scale image-text datasets it relies on are often highly redundant. Existing data selection methods are commonly tied to specific datasets, target models, or training states, requiring retraining or recomputation when transferred to new settings. We ask whether a selection signal can instead be learned once and reused across datasets and models. We propose OnceSelect, a reusable data selector trained once and directly applied to unseen datasets. OnceSelect encodes image-instruction pairs in a frozen joint multimodal space, clusters them into pseudo-labels capturing coarse semantic structure, and trains a lightweight selector using a fixed validation macro-accuracy criterion. Low-confidence samples under the learned partition are treated as informative candidates. As the selector is independent of the downstream MLLM, it can be transferred across datasets without retraining or model-dependent scoring, while selected subsets can be reused across model architectures. Using only 15% of LLaVA-625K, OnceSelect retains 99.2% of full-data aggregate performance across nine evaluation metrics. Without retraining, it achieves 103.4% and 102.5% relative performance on unseen Vision-Flan-186K and LRV-Sub-180K. The same 93.7K LLaVA subset further achieves 99.9-106.7% relative performance across Qwen3-VL and InternVL3 models without model-specific scoring or reselection. These results demonstrate efficient and reusable multimodal data selection across datasets and model architectures. Code is available at https://github.com/DMK041218/OnceSelect
Sources
- Visual Instruction Tuning
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
- ICONS: Influence Consensus for Vision-Language Data Selection
- From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization
- Deep Learning on a Data Diet: Finding Important Examples Early in Training
- Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets
- Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
- Concept-skill Transferability-based Data Selection for Large Vision-Language Models
- Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning
- MergeIT: From Selection to Merging for Efficient Instruction Tuning
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Evaluating Object Hallucination in Large Vision-Language Models
- Learning Transferable Visual Models From Natural Language Supervision
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models