Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

arXiv:2603.12478 · cs.CV, cs.LG · Submitted 2026-03-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning".

Jane: The paper was written by Rujie Wu, Haozhe Zhao, Hai Ci and Yizhou Wang from Peking University and University of Illinois Urbana-Champaign and National University of Singapore.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone! We've got a paper that's been making the rounds, and the title alone is a promise: "Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning." Jane, my first reaction is, is that even legal in the AI world?

Jane: Ha! It feels like it shouldn't be, right? We're so used to the mantra that more data is always the answer. But this paper is basically saying, "Hold on, what if we just pick the *right* data instead of all the data?" And they've got the results to back it up.

Tom: And I love that it's coming from a team at Peking University, with Rujie Wu as the corresponding author. They're not just theorizing; they ran this on real hardware with a real model.

Jane: Exactly. They used Qwen3-VL-8B-Instruct, which is a serious multimodal model, and they trained it on a huge pool of mixed image and video data. The pool is massive, like five hundred twelve thousand samples in their baseline.

Tom: But here's the kicker, Jane. They didn't use all of it. Their whole framework, GDO, builds these tiny, optimized subsets. We're talking about twelve thousand to fifty-three thousand samples. That's like a tenth of the baseline.

Jane: And they're not just getting away with it; they're *beating* the baseline. On MVBench, they went from sixty-two point two seven percent accuracy to sixty-three point six five percent. On MLVU, it's even more dramatic, jumping from forty-three point eight one percent to forty-six point eight nine percent.

Tom: That's a huge jump for using a fraction of the data. It's like cleaning your room and suddenly finding you can walk faster because there's nothing in the way.

Jane: That's a good way to put it. The paper's core argument is that most of that 512k dataset is redundant or just not that useful for the specific skills you want to teach. So they built a system to figure out which samples are actually valuable.

Tom: And that's what we're going to dig into today. How do they decide what's "valuable"? What does "goal-driven" even mean here? We've got the whole crew here to break it down.

Jane: We do! And I think the most exciting part is that this isn't just about saving compute. It's about getting *better* results by being smarter about what you feed the model. It's a paradigm shift.

Tom: A paradigm shift with numbers attached. I'm sold already. Let's get into the nitty-gritty of how they actually pull this off.

Summary: Tom: So, Jane, we've established that "Less Data" is possible. But the paper's real meat is in the *how*. They call it Goal-Driven Data Optimization, or GDO. It's not just one magic filter.

Jane: Right, it's a framework. And the clever part is that they separate the "scoring" from the "goal." Think of it like this: they have a universal quality score for every sample in the pool, but then they have different "presets" that decide how to use that score based on what you want the model to be good at.

Tom: Okay, so it's like a chef rating ingredients on freshness, but then deciding whether to make a salad or a stew based on the customer's request.

Jane: Exactly! And they have four different "recipes" in the paper. There's "MinLoss," which just wants the easiest samples to train on fast. Then there's "Diverse," which tries to cover as many different topics as possible.

Tom: And then we get to the interesting ones for video: "Temp" and "Temp+." These are the ones that really push for temporal understanding, meaning the model needs to understand change over time, ordering, and motion.

Jane: And the results show that the goal matters a lot. The "Temp+" profile, which has the strongest temporal pressure, gives the best overall results. It's not just about picking high-quality data; it's about picking the *right kind* of high-quality data for the task.

Tom: So, how do they actually score these samples? They have six descriptors, right? It's not just one number.

Jane: Right, it's a six-dimensional vector. They look at things like optical flow to measure motion, a "video-dependence score" to see if the answer actually requires watching the video, and something they call "self-consistency" to check if the model gives stable answers.

Tom: That self-consistency one is interesting. It's basically a reliability check. If you ask the model the same question multiple times and it gives wildly different answers, the sample might be too ambiguous to be useful for training.

Jane: Precisely. And then they have a "PPL-like difficulty" score, which measures how hard the sample is for the model to learn. So they're balancing a lot of different signals, not just "is this a clean question?"

Tom: It sounds like they're building a really rich profile of each data point. And then the "goal" preset decides how to weight all that.

Jane: Exactly. And the beauty is that the whole process is benchmark-blind. They're not looking at the test set to pick the training data. They're just using these general descriptors to build a better training set.

Tom: That's a crucial point for credibility. They're not cheating by peeking at the answers. So, we've got the framework, we've got the goals. But what's the actual impact? What does this mean for people trying to build these models?

Improvements: Tom: Okay, so we've got the framework and the goals. But let's talk about the actual improvements, because that's where it gets really tangible. We have Lu and Meng on the line to help us break this down. Lu, you're the researcher here—what's the most exciting part of these results for you?

Lu: Thanks, Tom. For me, it's the convergence speed. The paper doesn't just show a better final score; it shows that GDO reaches the baseline's performance *much* earlier in training. On VideoMME, it hits the Uni-ten times reference after just 26 point 6k samples, versus the full 512k. That's a nineteen point two times reduction in data needed to get to the same point.

Meng: And that's not just a lab curiosity. That's a direct cost saving. If you're training on thirty-two H20 GPUs, like they did, cutting the training data by that much means you're done in a fraction of the time. You're freeing up expensive hardware for other experiments.

Tom: So it's not just "less data," it's "faster iteration." You can test more ideas in the same amount of time.

Lu: Exactly. And the gains are concentrated where you'd hope. The biggest improvements are on MVBench and MLVU, which are benchmarks that heavily test temporal reasoning—things like ordering events, counting actions, and understanding state changes. That's the whole point of the "Temp+" profile.

Meng: But I want to push back on that a little. The paper shows LVBench, which is about ultra-long videos, only gets a +zero point eight four pp improvement. That's a lot smaller than the others. So is this just a case of the method failing on longer videos?

Jane: That's a great point, Meng. The paper addresses that directly. LVBench is a different beast. The training pool is mostly short videos and images. So you're asking the model to learn ultra-long video understanding from data that doesn't really contain it.

Lu: Right, it's a distribution mismatch. The tool is working, but it can't create information that isn't in the pool. It's like trying to teach someone to write a novel by only giving them short stories. You can improve their prose, but you can't teach them the structure of a book.

Tom: So the improvement is still positive, but it's capped by the source material. That's a really honest finding, I think.

Meng: It is. And it tells me that if you want better long-video performance, you need to fix the data pool first. GDO is a great allocator, but it's not a miracle worker.

Jane: And that's the key takeaway for me. This isn't a magic bullet; it's a smarter way to use what you have. And the paper shows that the "goal" is a real lever. You can tune the model's behavior by changing the allocation goal.

Lu: Exactly. The ablation study shows that removing the self-consistency score hurts a lot on MLVU, while removing the video-dependence score hurts more on VideoMME. So different benchmarks rely on different cues. It's a very nuanced picture.

Conclusion: Tom: We've covered a lot of ground on "Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning." Jane, if you had to boil this whole paper down to one sentence for someone who just tuned in, what would it be?

Jane: I'd say it's proof that in multimodal AI, being smart about *which* data you use is just as important as *how much* data you use. They've built a system that lets you dial in the exact skills you want to teach, and it does it with a fraction of the data.

Tom: And it's a really clean piece of work. They held everything else constant—the model, the training recipe, the evaluation—and only changed the data. That makes the results super interpretable. It's a controlled experiment.

Meng: And from a practical standpoint, that's huge. It means I can take this framework and apply it to my own training pipeline without having to reinvent the wheel. The code is even available on GitHub.

Lu: It also opens up a new research direction. Instead of just scaling up data, we can now think about "goal-driven" data curation as a first-class design choice. What other goals could we optimize for? Robustness? Fairness? The possibilities are exciting.

Tom: And that's the real takeaway for me. This isn't the end of the story; it's a new beginning. We're moving from "throw everything at the wall" to "let's be architects of our training data."

Jane: Well said, Tom. It's a powerful idea, and the evidence is solid. We'll be watching to see how this framework evolves and how others build on it.

Tom: Absolutely. So, we're going to say goodbye to this paper and get ready to dive into the next one. Thanks to Lu and Meng for joining us, and to all our listeners out there.

Jane: Thanks, everyone. Keep asking big questions, and we'll see you on the next episode!

Peking University · University of Illinois Urbana-Champaign · National University of Singapore

cs.CV, cs.LG

Submitted: 2026-03-12

Updated: 2026-09-27

Comments: Accepted to ECCV 2026

Code: https://github.com/rujiewu/GDO

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: "Modern multimodal assistants have moved from broad visual-language representation learning to instruction-following interaction...

Key concepts

Goal-Driven Data Optimization (GDO)
A framework that separates sample scoring from the desired training goal. It uses a universal quality score for every data sample and then applies different "presets" based on what the model needs to learn, such as focusing on speed or temporal understanding.
Scoring Descriptors
A six-dimensional vector used to evaluate each data sample. These descriptors look at various aspects like optical flow to measure motion, a video-dependence score to check if watching the video is necessary, and self-consistency for answer stability.
Goal Presets
Different configurations within the GDO framework that dictate how the quality scores are weighted. Examples include "MinLoss" for fast training on easy samples and "Temp+" which prioritizes temporal understanding like ordering events and motion.

Terminology

Summary

Summary

The paper introduces Goal-Driven Data Optimization (GDO), a framework designed to improve the sample efficiency and convergence speed of multimodal instruction tuning by optimizing the selection of training data under a fixed training and evaluation contract. The central problem addressed is that multimodal instruction tuning is often compute-inefficient because training budgets are spread across large, mixed image-video pools where the utility of individual samples is highly uneven.

The authors formulate the problem as a data allocation issue: "Modern multimodal assistants have moved from broad visual-language representation learning to instruction-following interaction... a central practical question is how to allocate a fixed supervision budget across samples with unequal training value. They argue that Mixed image-video instruction pools are large, redundant, and heterogeneous, and the value of one more training example is far from uniform."

The GDO method is described as follows: "GDO keeps the model, training recipe, checkpoints, and evaluation fixed, and changes only the training data. It computes six sample descriptors for each candidate and builds optimized 1× subsets under explicit budget, mixture, and source-coverage controls." The six sample descriptors are: Flow (average optical-flow magnitude), VDS (video-dependence score, measured by the loss gap between blind and video-conditioned inputs), Temporal necessity (a question-level proxy for temporal reasoning), Self-consistency (agreement across stochastic decodes), PPL-like difficulty (exponentiated teacher-forced video loss), and Coverage (semantic clustering and source statistics). These descriptors are combined into a shared score that ranks candidates, while goal-specific feasibility presets control budget, video ratio, temporal-positive coverage, and source floors.

The paper reports four released goal profiles: MinLoss (12.9k samples, targets lowest training loss), Diverse (42.9k samples, broader coverage), Temp (33.3k samples, stronger temporal usefulness), and Temp+ (53.3k samples, strongest temporal pressure). All profiles are compared against a fixed 512k-sample Uni-10x baseline under one fixed one-epoch Qwen3-VL-8B-Instruct training recipe on 32 H20 GPUs.

The main results show that GDO achieves higher benchmark accuracy with far fewer samples. Specifically, "Relative to the fixed 512k-sample Uni-10x baseline, GDO reaches the Uni-10x reference after 35.4k samples on MVBench, 26.6k on VideoMME, 27.3k on MLVU, and 34.7k on LVBench, while improving Accuracy by +1.38, +1.67, +3.08, and +0.84 pp, respectively. This corresponds to data reductions of 14.5×, 19.2×, 18.8×, and 14.8×, respectively. The paper states: These peak-match points quantify 'less data, faster convergence' by showing that the optimized subsets reach the useful baseline much earlier while also finishing higher."

The gains are not uniform across benchmarks. The authors note: "MVBench and MLVU improve the most because they are subtask-focused and align well with targeted filtering, while LVBench improves more modestly because it emphasizes ultra-long videos that are only weakly matched by the short-video/image-heavy training pool. The LVBench result is explained by a distribution mismatch: LVBench evaluates ultra-long videos, often much longer than the clips represented in LLaVA-Video, while the training pool also contains a large amount of image QA from LLaVA-OneVision."

Subtask analysis reveals that gains cluster around temporal capabilities. Table 3 shows the largest gains on subtasks like VideoMME Temporal Perception (+18.2 pp for Temp+), MVBench Character Order (+7.00 pp), MLVU Order (+7.14 pp), and MLVU SportsQA (+11.11 pp). The paper notes: the gains are not scattered across arbitrary categories; they cluster around the capabilities targeted by stronger temporal data optimization.

The ablation study (Table 5) shows that Temp+ relies on multiple scoring ingredients rather than one dominant term. Removing self-consistency hurts most on MVBench (-1.40 pp) and MLVU (-3.16 pp), while removing VDS or PPL hurts most on VideoMME (-1.11 and-1.07 pp, respectively). The joint removal is most damaging on MVBench (-2.00 pp), suggesting that Temp+ relies on multiple scoring ingredients and feasibility controls that reinforce the same temporal allocation direction.

The paper's three main claims are: (1) Less Data: Optimized 1× subsets can outperform the fixed Uni-10x baseline with far fewer training samples. (2) Faster Convergence: The gains appear as earlier frontier crossings, not only as stronger reported endpoints. (3) Goal-Driven Data Optimization: Different goals produce different capability profiles, and stronger temporal emphasis yields stronger long-video understanding behavior.

The authors conclude: "The broader implication is that multimodal SFT data should be treated as an allocation problem, not only as a scaling problem. Sample descriptors, shared scoring, and goal-specific feasibility controls provide a practical way to expose this allocation choice while keeping the rest of the training stack fixed. They also acknowledge a key boundary: Data optimization is most effective when the candidate pool contains supervision aligned with the target capability, and native ultra-long-video pools remain an important direction for benchmarks such as LVBench."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and what the improved AI system can do:

Implementation: Add a data-selection layer before SFT that computes six sample descriptors (flow, video-dependence, temporal necessity, self-consistency, PPL-like difficulty, coverage) and builds optimized training subsets under explicit budget, video-ratio, and source-coverage constraints.

What the improved system can do:

  • Train a multimodal model on 12.9k–53.3k samples instead of 512k, reaching the same or better benchmark accuracy.

  • Reduce training compute by 10–19× while improving final accuracy by +0.84 to +3.08 percentage points on video benchmarks.

  • Reach baseline performance after 26.6k–35.4k samples instead of 512k, enabling faster iteration cycles.

Implementation: Add four goal profiles (MinLoss, Diverse, Temp, Temp+) that adjust budget size, selected video ratio, VDS-positive coverage, and temporal-positive floor. Use Temp+ as the default for video-centric tasks.

Implementation: Replace single-quality-score filtering with a six-dimensional descriptor vector per sample, combined into a shared score with separate weights for video and image samples.

Implementation: Keep backbone, optimizer, checkpoint cadence, and evaluation fixed while varying only data allocation. Report peak-match sample counts and reduction ratios alongside final accuracy.

Implementation: Use the subtask delta analysis to monitor and steer capability allocation across motion, ordering, reasoning, and temporal perception categories.

Implementation: Run ablations that remove VDS, PPL, or self-consistency terms to identify which descriptors matter most per benchmark before deployment.


Capability Before (Uni-10x) After (GDO Temp+) Improvement


MVBench accuracy 62.27% 63.65% +1.38 pp

VideoMME accuracy 61.22% 62.89% +1.67 pp

MLVU accuracy 43.81% 46.89% +3.08 pp

LVBench accuracy 40.22% 41.06% +0.84 pp

Training samples needed 512k 53.3k 9.6× fewer

Samples to reach baseline — 26.6k–35.4k 14.5–19.2× reduction

Temporal perception (VideoMME) 67.27% 85.45% +18.18 pp

Character order (MVBench) 67.50% 74.50% +7.00 pp

Order (MLVU) 25.71% 32.86% +7.14 pp

The improved system can now train video-language models with a fraction of the data, converge faster, and achieve higher accuracy on temporal reasoning tasks, while providing interpretable control over the data-allocation strategy.

Sources

Related papers