What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs".
Jane: The paper was written by Ziran Li, Qiang Wang, Zhengyu Chen, Shanglin Lei, Borun Chen et al. from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary and Implications: Tom: Now that we understand the core concept of the Capability-Driven Multimodal Scaling Law, let's look at what this framework tells us about how to actually use it.
Jane: The authors provide some very actionable insights, such as identifying a specific "transfer tax" in certain benchmarks.
Lu: It turns out that certain textual benchmarks, for example those involving complex symbolic execution or rigid formats, have negative transfer coefficients (lambda j < zero).
Meng: That's a critical warning to share with other engineers because it means you can't just look at the raw leaderboard score and assume high numbers are helping your VLM performance.
Lalam: The concept of the "transfer tax" highlights that optimizing for text-specific leaderboards can be counterproductive when we need continuous visual alignment, making the whole process much more nuanced.
Tom: We also found something surprising: base LLMs tend to perform better as VLM backbones than instruction-tuned versions, despite the instruction tuning being widely accepted.
Jane: They attribute this to what they call the "alignment tax," where the extra effort of instruction tuning reduces our data-scaling efficiency compared to a simpler base model.
Lu: This suggests that for large, expensive multimodal training runs, we might actually want a simpler, less refined base LLM backbone than we initially expected based on its textual scores.
Meng: It’s a very practical trade-off: the overhead of instruction tuning might be negated by the fact that it requires significantly more data to achieve the same performance level as its base counterpart.
Lalam: The paper gives us a quantitative way to look at this, mapping out these distinct transfer and absorption positions so we can make an informed choice based on our resource constraints.
Tom: This focus on actionable insights truly changes how we approach backbone selection in a way that is principled rather than experimental.
Improvements and Insights: Tom: So, we’ve seen how the Capability-Driven Multimodal Scaling Law works conceptually, and now we're digging into its specific findings for real-world model selection.
Jane: The paper suggests several improvements to how we think about performance prediction, such as showing that the LLM backbone isn't just a static input.
Lu: The authors demonstrate that the capacity of this foundational knowledge can be directly predicted even for models far larger than those used in testing, which is a huge jump in predictive power.
Meng: This is a massive step toward planning because we can now extrapolate performance for entirely unseen backbones, which is invaluable when budgeting or scope changes.
Lalam: The way the knowledge transfers from the text-only phase to the multimodal phase shows that we are no longer just looking at raw model size, but at a true transfer mechanism.
Tom: And we've seen how this works across thirty-four different LLM backbones, which is an impressive dataset for reliability.
Jane: It has also shown us that different model families occupy very distinct positions in the transfer-absorption space, giving us a roadmap to guide our choices.
Lu: The data shows that Llama has a high-transfer, low-capacity regime compared to Qwen3’s low-transfer, high-capacity regime, which is fascinating from an architectural perspective.
Meng: This implies that we should tailor our expectations based on the family type rather than just the size, making implementation much more predictable.
Lalam: The paper gives us confidence in how these foundations translate into a future that is not just brute force, but intelligently designed.
Tom: It’s great to see how this study provides a unified framework for multimodal scaling across different architectural types and scales at all times.
Conclusion and Wrap-up: Tom: We've covered so much ground today, from the initial idea of what transfers from text to vision, all the way to the practical implications for backbone selection.
Jane: It is clear that this new law is providing a truly principled way to approach VLM development, giving us confidence in our choices.
Lu: The consistency across different model families shown in their analysis really validates this, proving that the framework isn't just a fluke of one specific architecture.
Meng: It allows us to extrapolate performance for entirely unseen backbones, which is a massive step forward for planning and budget allocation on these massive projects.
Lalam: We are moving away from guessing toward the understanding that the foundational intelligence of a model is its true predictor of its potential, regardless of how many more parameters it has.
Tom: Before we wrap up this discussion, does anyone have a final thought on the impact of this work?
Lu: I think this opens up incredible possibilities for how we can design future AI systems by using the statistical laws to guide our creativity.
Meng: It gives us a clear roadmap for optimization, ensuring that when we spend resources, they are spent in the most effective way possible.
Lalam: The entire process of creating vision-language models is becoming more elegant and less reliant on trial and error, thanks to this work on transfer dynamics.
Tom: A great summary of the paper's contribution to a final time; I want to thank all of you for discussing "What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs" with me today.
Jane: It has truly shifted how we think about the relationship between language and vision, Tom, and I'm genuinely excited to see where this research takes us next.
Lu: The potential is immense; I feel that way when I look at these new dynamics.
Meng: We should be able to build smarter, more efficient systems now that we have these tools.
Lalam: To a more informed and powerful future for AI, everyone!
Conclusion: Tom: We've been exploring a really profound shift in how we approach multimodal AI, moving away from simply guessing which large language model backbone works best.
Jane: Exactly, and by using the Capacity-Driven Multimodal Scaling Law, we can now make those decisions with a level of quantitative certainty that feels incredibly empowering.
Meng: From an engineering standpoint, this means we can plan our entire compute budget for VLM training based on measurable textual benchmarks rather than just relying on arbitrary model size.
Lu: It’s exciting to realize that the cross-family success—seeing how thirty-four different models behave—suggest that we are moving toward a truly universal understanding of knowledge transfer across different architectural families.
Lalam: This research allows us to build systems whose foundations are not just powerful, but also highly reflective of human reasoning capabilities and cultural contexts.
Meng: It gives us a practical way to optimize resource allocation, knowing exactly how much data we need based on the LLM's initial performance score.
Lu: I think the sheer scope of this work highlights that capability is truly independent of size, which is a wild idea when looking at all those different model lines.
Jane: It’s comforting to see this predictability, knowing that we aren't just hoping for the best outcome, but aiming for it through a principled design process.
Tom: The full title of the paper, "What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs," really captures this transition from guesswork to a powerful framework.
Lalam: A more informed and culturally aligned future for AI, it is.
Meng: We should be able to build smarter, more efficient systems now that we have these tools in our hands.
Tom: I’m genuinely excited to see what the next paper has in store for us as we head into our break.
cs.CL, cs.AI, cs.CV
Submitted: 2026-06-24
Updated: 2026-10-08
Code: https://github.com/wangq-dev/CDMScaling
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 89/100
The gist: This paper introduces the Capability-Driven Multimodal Scaling Law, a novel framework designed to predict the performance of Vision-Language Models (VLMs) based on the textual capabilities of their
Key concepts
- Capability-Driven Multimodal Scaling Law
- This framework offers a unified way to understand how knowledge transfers from text to vision in large language models (VLMs). It allows researchers to predict performance and map out distinct transfer and absorption positions across different model architectures.
- Transfer Tax
- This refers to negative transfer coefficients found when certain textual benchmarks are used. Optimizing for these specific text-based leaderboards can be counterproductive when a continuous visual alignment is required, indicating that raw score increases do not always translate to better VLM performance.
- Alignment Tax
- This concept explains why base LLMs sometimes perform better as VLM backbones than instruction-tuned versions. The extra effort and refinement involved in instruction tuning can reduce the data-scaling efficiency compared to a simpler, less refined base model.
Terminology
Summary
This paper introduces the Capability-Driven Multimodal Scaling Law, a novel framework designed to predict the performance of Vision-Language Models (VLMs) based on the textual capabilities of their Large Language Model (LLM) backbones. By moving away from unreliable compute-based metrics and parameter counts, which fail to generalize across model families,
this research provides a principled, quantitative method for backbone selection and resource allocation in multimodal development.
The Proposed Framework
The authors propose a framework that predicts multimodal benchmark accuracy by coupling the textual capability of an LLM backbone with the volume of multimodal training data. To avoid the redundancy of using raw multidimensional benchmarks, they apply Principal Component Analysis (PCA) to extract a low-dimensional capability score S
from over 200 textual benchmarks. This score serves as a proxy for intrinsic capability, capturing both parameter count and pre-training data volume.
The core of the framework is an end-to-end predictor that models multimodal performance through two primary mechanisms:
((
-
Transfer: The first term,
 · S,
represents thecapability transferred from the textual backbone to the multimodal setting,
acting as a starting point for training. -
Absorption: The second term,
B̂ · ln Dmm,
models performance gains from multimodal data, where B̂ is an effective absorption rate that captures how efficiently a model turns additional data into improvements.
)
Empirical Validation and Extrapolation
To validate the law, the researchers trained over 150 VLMs on 34 LLMs spanning seven different model families under a strictly controlled recipe.
The results demonstrate that the capability-driven scaling law achieves significantly higher fidelity than classical compute-based laws. Specifically, it can accurately extrapolate transfer rates from models up to 8B parameters to much larger 72B-scale backbones
and predict full training trajectories with high precision. The framework's robustness is further proven through leave-one-family-out validation,
where it successfully predicts the performance of entirely held-out model lineages.
Key Insights into Transfer Dynamics
The analysis surfaces several actionable insights regarding how textual intelligence translates to visual understanding:
((
-
Transfer Tax: The researchers identified certain textual benchmarks that act as a
transfer tax,
meaning improvements in these specific dimensions actually fail to translate to multimodal performance, potentially due tolatent benchmark-gaming behavior.
-
The Instruction-Tuning Disparity: Contrary to intuition, base LLMs often serve as superior VLM backbones compared to instruction-tuned counterparts. While instruction tuning provides a higher initial transfer slope, it incurs an
alignment tax
that results in a higherabsorption decay rate,
meaning these models saturate faster during multimodal training. -
Family-Specific Profiles: Different model families occupy distinct positions in the
transfer–absorption space,
such as the Llama family'shigh-transfer, low-capacity
regime versus the Qwen3 family'slow-transfer, high-capacity
regime.
)
Practical Applications
Beyond theoretical modeling, the framework enables practical engineering decisions. It supports optimal joint selection of backbone and data budget under a compute constraint
and allows for hyperparameter extrapolation.
For instance, the researchers demonstrated that the optimal batch size for multimodal fine-tuning can be estimated directly from a model's composite scaling metric, allowing practitioners to avoid costly empirical sweeps
when scaling up to massive model sizes.
)
Improvements for AI systems
To implement the findings of this paper, I would transition our development pipeline from an empirical trial-and-error
approach to a principled, predictive engineering workflow.
Here are the specific improvements and their capabilities:
- Implement a
Capability-Driven Backbone Selection
Engine
Instead of selecting LLM backbones based on parameter count or unweighted leaderboard averages, I will integrate a PCA-based capability scoring system (using the paper's method of extracting latent factors from 200 textual benchmarks).
The improved system can:
-
Predict the multimodal accuracy of a candidate VLM before any multimodal training begins.
-
Identify
Transfer Taxes
—textual capabilities that appear to improve LLM scores but actually degrade VLM alignment (e.g., over-optimizing for rigid symbolic execution or specific text formats)—allowing us to discard backbones that aregaming
benchmarks at the expense of visual reasoning.
- Deploy an Adaptive Data-Scaling Strategy (Absorption-Aware Training)
I will replace static training schedules with a dynamic scheduler based on the Absorption Rate
formula:
If a backbone has high textual capability, we will trigger early saturation protocols to avoid wasting compute; if it has low absorption, we will allocate significantly more multimodal tokens.
The improved system can:
-
Optimize the multimodal data budget (Dmm) for any given backbone scale to ensure maximum ROI on GPU hours.
-
Predict exactly when a model will hit performance saturation, preventing
diminishing returns
compute waste.
- Adopt
Base-Model First
Multimodal Pre-training
Based on the paper’s discovery of the Alignment Tax,
I will mandate that all high-scale multimodal training uses Base LLMs rather than Instruct/Chat versions as the backbone.
The improved system can:
-
Achieve superior data-scaling efficiency by utilizing backbones with higher absorption rates and lower decay.
-
Avoid the
geometric flexibility
loss caused by text-specific instruction tuning, ensuring more residual capacity for continuous visual embedding alignment.
- Automated Hyperparameter Extrapolation (Zero-Shot Tuning)
I will implement the paper’s composite scaling metric to automate the selection of training hyperparameters like Batch Size across different model scales.
The improved system can:
- Predict the optimal batch size for a massive 72B+ parameter VLM based solely on its textual benchmark scores and planned data volume, eliminating the need for expensive, multi-scale hyperparameter sweeps.
Sources
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Qwen3-VL Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Seed1.5-VL Technical Report
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- The Llama 3 Herd of Models
- Gemma 2: Improving Open Language Models at a Practical Size
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Mistral 7B
- LLaVA-OneVision: Easy Visual Task Transfer
- Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Gemma 3 Technical Report
- Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering