Scaling Native Multimodal Pre-Training From Scratch
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Scaling Native Multimodal Pre-Training From Scratch".
Tom: The gist: Native multimodal pre-training establishes compute-optimal scaling laws for transformer-based vision-language models by showing that language and multimodal objectives follow distinct allocation laws, one composition-invariant and the other composition-variant.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So we’ve talked about how native multimodal pre-training is trying to find its optimal scaling laws. This paper "Scaling Native Multimodal Pre-Training From Scratch" tackles the challenge of characterizing these scaling properties when you train vision-language models from scratch on both modalities together.
Jane: The core thesis they are making is that language and multimodal objectives follow different allocation laws. Specifically, they claim the language allocation law is largely invariant to how you mix the multimodal data, while the multimodal allocation law is highly sensitive to that composition.
Lu: This sensitivity means that if you use a lot of text data, you need bigger models to be compute-efficient. If you use more vision data, it changes things differently.
Meng: So they are essentially proving that there isn't just one way to allocate the computational budget when training these kinds of models from scratch.
Tom: They investigate this by looking at the optimal model size and token count under a fixed computational budget C, showing that the minimal objective loss follows a predictable compute law.
Jane: And they use two estimators for this, an IsoFLOP profile and the training-curve envelope to cross-validate their findings on where those optimal points are located.
Lu: Beyond just the scaling dynamics, they also evaluate how native multimodal pre-training impacts downstream tasks like language understanding and spatial reasoning.
Tom: The key finding here is that this method improves pure-text spatial reasoning, showing that the spatial understanding you get from multimodal training actually generalizes to unimodal text tasks.
Jane: So for anyone listening, it matters because this provides the foundational infrastructure needed to guide native multimodal pre-training from scratch for future scaling efforts.
Meng: It’s a lot of complexity to manage all those variables at once while trying to find the right configuration. That sounds like a heavy lift for real-world implementation.
Lu: The authors provide a language-multimodal Pareto frontier, which offers quantitative guidance on the optimal model size, text token count, and multimodal token count for any given budget C.
Tom: So essentially, they give researchers a clear map to follow when designing native multimodal pre-training from scratch based on computational constraints.
Jane: And we need to keep an eye on how this scaling works because it sets the baseline for building the next generation of multimodal foundation models.
Conclusion: Tom: To wrap up, we’ve looked at the paper "Scaling Native Multimodal Pre-Training From Scratch". The authors are Haoyuan Wu, Aoqi Wu, Hai Wang, Jiajia Wu, Jinxiang Ou, and Bei Yu.
Jane: The big implication is that they’ve mapped out how to scale these models predictably by showing that the language objective scales differently than the multimodal one.
Lu: So it gives us a concrete way to understand which scaling levers—model size or token count—to pull based on whether we are emphasizing text or vision data.
Meng: It helps ground the theoretical scaling laws in practical considerations, which is useful for engineers trying to deploy these systems efficiently within a fixed budget Ctotal.
Tom: And that’s what they did: they derived compute-optimal scaling laws by showing how minimal objective loss behaves under different conditions, leading to power laws for model sizes and token counts.
Jane: It moves the conversation from just *if* native multimodal pre-training works to *how* we can do it efficiently when training from scratch.
Lu: The language-multimodal Pareto frontier is a concrete tool that researchers can use right away to pinpoint the exact configuration for their training runs.
Meng: I think it gives us a practical target for our engineers to start testing configurations rather than just guessing and hoping they work out on the first try.
Tom: It’s about moving from guesswork to knowing where the optimal scaling lies, which is what this research achieves with "Scaling Native Multimodal Pre-Training From Scratch".
The Chinese University of Hong Kong
cs.CL, cs.CV
Submitted: 2026-07-24
Updated: 2026-10-08
Importance score: 92/100
The gist: The gist: Native multimodal pre-training establishes compute-optimal scaling laws for transformer-based vision-language models by showing that language and multimodal objectives follow distinct
Key concepts
- Language Allocation Law
- This law describes how computational resources are distributed for language tasks during training. The study found this allocation is largely independent of the multimodal data used, meaning learning text representations scales consistently regardless of how much visual data is present in the training set.
- Multimodal Allocation Law
- This law governs resource allocation for multimodal objectives, which is highly dependent on the composition of the training data. Specifically, mixtures with more text require larger model capacities to be compute-efficient, shifting optimal scaling towards bigger models.
- Language-Multimodal Pareto Frontier
- This frontier represents a quantitative guide showing the trade-off between optimizing language and multimodal objectives under a fixed budget. It helps researchers determine the best model size and token count by balancing the distinct scaling behaviors of these two learning goals.
Terminology
Summary
The gist: Native multimodal pre-training establishes compute-optimal scaling laws for transformer-based vision-language models by showing that language and multimodal objectives follow distinct allocation laws, one composition-invariant and the other composition-variant.
How it works
The research investigates the optimal model size (N) and token count (D) for training a transformer-based vision-language model under a fixed computational budget C, demonstrating that minimal objective loss adheres to a predictable compute law while compute-optimal model sizes and token counts scale as power laws <ref:2607.22043#pg2> The language allocation law is largely invariant to the composition of the data, indicating stable language learning regardless of the multimodal data ratio <ref:2607.22043#pg2> Conversely, the multimodal allocation law is highly sensitive to this composition <ref:2607.22043#pg2> Specifically, text-heavy mixtures become compute-efficient only at larger model scales, shifting the optimal resource allocation toward greater model capacity <ref:2607.22043#pg2>
The study employs two estimators to derive scaling laws for the compute-optimal frontier:
-
IsoFLOP profiles serve as the primary estimator, where plotting the loss of each model against log N generates an IsoFLOP profile, and a parabola accurately models each profile with its minimum identifying the optimal model size Nopt(C) <ref:2607.22043#pg4>.
-
The training-curve envelope provides an independent cross-validation mechanism, where pooling the trajectories of all evaluated models and extracting the lower envelope yields the minimum achievable loss for any given budget C <ref:2607.22043#pg6>.
Scaling Dynamics of Objectives
The analysis reveals sharp differences in scaling behaviors between language and multimodal objectives <ref:2607.22043#pg2> The language allocation law remains largely invariant to data composition, suggesting that text representations are learned consistently per token, regardless of the accompanying multimodal data <ref:2607.22043#pg2> In contrast, the multimodal allocation law is highly sensitive to this composition <ref:2607.22043#pg2> Text-heavy data mixtures are only compute-efficient for larger models, which shifts the optimal allocation toward increased model capacity <ref:2607.22043#pg2> This relationship is captured by a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
Joint Optimization and Trade-offs
While decoupling the allocation problem yields mechanistic insights, practical native multimodal pre-training must ultimately optimize the shared parameter count N under a unified computational budget Ctotal <ref:2607.22043#pg2> The researchers employ an asymmetric modeling framework, pairing a composition-invariant language objective with a composition-variant multimodal objective <ref:2607.22043#pg2> This asymmetry reflects the inherent structural divergence between the two modalities <ref:2607.22043#pg2> Consequently, the computational efficiency of the language objective remains largely independent of r; forcing it to fluctuate would introduce overfitting and localized optimization noise <ref:2607.22043#pg2> Conversely, cross-modal alignment is highly sensitive to data composition <ref:2607.22043#pg2>.
Downstream Implications
The downstream evaluations reveal that native multimodal pre-training induces positive cross-modal transfer, thereby enhancing pure-text spatial reasoning and enabling robust multimodal in-context learning <ref:2607.22043#pg2> This improvement in pure-text spatial reasoning demonstrates that spatial understanding acquired through multimodal training successfully generalizes to unimodal text tasks <ref:2607.22043#pg2> Furthermore, this paradigm exhibits multimodal in-context learning capabilities analogous to those observed in traditional LLMs <ref:2607.22043#pg2>. The few-shot gains concentrate on spatial reasoning, showing the most pronounced and consistent improvements in spatial reasoning benchmarks <ref:2607.22043#pg2> This suggests that the templates primarily assist the model in resolving the spatial and relational structure of a query, rather than enhancing fine-grained visual recognition or text extraction <ref:2607.22043#pg2>.
The paper establishes the essential groundwork for predictably scaling multimodal foundation models by deriving compute-optimal scaling laws for native multimodal pre-training <ref:2607.22043#pg2>. The main contributions are summarized as follows:
• We independently analyze the scaling behaviors of language and multimodal objectives, revealing distinct allocation laws for each modality <ref:2607.22043#pg2>.
• We identify a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
• We empirically demonstrate that native multimodal pre-training improves pure-text spatial reasoning and enables robust multimodal in-context learning <ref:2607.22043#pg2>. This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models >
--- Page 19 ---
The paper establishes the essential groundwork for predictably scaling multimodal foundation models by deriving compute-optimal scaling laws for native multimodal pre-training <ref:2607.22043#pg2>. The main contributions are summarized as follows:
• We independently analyze the scaling behaviors of language and multimodal objectives, revealing distinct allocation laws for each modality <ref:2607.22043#pg2>.
• We identify a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
• We empirically demonstrate that native multimodal pre-training improves pure-text spatial reasoning and enables robust multimodal in-context learning <ref:2607.22043#pg2>. This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models >
--- Page 18 ---
The paper establishes the essential groundwork for predictably scaling multimodal foundation models by deriving compute-optimal scaling laws for native multimodal pre-training <ref:2607.22043#pg2>. The main contributions are summarized as follows:
• We independently analyze the scaling behaviors of language and multimodal objectives, revealing distinct allocation laws for each modality <ref:2607.22043#pg2>.
• We identify a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
• We empirically demonstrate that native multimodal pre-training improves pure-text spatial reasoning and enables robust multimodal in-context learning <ref:2607.22043#pg2>. This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models >
--- Page 17 ---
The paper establishes the essential groundwork for predictably scaling multimodal foundation models by deriving compute-optimal scaling laws for native multimodal pre-training <ref:2607.22043#pg2>. The main contributions are summarized as follows:
• We independently analyze the scaling behaviors of language and multimodal objectives, revealing distinct allocation laws for each modality <ref:2607.22043#pg2>.
• We identify a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
• We empirically demonstrate that native multimodal pre-training improves pure-text spatial reasoning and enables robust multimodal in-context learning <ref:2607.22043#pg2>. This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models >
--- Page 16 ---
The paper establishes the essential groundwork for predictably scaling multimodal foundation models by deriving compute-optimal scaling laws for native multimodal pre-training <ref:2607.22043#pg2>. The main contributions are summarized as follows:
• We independently analyze the scaling behaviors of language and multimodal objectives, revealing distinct allocation laws for each modality <ref:2607.22043#pg2>.
• We identify a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
• We empirically demonstrate that native multimodal pre-training improves pure-text spatial reasoning and enables robust multimodal in-context learning <ref:2607.22043#pg2>. This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models >
--- Page 15 ---
The paper establishes the essential groundwork for predictably scaling multimodal foundation models by deriving compute-optimal scaling laws for native multimodal pre-training <ref:2607.22043#pg2>. The main contributions are summarized as follows:
• We independently analyze the scaling behaviors of language and multimodal objectives, revealing distinct allocation laws for each modality <ref:2607.22043#pg2>.
• We identify a language-multimodal Pareto frontier, offering quantitative guidance for the optimal scaling of native multimodal pre-training <ref:2607.22043#pg2>.
• We empirically demonstrate that native multimodal pre-training improves pure-text spatial reasoning and enables robust multimodal in-context learning <ref:2607.22043#pg2>.
Improvements for AI systems
-
Cross-modal transfer enhancement: The improved system will demonstrate
positive cross-modal transfer, thereby enhancing pure-text spatial reasoning
because native multimodal pre-training enablesspatial understanding acquired through multimodal training successfully generalizes to unimodal text tasks.
-
Robust in-context learning: The AI can exhibit
multimodal in-context learning capabilities analogous to those observed in traditional LLMs
by leveraging the findings that models show a performance gain thatemerges with model scaling,
specifically noting that for larger models,the performance gap emerges early and widens steadily as the training progresses.
-
Optimized resource allocation: The system will utilize a
joint compute-optimal frontier
derived from the analysis of decoupled objectives to determine the optimal configurations ofmodel size, token count, and data mixture
for a fixed budget. This allows precise determination of where to allocate resources based on composition, as shown by the finding thattext-heavy data mixtures become compute-efficient only at larger model scales.
Sources
- Training Verifiers to Solve Math Word Problems
- Emu3.5: Native Multimodal Models are World Learners
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- Measuring Massive Multitask Language Understanding
- Training Compute-Optimal Large Language Models
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Scaling Laws for Neural Language Models
- Kimi K2.5: Visual Agentic Intelligence
- DeepSeek-V3 Technical Report
- SocialIQA: Commonsense Reasoning about Social Interactions
- Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
- CountQA: How Well Do MLLMs Count in the Wild?
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- HellaSwag: Can a Machine Really Finish Your Sentence?
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering