Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging
summary
In short
The episode discusses a paper titled "Reasoning Resides in Layers," which proposes a method called MERIT for restoring temporal reasoning in video-language models by selectively merging parameters from an original text model. The hosts explain that this training-free approach improves temporal reasoning by nearly twenty-four percent on one benchmark, showing that reasoning is concentrated in specific layers.
Key concepts
- Video-Language Models (VLMs)
- These are models created by taking a text-only language model and adding a visual encoder to understand videos. The paper suggests that adapting these models can damage their original temporal reasoning capabilities.
- Layer-Selective Merging
- This is the solution proposed by MERIT. Instead of uniformly merging all parameters, the method searches for specific layers in the network where parameters from the original text model should be merged back into the video model to restore lost reasoning.
- Evolutionary Search Algorithm (CMA-ES)
- The paper uses this black-box optimizer to find which layer configurations are best. It searches over layer-wise recipes, maximizing temporal reasoning while penalizing any drop in visual perception.
Terminology used across episodes
This episode discusses
- Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging · Paper Radio
- GPT-4 Technical Report
- Qwen3-VL Technical Report
- ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- Lost in Time: A New Temporal Benchmark for VideoLLMs
- Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
- Arcee's MergeKit: A Toolkit for Merging Large Language Models
- Temporal Reasoning Transfer from Text to Video
- Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
- Qwen2.5 Technical Report
- TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
- Qwen2 Technical Report
- Qwen3 Technical Report
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
- Long Context Transfer from Language to Vision
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
The paper
Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging · Read on arXiv
Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi, Jiaying Wu
National University of Singapore · MBZUAI
Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from language-only pretraining. This trade-off is especially pronounced in video-language models (VLMs), where visual alignment can impair temporal reasoning (TR) over sequential events. We propose MERIT, a training-free, task-driven model merging framework for restoring TR in VLMs. MERIT searches over layer-wise self-attention merging recipes between a VLM and its paired text-only backbone using an objective that improves TR while penalizing degradation in temporal perception (TP). Across three representative VLMs and multiple challenging video benchmarks, MERIT consistently improves TR, preserves or improves TP, and generalizes beyond the search set to four distinct benchmarks. It also outperforms uniform full-model merging and random layer selection, showing that effective recovery depends on selecting the right layers. Interventional masking and frame-level attribution further show that the selected layers are disproportionately important for reasoning and shift model decisions toward temporally and causally relevant evidence. These results show that targeted, perception-aware model merging can effectively restore TR in VLMs without retraining.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging".
Jane: The paper was written by Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi and Jiaying Wu from National University of Singapore and MBZUAI.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone! Today we're digging into a paper that's got a title that really grabs you: "Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging." Jane, I have to say, that title alone tells you the authors believe something very specific about where intelligence lives in these models.
Jane: It really does, Tom. And that's exactly what got me excited when I first read it. The team at National University of Singapore and MBZUAI—Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi, and Jiaying Wu—they're making a bold claim. They're saying that when you take a large language model and turn it into a video model, you might be damaging the reasoning circuits that were already there.
Tom: Right, and that's the puzzle. You take a perfectly good text model, you bolt on a visual encoder, you train it to understand videos, and suddenly it can't do temporal reasoning anymore. It fails on logic questions about time that the original text model handled fine.
Jane: Exactly. And the title tells you where they think the problem lives—in specific layers of the network. Not everywhere. Not uniformly. Just in certain spots. So their solution is to go in and surgically replace those specific layers with the ones from the original text model.
Tom: It's like if you had a great chef who could cook amazing meals, and then you taught them to also be a great sommelier, but in the process they forgot how to season food properly. The fix isn't to retrain the whole chef—it's to remind them of the seasoning techniques they already knew.
Jane: That's a great way to put it, Tom. And the paper's called MERIT, which stands for Merging for Enhanced Reasoning In Temporal Tasks. The whole idea is that you don't need to retrain anything. You just find the right layers to swap back.
Tom: And that's the part that gets me excited. No retraining means no massive GPU bills, no new data collection, no weeks of fine-tuning. You just merge parameters from the original model back into the video model at the right spots.
Jane: But finding the right spots is the hard part, right? That's what the whole paper is about. They use an evolutionary search algorithm to figure out which layers matter most for temporal reasoning.
Tom: And we're going to get into exactly how that works in a minute. But first, let me just say—the fact that they can do this without any training at all is a big deal for the field.
Jane: It really is. And the results are pretty striking. We'll talk about the numbers soon, but let me just tease this—they improved temporal reasoning by nearly twenty-four percent on one benchmark for one of their models.
Tom: Twenty-four percent! That's not a small bump. That's a real recovery of lost capability. So stick around, because we're about to get into the meat of how they actually pulled this off.
Summary: Jane: So we've set the stage with the title, and now let's actually walk through what this paper does. "Reasoning Resides in Layers" is built on a really simple observation that has big consequences.
Tom: And that observation is that video-language models—VLMs, as the paper calls them—they're built by taking a text-only language model and adding a visual encoder. The language model already knows how to reason about time when it's just reading text.
Jane: But once you train it on videos, something breaks. The paper shows this really clearly in their opening figure. You give the base language model a temporal logic question—something like "Emma mailed the letter after buying a stamp but before calling a taxi. What happened last?"—and it gets it right.
Tom: But then you take the same language model after it's been adapted into a video model, give it the exact same text question, and it gets it wrong. Same model, same question, different answer. That's wild.
Jane: It really is. And the authors argue this happens because multimodal training is dominated by spatial perception and object recognition. The model gets really good at seeing things, but the supervision for temporal abstraction and causal inference is much weaker.
Tom: So you end up with a model that can tell you there's a kettle and a sink in the video, but can't figure out why the person walked over to the sink in the first place.
Jane: Exactly. And that's where MERIT comes in. Instead of retraining the model to fix this, they merge parameters from the original text-only backbone back into the video model. But they don't merge everything uniformly—that's the key insight.
Tom: Right, because if you just averaged all the parameters together, you'd probably lose some of the visual capabilities you just gained. The paper shows that uniform merging is unreliable.
Jane: So they do something smarter. They search over layer-wise recipes. Each layer of the self-attention mechanism in the language backbone can either be taken mostly from the video model or mostly from the text model.
Tom: And they use an evolutionary algorithm called CMA-ES to search this space. It's a black-box optimizer, which is perfect here because the objective—temporal reasoning accuracy—isn't differentiable.
Jane: But here's the clever part. Their objective isn't just to maximize reasoning. It's to maximize reasoning while penalizing any drop in temporal perception. So they're explicitly protecting the visual capabilities while restoring the reasoning ones.
Tom: And the results speak for themselves. On Video-MME, they improved temporal reasoning by nearly twenty-four percent for LongVA-7B, about eleven percent for InternVL3-8B, and almost four percent for Qwen3-VL-4B. All while keeping perception roughly the same.
Jane: And those gains transferred to other benchmarks too—LongVideoBench, LVBench, MMBench-Video, Video-Holmes. The recipes they found on one dataset worked on others.
Tom: So the reasoning recovery isn't just memorizing the search set. It's actually restoring a general capability. That's the part that makes me think this could be really important.
Jane: And it's all training-free. No gradient updates, no new data, no fine-tuning. Just parameter merging at the right layers.
Improvements: Tom: So we've covered what MERIT does, but let's get into the improvements it brings over existing approaches. Jane, what stood out to you?
Jane: The biggest improvement is the layer selectivity itself. Before this paper, if you wanted to merge a text model back into a vision model, you'd typically merge everything uniformly. The paper shows that's just not good enough.
Tom: And they prove it with their baselines. They compared against merging all layers uniformly, and against randomly selecting the same number of layers. Both performed worse than MERIT. Random selection had high variance—sometimes it helped, sometimes it hurt.
Jane: That's the key finding. It's not just about how many layers you merge, but which ones. The paper shows that reasoning-related behavior is concentrated in specific layers, not spread evenly across the network.
Tom: And they have a really nice experiment to prove this. They took the layers that MERIT selected and masked them out in the original video model. When they did that, reasoning performance dropped much faster than overall performance or non-reasoning metrics.
Jane: For LongVA-7B, when they fully masked the selected layers, reasoning dropped by over fifty-one percent, while overall performance only dropped by about seventeen percent. That's a huge gap.
Tom: So those layers are disproportionately important for reasoning. That's direct evidence that reasoning does reside in specific layers, just like the title says.
Jane: And there's another improvement that I think is really elegant. They use a continuous gating vector to represent which layers to merge, then threshold it to get a discrete recipe. This lets the evolutionary search explore smoothly while still producing concrete merged models.
Tom: They also restricted the interpolation weight to a small set of values and ran separate searches for each. That keeps the search tractable while still giving good results.
Jane: And the efficiency gains are real. They cached objective scores for identical layer configurations and cached visual processor outputs, since merging only changes the language backbone. That gave them a two to three times speedup.
Tom: But the improvement that excites me most is the behavioral shift they observed. They did frame-level attribution analysis on Video-Holmes, which is a reasoning-heavy benchmark with suspense films.
Jane: And they found that the base model tends to anchor on locally salient but misleading events. Like in one case, the base model focused on a character making a phone call and concluded the victim died because help didn't arrive in time.
Tom: But MERIT focused on two temporally separated but causally informative moments—the victim being restrained, and later being found motionless. And it correctly inferred the victim was killed by the murderer, even though the killing itself wasn't shown.
Jane: That's a fundamental shift in how the model grounds its reasoning. It's not just getting more answers right—it's attending to the right evidence.
Tom: And that's the kind of improvement that could generalize beyond benchmarks. If the model is actually reasoning over temporal structure, it should handle new situations better.
Conclusion: Jane: Alright, Tom, let's wrap this up. We've been talking about "Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging" for a while now, and I think we should pull together what we've learned.
Tom: The core message is that temporal reasoning in video-language models isn't lost forever when you adapt a text model. It's just buried under layers that got overwritten by perception-focused training. And you can dig it back out by selectively merging parameters from the original text backbone.
Jane: No retraining. No new data. Just a smart search over which layers to restore. And the results are consistent across three different model families and five different benchmarks.
Tom: The gains were especially strong for LongVA-7B—nearly twenty-four percent improvement in temporal reasoning on Video-MME. And those gains transferred to other benchmarks, which tells us the recovery is real and general.
Jane: And the interventional experiments were really convincing. Masking the layers MERIT selected caused reasoning to collapse much faster than other capabilities. That's strong evidence that reasoning really does concentrate in specific layers.
Tom: The attribution analysis showed something even more interesting. The merged model doesn't just answer more correctly—it actually looks at different parts of the video. It focuses on causally relevant moments rather than locally salient but misleading ones.
Jane: That's the kind of change that suggests the model is genuinely reasoning better, not just memorizing patterns.
Tom: So what does this mean for the field? Well, it suggests that a lot of capability loss during multimodal adaptation might be recoverable with targeted interventions. We might not need to retrain everything every time we want to add a new capability.
Jane: And that's a big deal for practical deployment. Training-free fixes are fast, cheap, and don't require massive compute budgets. They could make video-language models much more practical for real-world applications.
Tom: There are limitations, of course. The recipes are model-specific, so you have to run the search for each new model. And the search depends on having good benchmarks for temporal reasoning and perception.
Jane: But the direction is really promising. This paper opens up a whole line of research into layer-selective merging as a general tool for capability recovery.
Tom: And with that, we're going to say goodbye to "Reasoning Resides in Layers" and get ready for our next paper. Thanks for listening, everyone!
Jane: See you next time!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language