Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging

arXiv:2604.11399 · cs.CV, cs.CL · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging".

Jane: The paper was written by Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi and Jiaying Wu from National University of Singapore and MBZUAI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we're digging into a paper that's got a title that really grabs you: "Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging." Jane, I have to say, that title alone tells you the authors believe something very specific about where intelligence lives in these models.

Jane: It really does, Tom. And that's exactly what got me excited when I first read it. The team at National University of Singapore and MBZUAI—Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi, and Jiaying Wu—they're making a bold claim. They're saying that when you take a large language model and turn it into a video model, you might be damaging the reasoning circuits that were already there.

Tom: Right, and that's the puzzle. You take a perfectly good text model, you bolt on a visual encoder, you train it to understand videos, and suddenly it can't do temporal reasoning anymore. It fails on logic questions about time that the original text model handled fine.

Jane: Exactly. And the title tells you where they think the problem lives—in specific layers of the network. Not everywhere. Not uniformly. Just in certain spots. So their solution is to go in and surgically replace those specific layers with the ones from the original text model.

Tom: It's like if you had a great chef who could cook amazing meals, and then you taught them to also be a great sommelier, but in the process they forgot how to season food properly. The fix isn't to retrain the whole chef—it's to remind them of the seasoning techniques they already knew.

Jane: That's a great way to put it, Tom. And the paper's called MERIT, which stands for Merging for Enhanced Reasoning In Temporal Tasks. The whole idea is that you don't need to retrain anything. You just find the right layers to swap back.

Tom: And that's the part that gets me excited. No retraining means no massive GPU bills, no new data collection, no weeks of fine-tuning. You just merge parameters from the original model back into the video model at the right spots.

Jane: But finding the right spots is the hard part, right? That's what the whole paper is about. They use an evolutionary search algorithm to figure out which layers matter most for temporal reasoning.

Tom: And we're going to get into exactly how that works in a minute. But first, let me just say—the fact that they can do this without any training at all is a big deal for the field.

Jane: It really is. And the results are pretty striking. We'll talk about the numbers soon, but let me just tease this—they improved temporal reasoning by nearly twenty-four percent on one benchmark for one of their models.

Tom: Twenty-four percent! That's not a small bump. That's a real recovery of lost capability. So stick around, because we're about to get into the meat of how they actually pulled this off.

Summary: Jane: So we've set the stage with the title, and now let's actually walk through what this paper does. "Reasoning Resides in Layers" is built on a really simple observation that has big consequences.

Tom: And that observation is that video-language models—VLMs, as the paper calls them—they're built by taking a text-only language model and adding a visual encoder. The language model already knows how to reason about time when it's just reading text.

Jane: But once you train it on videos, something breaks. The paper shows this really clearly in their opening figure. You give the base language model a temporal logic question—something like "Emma mailed the letter after buying a stamp but before calling a taxi. What happened last?"—and it gets it right.

Tom: But then you take the same language model after it's been adapted into a video model, give it the exact same text question, and it gets it wrong. Same model, same question, different answer. That's wild.

Jane: It really is. And the authors argue this happens because multimodal training is dominated by spatial perception and object recognition. The model gets really good at seeing things, but the supervision for temporal abstraction and causal inference is much weaker.

Tom: So you end up with a model that can tell you there's a kettle and a sink in the video, but can't figure out why the person walked over to the sink in the first place.

Jane: Exactly. And that's where MERIT comes in. Instead of retraining the model to fix this, they merge parameters from the original text-only backbone back into the video model. But they don't merge everything uniformly—that's the key insight.

Tom: Right, because if you just averaged all the parameters together, you'd probably lose some of the visual capabilities you just gained. The paper shows that uniform merging is unreliable.

Jane: So they do something smarter. They search over layer-wise recipes. Each layer of the self-attention mechanism in the language backbone can either be taken mostly from the video model or mostly from the text model.

Tom: And they use an evolutionary algorithm called CMA-ES to search this space. It's a black-box optimizer, which is perfect here because the objective—temporal reasoning accuracy—isn't differentiable.

Jane: But here's the clever part. Their objective isn't just to maximize reasoning. It's to maximize reasoning while penalizing any drop in temporal perception. So they're explicitly protecting the visual capabilities while restoring the reasoning ones.

Tom: And the results speak for themselves. On Video-MME, they improved temporal reasoning by nearly twenty-four percent for LongVA-7B, about eleven percent for InternVL3-8B, and almost four percent for Qwen3-VL-4B. All while keeping perception roughly the same.

Jane: And those gains transferred to other benchmarks too—LongVideoBench, LVBench, MMBench-Video, Video-Holmes. The recipes they found on one dataset worked on others.

Tom: So the reasoning recovery isn't just memorizing the search set. It's actually restoring a general capability. That's the part that makes me think this could be really important.

Jane: And it's all training-free. No gradient updates, no new data, no fine-tuning. Just parameter merging at the right layers.

Improvements: Tom: So we've covered what MERIT does, but let's get into the improvements it brings over existing approaches. Jane, what stood out to you?

Jane: The biggest improvement is the layer selectivity itself. Before this paper, if you wanted to merge a text model back into a vision model, you'd typically merge everything uniformly. The paper shows that's just not good enough.

Tom: And they prove it with their baselines. They compared against merging all layers uniformly, and against randomly selecting the same number of layers. Both performed worse than MERIT. Random selection had high variance—sometimes it helped, sometimes it hurt.

Jane: That's the key finding. It's not just about how many layers you merge, but which ones. The paper shows that reasoning-related behavior is concentrated in specific layers, not spread evenly across the network.

Tom: And they have a really nice experiment to prove this. They took the layers that MERIT selected and masked them out in the original video model. When they did that, reasoning performance dropped much faster than overall performance or non-reasoning metrics.

Jane: For LongVA-7B, when they fully masked the selected layers, reasoning dropped by over fifty-one percent, while overall performance only dropped by about seventeen percent. That's a huge gap.

Tom: So those layers are disproportionately important for reasoning. That's direct evidence that reasoning does reside in specific layers, just like the title says.

Jane: And there's another improvement that I think is really elegant. They use a continuous gating vector to represent which layers to merge, then threshold it to get a discrete recipe. This lets the evolutionary search explore smoothly while still producing concrete merged models.

Tom: They also restricted the interpolation weight to a small set of values and ran separate searches for each. That keeps the search tractable while still giving good results.

Jane: And the efficiency gains are real. They cached objective scores for identical layer configurations and cached visual processor outputs, since merging only changes the language backbone. That gave them a two to three times speedup.

Tom: But the improvement that excites me most is the behavioral shift they observed. They did frame-level attribution analysis on Video-Holmes, which is a reasoning-heavy benchmark with suspense films.

Jane: And they found that the base model tends to anchor on locally salient but misleading events. Like in one case, the base model focused on a character making a phone call and concluded the victim died because help didn't arrive in time.

Tom: But MERIT focused on two temporally separated but causally informative moments—the victim being restrained, and later being found motionless. And it correctly inferred the victim was killed by the murderer, even though the killing itself wasn't shown.

Jane: That's a fundamental shift in how the model grounds its reasoning. It's not just getting more answers right—it's attending to the right evidence.

Tom: And that's the kind of improvement that could generalize beyond benchmarks. If the model is actually reasoning over temporal structure, it should handle new situations better.

Conclusion: Jane: Alright, Tom, let's wrap this up. We've been talking about "Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging" for a while now, and I think we should pull together what we've learned.

Tom: The core message is that temporal reasoning in video-language models isn't lost forever when you adapt a text model. It's just buried under layers that got overwritten by perception-focused training. And you can dig it back out by selectively merging parameters from the original text backbone.

Jane: No retraining. No new data. Just a smart search over which layers to restore. And the results are consistent across three different model families and five different benchmarks.

Tom: The gains were especially strong for LongVA-7B—nearly twenty-four percent improvement in temporal reasoning on Video-MME. And those gains transferred to other benchmarks, which tells us the recovery is real and general.

Jane: And the interventional experiments were really convincing. Masking the layers MERIT selected caused reasoning to collapse much faster than other capabilities. That's strong evidence that reasoning really does concentrate in specific layers.

Tom: The attribution analysis showed something even more interesting. The merged model doesn't just answer more correctly—it actually looks at different parts of the video. It focuses on causally relevant moments rather than locally salient but misleading ones.

Jane: That's the kind of change that suggests the model is genuinely reasoning better, not just memorizing patterns.

Tom: So what does this mean for the field? Well, it suggests that a lot of capability loss during multimodal adaptation might be recoverable with targeted interventions. We might not need to retrain everything every time we want to add a new capability.

Jane: And that's a big deal for practical deployment. Training-free fixes are fast, cheap, and don't require massive compute budgets. They could make video-language models much more practical for real-world applications.

Tom: There are limitations, of course. The recipes are model-specific, so you have to run the search for each new model. And the search depends on having good benchmarks for temporal reasoning and perception.

Jane: But the direction is really promising. This paper opens up a whole line of research into layer-selective merging as a general tool for capability recovery.

Tom: And with that, we're going to say goodbye to "Reasoning Resides in Layers" and get ready for our next paper. Thanks for listening, everyone!

Jane: See you next time!

Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi, Jiaying Wu

National University of Singapore · MBZUAI

cs.CV, cs.CL

Submitted: 2026-08-16

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 76/100

Key concepts

Video-Language Models (VLMs)
These are models created by taking a text-only language model and adding a visual encoder to understand videos. The paper suggests that adapting these models can damage their original temporal reasoning capabilities.
Layer-Selective Merging
This is the solution proposed by MERIT. Instead of uniformly merging all parameters, the method searches for specific layers in the network where parameters from the original text model should be merged back into the video model to restore lost reasoning.
Evolutionary Search Algorithm (CMA-ES)
The paper uses this black-box optimizer to find which layer configurations are best. It searches over layer-wise recipes, maximizing temporal reasoning while penalizing any drop in visual perception.

Terminology

Summary

Summary

This paper introduces MERIT (Merging for Enhanced Reasoning In Temporal Tasks), a training-free, task-driven model merging framework designed to restore temporal reasoning (TR) in video-language models (VLMs) by selectively merging self-attention layers between a VLM and its paired text-only backbone.

The paper identifies a core problem: Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from language-only pretraining. This trade-off is especially pronounced in video-language models (VLMs), where visual alignment can impair temporal reasoning (TR) over sequential events. The authors attribute this to "a mismatch in adaptation objectives: existing multimodal training pipelines are dominated by spatial perception and object-level alignment, while providing much weaker supervision for temporal abstraction and causal inference."

The proposed method, MERIT, searches over layer-wise self-attention merging recipes between a VLM and its paired text-only backbone using an objective that improves TR while penalizing degradation in temporal perception (TP). The objective function is defined as: g* = g F(g), where F(g) = Acc TR(g) - lambda times D TP(g), and D TP(g) = (0, Acc base TP - Acc TP(g)). This formulation favors recipes that improve temporal reasoning while preserving temporal perception.

The search is performed using CMA-ES (Covariance Matrix Adaptation Evolution Strategy) over a continuous gating vector g in [0,1] L, where each dimension corresponds to a self-attention layer. The gates are thresholded to produce discrete recipes, and merging uses directional interpolation: theta = alpha theta M + (1-alpha) theta N if = 1, otherwise theta = (1-alpha) theta M + alpha theta N, with alpha in [0.5, 1].

Experiments were conducted on three VLMs: LongVA-7B, InternVL3-8B, and Qwen3-VL-4B-Instruct, each paired with its corresponding text-only backbone. The search was performed on a targeted subset of Video-MME (55 TP examples, 177 TR examples), and evaluation was extended to full Video-MME, LongVideoBench, LVBench, MMBench-Video, and Video-Holmes.

Key results show that MERIT consistently improves TR, preserves or improves TP, and generalizes beyond the search set to four distinct benchmarks. On Video-MME, MERIT achieved relative TR gains of +23.9% for LongVA-7B, +10.8% for InternVL3-8B, and +3.8% for Qwen3-VL-4B, while preserving or improving TP. The method also outperformed uniform full-model merging (ALL-LAYER) and random layer selection (RANDOM-K), demonstrating that effective recovery depends on selecting the right layers.

Interventional masking experiments showed that the layers selected by MERIT are disproportionately important for reasoning. For LongVA-7B at kappa = 0.4, masking selected layers caused Reasoning to drop by 35.1%, compared with 16.8% for Overall and 2.7% for Others. Frame-level attribution analysis revealed that MERIT shifts model decisions toward temporally and causally relevant evidence, moving away from locally salient but misleading cues.

The paper concludes that targeted, perception-aware model merging can effectively restore TR in VLMs without retraining, and that "reasoning deficits introduced during multimodal adaptation can be addressed through targeted architectural intervention, pointing toward a scalable path for recovering specialized capabilities in future multimodal systems."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems and the resulting capabilities:

1. Layer-Selective Reasoning Restoration Module

  • Implement a training-free, task-driven merging framework that selectively interpolates self-attention parameters between a VLM and its paired text-only backbone at specific layers

  • Use CMA-ES evolutionary search with an objective function: F(g) = Acc TR(g) - λ · max(0, Acc TP base - Acc TP(g))

  • This restores temporal reasoning without degrading temporal perception

Resulting capability: The improved VLM can correctly answer temporal reasoning questions (e.g., What happened last? or Why did the person go to the sink?) with relative gains of +23.9% (LongVA-7B), +10.8% (InternVL3-8B), and +3.8% (Qwen3-VL-4B) on Video-MME, while preserving or improving temporal perception.

2. Reasoning-Critical Layer Identification

  • Add interventional layer-masking analysis to identify which self-attention layers disproportionately support temporal reasoning

  • Scale selected layer outputs by coefficient κ ∈ [0,1] to measure reasoning degradation

3. Temporally-Grounded Decision Attribution

  • Implement frame-level gradient-activation attribution at the answer token to measure which visual evidence supports the final decision

  • Use logit margin between selected option and strongest alternative as attribution target

4. Capability-Recovery Framework for Multimodal Adaptation

  • Apply the layer-selective merging principle to any multimodal model (not just video) to recover reasoning abilities lost during visual/audio alignment

  • Use the same objective: reward target capability gains while penalizing degradation in the perceptual capability

5. Efficient Black-Box Optimization for Model Merging

  • Use CMA-ES with continuous gating vectors thresholded to discrete layer selections

  • Cache objective scores for identical layer configurations and visual processor outputs (2-3× speedup)

6. Generalizable Reasoning Recovery

  • Discover merging recipes on a small targeted subset (55 TP + 177 TR examples from Video-MME) that transfer to four additional benchmarks (LongVideoBench, LVBench, MMBench-Video, Video-Holmes)

  • Merging formula: θ l = α·θ l M + (1-α)·θ l N when layer selected (ĝ=1), with α ∈ 1.0, 0.9, 0.8, 0.7, 0.6, 0.5

  • Search space: L-dimensional gating vector where L = number of self-attention layers (28 for LongVA/InternVL3, 36 for Qwen3-VL)

  • Population size: λ pop = 4 + ⌊3·ln(L)⌋

  • Evaluation protocol: Decoding temperature = 0, frame caps of 128 (LongVA) or 64 (InternVL3, Qwen3-VL)

Abstract

Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from language-only pretraining. This trade-off is especially pronounced in video-language models (VLMs), where visual alignment can impair temporal reasoning (TR) over sequential events. We propose MERIT, a training-free, task-driven model merging framework for restoring TR in VLMs. MERIT searches over layer-wise self-attention merging recipes between a VLM and its paired text-only backbone using an objective that improves TR while penalizing degradation in temporal perception (TP). Across three representative VLMs and multiple challenging video benchmarks, MERIT consistently improves TR, preserves or improves TP, and generalizes beyond the search set to four distinct benchmarks. It also outperforms uniform full-model merging and random layer selection, showing that effective recovery depends on selecting the right layers. Interventional masking and frame-level attribution further show that the selected layers are disproportionately important for reasoning and shift model decisions toward temporally and causally relevant evidence. These results show that targeted, perception-aware model merging can effectively restore TR in VLMs without retraining.

Sources

Related papers