Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs".
Jane: Detailed Research Summary: "Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs" This paper investigates a critical failure mode in Video Large Language Models (Video-LLMs):
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Speaking of that struggle, let’s look at what the paper actually proposes to fix this issue, which is where they get really inventive. They introduce a training objective called DeltaDirect to help solve the direction binding gap. Jane, can you explain what DeltaDirect actually does in simple terms?
Jane: Sure. The core idea of DeltaDirect is to force the projector output—that's the part that translates the video data into something the LLM can read—to carry a much stronger signal about how things are moving. They achieve this by training it to predict two-D motion vectors based on small changes in features between adjacent frames <ref:2605.22823#pg0>.
Lu: That’s clever because it takes information that's already there, like feature deltas from neighboring frames, and uses that to create a more reliable signature for the movement direction before it even enters the LLM. It’s essentially boosting the projector's ability to communicate motion explicitly.
Meng: From an engineering standpoint, that sounds like a way to engineer robustness into the intermediate representation rather than just tweaking the final readout layer, which is where we often end up wasting time. So they are trying to fix it upstream in the pipeline.
Lalam: If we think about our development, this suggests that strengthening those intermediate signal carriers is a more effective path than just relying on massive instruction tuning alone to fix every single failure point in the final decoding step. It’s about making sure the input to the LLM is as clear as possible regarding direction.
The paper's summary: Tom: So, summarizing what we’ve discussed so far, the paper explains that directional motion blindness is a failure where the necessary information for motion direction is present throughout the entire processing pipeline, yet it fails to bind correctly during final readout. Jane, how would you put that 'direction binding gap' simply?
Jane: Think of it like this: imagine you have a map showing every road and direction, but when you get to the exit sign, the sign itself is blurry or misaligned so your brain can't link "this path" to "this destination." The paper says the visual information about motion direction is available everywhere along the route, but at that final binding point, that connection breaks down.
Lu: It’s interesting because they didn't find that the information was missing from any single component; it’s a systemic mismatch in signal strength across those different parts of the system. They used difference-in-means concept vectors to show this magnitude deficit when testing on more complex, out-of-domain scenarios.
Meng: That magnitude deficit idea is important because it explains why models trained only on easy, synthetic data don't translate well when they see something really messy in the real world. It suggests the learned understanding of motion direction structure isn't strong enough to hold up under varied visual noise.
Lalam: I see how that relates to our work with robustness; if the signal strength drops significantly in complex domains, then any defense we build must account for that drop, not just assume the signal is always there at full strength.
The paper's improvements: Tom: Now let’s talk about the actual fixes they propose—the improvements they test. They suggest using instruction tuning alongside DeltaDirect to tackle this. Jane, what does this combined approach achieve in terms of measurable results?
Jane: When they combine instruction tuning on a specific dataset like MODIRECT-INST with the addition of DeltaDirect, their average motion direction accuracy gets a boost of six point five points over the baseline model <ref:2605.22823#pg1>. Most notably, adding DeltaDirect provides the largest gain, which is an eleven point two point improvement when testing on Cutout-on-Real videos <ref:2605.22823#pg1>.
Lu: That eleven point two point jump on Cutout-on-Real is substantial because that benchmark tests generalization on real-world scenes that aren't just clean synthetic images <ref:2605.22823#pg1>. It suggests DeltaDirect helps the model generalize its understanding of motion direction much better than instruction tuning alone would by itself.
Meng: I’m looking at the transferability results, and it seems promising; DeltaDirect improves accuracy by twenty-one point nine points over the vanilla baseline when tested on MODIRECT-REALBENCH, even without seeing real videos during the initial training phase <ref:2605.22823#pg1>. That points to a really solid mechanism for transferring knowledge across different visual domains.
Lalam: It’s encouraging because it shows that this technique isn't just overfitting to synthetic motion; it’s actually building a more resilient representation of direction that can be applied broadly. This kind of targeted signal strengthening feels like a practical step toward making these systems more reliable for general use.
Conclusion: Tom: So, wrapping up the discussion on "Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs," we see that the core issue is a readout binding gap caused by signal weakness in complex domains, and DeltaDirect is a method to strengthen the projector output to fix it. Jane, what’s your final thought on the impact of closing this gap?
Jane: I think solving this means we move past models that are good at identifying objects but fail at understanding their physical relationship in motion. It allows these systems to perform more reliably when faced with the messy, unpredictable videos we actually see every day.
Lu: The implications for creative applications are huge; imagine video understanding tools that can accurately interpret complex, subtle movements in dynamic scenes without those frustrating blind spots. We could unlock new ways for AI to interact with and describe the physical world in a much richer way.
Meng: From an engineering view, this means we can design training objectives that specifically target these types of signal weaknesses rather than just hoping they emerge naturally during standard fine-tuning. It gives us a clearer roadmap for improving system stability when moving from controlled lab settings to real applications.
Lalam: For me, it means our culture shifts toward demanding higher fidelity in multimodal understanding; we need systems that don't just see things, but truly perceive the dynamics of their environment reliably. This work points toward a necessary evolution in how we build these sophisticated AI companions.
Kyung Hee University · Princeton University
cs.CV
Submitted: 2026-05-21
Updated: 2026-10-07
Comments: NeurIPS 2026 (Accept). 50 pages including Appendix. Project page: https://jong980812.github.io/which-way-did-it-move/
Code: https://github.com/KHU-VLL/DeltaDirect
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: This paper investigates a critical failure mode in Video Large Language Models (Video-LLMs): directional motion blindness, which is formally diagnosed as a "direction binding gap." This gap signifies
Key concepts
- Direction Binding Gap
- This is a failure where the model knows how to perceive motion direction from visual data, but it cannot reliably map that direction to the correct verbal answer choice during final output. It's a mismatch between available information and the final decision-making step.
- Magnitude Deficit Hypothesis
- This hypothesis explains why models fail on new, complex data. It suggests that while the model learns motion direction structures generally, the actual strength or 'magnitude' of those learned signals drops significantly when tested in novel contexts, leading to weak directional cues.
- DeltaDirect Objective
- This is a novel training method applied to the projector output. It uses derived 2-D motion vectors from adjacent frames to train the model to predict stronger, more reliable signed displacement signals. This forces the visual-language interface to carry a clearer directional signal before it reaches the main LLM.
Terminology
Summary
This paper investigates a critical failure mode in Video Large Language Models (Video-LLMs): directional motion blindness, which is formally diagnosed as a direction binding gap.
This gap signifies that while the necessary information for motion direction is linearly accessible across the entire processing pipeline—from the initial vision encoder, through the projector output, to intermediate LLM hidden states—this directional signal fails to reliably bind to the correct verbal answer option during final readout.
The research systematically diagnoses this problem and proposes a multi-faceted intervention strategy involving instruction tuning and a novel training objective.
The core finding is that the failure is not due to missing information, but rather a mismatch in signal strength across domains.
-
Linear Accessibility vs. Readout Failure: Motion direction remains linearly decodable from the vision encoder, projector output, and LLM hidden states. However, the final readout token fails to encode the correct mapping from that direction to the corresponding answer option (the
direction binding gap
). -
Instruction Tuning Effects: Instruction tuning on a controlled dataset (MODIRECT-INST) substantially improves recognition on the source domain. However, this improvement is not universal; models trained only on source data degrade significantly when tested on more complex, out-of-domain (OOD) scenarios, such as Cutout-on-Real.
-
The Magnitude Deficit Hypothesis: Analysis using difference-in-means motion direction concept vectors reveals the root cause of the OOD generalization gap. While the model learns the structure of motion direction across domains, the magnitude of this signal drops sharply on complex domains. This suggests that out-of-domain failure is not a lack of geometric understanding, but a signal weakness—the learned motion direction structure is too weak to be reliably read out when tested in novel contexts.
The paper introduces two key components to address this magnitude deficit: instruction tuning and the DeltaDirect objective.
This serves as a baseline improvement, showing that targeted instruction can align motion direction concept vectors across different domains, even if it doesn't fully solve the OOD magnitude issue on its own.
To directly combat the signal weakness identified in Section 4.4, DeltaDirect is introduced as a training-only, projector-level objective designed to strengthen signed displacement cues at the visual-language interface before the information enters the LLM.
-
Mechanism: DeltaDirect uses analytically available 2-D motion vectors derived from adjacent-frame feature deltas (supervised by MODIRECT-INST) to predict normalized 2-D motion vectors from these projector-feature deltas.
-
Goal: This forces the projector output to carry a stronger, more reliable signed displacement signal.
-
Performance Gains: DeltaDirect is shown to substantially improve performance: it nearly solves Primitive-on-Syn (99.5% accuracy) and significantly boosts performance on challenging domains. Specifically, it improves Cutout-on-Real accuracy from 60.5% (baseline) to 71.7% and raises the MODIRECT-SYNBENCH average accuracy from 78.9% to 85.4%. Crucially, DeltaDirect is shown to narrow the direction binding gap across visual complexity.
The paper employs rigorous causal interventions to confirm why the on-axis projection matters:
-
Axis Specificity: Interventions are performed at specific LLM layers (l=21 in LLaVA-Video) by rescaling the readout state along different concept vectors: Canonical (motion direction d,A), Random, Magnitude-only (rescaling norm without axis reprojection), and Wrong-anti (antipodal direction d-, A).
-
Key Finding: Only the canonical axis recovers MCQ accuracy across all OOD domains with a large gain (+14.6 pp on Cutout-on-Real). Random and magnitude-only interventions yield negligible effects (+1.4 pp), ruling out generic norm-scaling as the solution. The antipodal intervention catastrophically degrades accuracy (up to-39.3 pp), confirming that pushing the readout state along an incorrect direction axis actively misleads the answer-option binding mechanism.
Improvements for AI systems
Based on this scientific paper, here are the specific improvements for AI systems and what these improved systems will be able to do:
) Improve core video understanding by resolving Directional Motion Blindness.
This means moving beyond simple object recognition or general motion identification to accurately determining the signed image-plane motion direction (left/right/up/down).
) Implement a multi-stage diagnostic and intervention framework:
-
Perform linear probing across the Vision Encoder, Projector, and LLM hidden states to precisely localize where directional information is lost.
-
Use Difference-in-Means concept vector analysis to diagnose why the signal weakens in out-of-domain (OOD) scenarios (magnitude deficit vs. orientation loss).
-
Introduce a targeted training objective (DeltaDirect) that strengthens the projector output by predicting normalized 2D motion vectors from adjacent frame feature deltas, ensuring signed displacement cues are robustly passed to the LLM before language processing.
) Increase generalization across diverse visual contexts: The improved system will maintain high accuracy on complex, real-world video scenes (Cutout-on-Real) where vanilla models fail, by explicitly learning a magnitude-invariant motion signal during training.
) Achieve superior performance in complex reasoning tasks: By closing the direction binding gap,
the model will reliably link perceived motion direction to the correct verbal answer option (e.g., binding leftward
to the correct letter/word), leading to significantly higher accuracy in multiple-choice and open-ended motion reasoning questions.
) Enhance robustness against input variations: The system will be less susceptible to noise and visual complexity, as it learns that direction information is encoded in low-variance dimensions of the visual features, making it robust even when foreground/background complexity changes drastically.
) Enable effective transfer learning: The DeltaDirect objective allows models trained on simple synthetic data (Primitive-on-Syn) to effectively transfer their motion understanding capabilities to complex real-world scenarios (MODIRECT-REALBENCH) without requiring expensive, domain-specific real-world tuning data.
Abstract
Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed image-plane motion direction. On simple videos of a single object moving left, right, up, or down, most Video-LLMs perform near chance, with above-chance cases largely attributable to prediction biases rather than genuine direction understanding. We call this failure directional motion blindness. We localize the failure by tracing motion direction information through the Video-LLM pipeline. Motion direction remains linearly accessible from the vision encoder, projector, and LLM hidden states, but the readout fails to bind this signal to the correct verbal answer option, revealing a direction binding gap. Although synthetic motion direction instruction tuning reduces this gap on the source domain, motion direction concept vector analysis shows that visual complexity weakens the signal magnitude and limits out-of-domain generalization. We introduce MoDirect, a dataset family for motion direction instruction tuning and evaluation, and DeltaDirect, a diagnosis-driven, projector-level objective that predicts normalized 2-D motion vectors from adjacent-frame feature deltas. On MoDirect-SynBench, instruction tuning with DeltaDirect improves motion direction accuracy from 25.9% to 85.9%. On MODIRECT-REALBENCH, DeltaDirect improves realworld motion direction accuracy by 21.4 points over the vanilla baseline without real-world tuning data, while preserving standard video-understanding performance. Our project page is available at https://jong980812.github.io/which-way-did-it-move/
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Lost in Time: A New Temporal Benchmark for VideoLLMs
- GPT-4o System Card
- MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Qwen2 Technical Report
- Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models