Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
summary
The gist
This paper investigates a critical failure mode in Video Large Language Models (Video-LLMs): directional motion blindness, which is formally diagnosed as a "direction binding gap." This gap signifies
In short
The research diagnosed directional motion blindness in Video-LLMs, where motion direction information is accessible but fails to bind correctly to answer options, especially in new scenarios. The solution introduced DeltaDirect, a training objective that strengthens signed displacement cues at the projector level. This intervention significantly improved performance on complex out-of-domain tasks by boosting the signal strength of motion direction.
Key concepts
- Direction Binding Gap
- This is a failure where the model knows how to perceive motion direction from visual data, but it cannot reliably map that direction to the correct verbal answer choice during final output. It's a mismatch between available information and the final decision-making step.
- Magnitude Deficit Hypothesis
- This hypothesis explains why models fail on new, complex data. It suggests that while the model learns motion direction structures generally, the actual strength or 'magnitude' of those learned signals drops significantly when tested in novel contexts, leading to weak directional cues.
- DeltaDirect Objective
- This is a novel training method applied to the projector output. It uses derived 2-D motion vectors from adjacent frames to train the model to predict stronger, more reliable signed displacement signals. This forces the visual-language interface to carry a clearer directional signal before it reaches the main LLM.
Terminology used across episodes
This episode discusses
- Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs · Paper Radio
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Lost in Time: A New Temporal Benchmark for VideoLLMs
- GPT-4o System Card
- MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Qwen2 Technical Report
- Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
The paper
Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs · Read on arXiv
Kyung Hee University · Princeton University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs".
Jane: Detailed Research Summary: "Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs" This paper investigates a critical failure mode in Video Large Language Models (Video-LLMs):
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Speaking of that struggle, let’s look at what the paper actually proposes to fix this issue, which is where they get really inventive. They introduce a training objective called DeltaDirect to help solve the direction binding gap. Jane, can you explain what DeltaDirect actually does in simple terms?
Jane: Sure. The core idea of DeltaDirect is to force the projector output—that's the part that translates the video data into something the LLM can read—to carry a much stronger signal about how things are moving. They achieve this by training it to predict two-D motion vectors based on small changes in features between adjacent frames <ref:2605.22823#pg0>.
Lu: That’s clever because it takes information that's already there, like feature deltas from neighboring frames, and uses that to create a more reliable signature for the movement direction before it even enters the LLM. It’s essentially boosting the projector's ability to communicate motion explicitly.
Meng: From an engineering standpoint, that sounds like a way to engineer robustness into the intermediate representation rather than just tweaking the final readout layer, which is where we often end up wasting time. So they are trying to fix it upstream in the pipeline.
Lalam: If we think about our development, this suggests that strengthening those intermediate signal carriers is a more effective path than just relying on massive instruction tuning alone to fix every single failure point in the final decoding step. It’s about making sure the input to the LLM is as clear as possible regarding direction.
The paper's summary: Tom: So, summarizing what we’ve discussed so far, the paper explains that directional motion blindness is a failure where the necessary information for motion direction is present throughout the entire processing pipeline, yet it fails to bind correctly during final readout. Jane, how would you put that 'direction binding gap' simply?
Jane: Think of it like this: imagine you have a map showing every road and direction, but when you get to the exit sign, the sign itself is blurry or misaligned so your brain can't link "this path" to "this destination." The paper says the visual information about motion direction is available everywhere along the route, but at that final binding point, that connection breaks down.
Lu: It’s interesting because they didn't find that the information was missing from any single component; it’s a systemic mismatch in signal strength across those different parts of the system. They used difference-in-means concept vectors to show this magnitude deficit when testing on more complex, out-of-domain scenarios.
Meng: That magnitude deficit idea is important because it explains why models trained only on easy, synthetic data don't translate well when they see something really messy in the real world. It suggests the learned understanding of motion direction structure isn't strong enough to hold up under varied visual noise.
Lalam: I see how that relates to our work with robustness; if the signal strength drops significantly in complex domains, then any defense we build must account for that drop, not just assume the signal is always there at full strength.
The paper's improvements: Tom: Now let’s talk about the actual fixes they propose—the improvements they test. They suggest using instruction tuning alongside DeltaDirect to tackle this. Jane, what does this combined approach achieve in terms of measurable results?
Jane: When they combine instruction tuning on a specific dataset like MODIRECT-INST with the addition of DeltaDirect, their average motion direction accuracy gets a boost of six point five points over the baseline model <ref:2605.22823#pg1>. Most notably, adding DeltaDirect provides the largest gain, which is an eleven point two point improvement when testing on Cutout-on-Real videos <ref:2605.22823#pg1>.
Lu: That eleven point two point jump on Cutout-on-Real is substantial because that benchmark tests generalization on real-world scenes that aren't just clean synthetic images <ref:2605.22823#pg1>. It suggests DeltaDirect helps the model generalize its understanding of motion direction much better than instruction tuning alone would by itself.
Meng: I’m looking at the transferability results, and it seems promising; DeltaDirect improves accuracy by twenty-one point nine points over the vanilla baseline when tested on MODIRECT-REALBENCH, even without seeing real videos during the initial training phase <ref:2605.22823#pg1>. That points to a really solid mechanism for transferring knowledge across different visual domains.
Lalam: It’s encouraging because it shows that this technique isn't just overfitting to synthetic motion; it’s actually building a more resilient representation of direction that can be applied broadly. This kind of targeted signal strengthening feels like a practical step toward making these systems more reliable for general use.
Conclusion: Tom: So, wrapping up the discussion on "Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs," we see that the core issue is a readout binding gap caused by signal weakness in complex domains, and DeltaDirect is a method to strengthen the projector output to fix it. Jane, what’s your final thought on the impact of closing this gap?
Jane: I think solving this means we move past models that are good at identifying objects but fail at understanding their physical relationship in motion. It allows these systems to perform more reliably when faced with the messy, unpredictable videos we actually see every day.
Lu: The implications for creative applications are huge; imagine video understanding tools that can accurately interpret complex, subtle movements in dynamic scenes without those frustrating blind spots. We could unlock new ways for AI to interact with and describe the physical world in a much richer way.
Meng: From an engineering view, this means we can design training objectives that specifically target these types of signal weaknesses rather than just hoping they emerge naturally during standard fine-tuning. It gives us a clearer roadmap for improving system stability when moving from controlled lab settings to real applications.
Lalam: For me, it means our culture shifts toward demanding higher fidelity in multimodal understanding; we need systems that don't just see things, but truly perceive the dynamics of their environment reliably. This work points toward a necessary evolution in how we build these sophisticated AI companions.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck