Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting

arXiv:2608.11655 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Xikai Sun, Kebin Liu, Haotian Wang, Li Liu, Xu Wang, Yunhao Liu

Tsinghua University · JD Logistics

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/SunVictor23/MaP

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation.

Terminology

Summary

Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However, multimodal large language models (MLLMs) typically process videos through sparse uniform sampling to control visual-token and attention costs. This strategy may discard critical transitions between sampled frames, limiting reasoning about object movement, collisions, and causal interactions. To mitigate this issue, we propose Motion-as-Prompt (MaP), a track-guided cross-frame visual prompting framework. MaP recovers dense point trajectories, selects motion-informative frames, and marks the trajectories accumulated between consecutive sampled frames directly onto the visual inputs, making otherwise hidden displacement, direction changes, and interactions observable to frozen MLLMs. Experiments on CLEVRER and Something-Something-v2 show that MaP consistently improves average motion-reasoning accuracy, yielding gains of 4.2% and 8.9% for GPT-5.5, respectively. Notably, these improvements are obtained without degrading non-motion understanding, highlighting the robustness of MaP. These results demonstrate that MaP provides a simple and effective solution for enhancing motion-centric video reasoning without model training or architectural modification.

The paper identifies inter-frame motion loss as an input-side bottleneck: sparse uniform sampling converts a continuous motion process into a sequence of snapshots, making critical transitions (acceleration, direction changes, collisions) invisible to the MLLM. Existing approaches do not directly mitigate this loss: motion-aware video models require additional training or architectural access, while existing pixel-level visual prompts are predominantly frame-local (annotating object regions, identities, or spatial relations within individual frames) and do not reveal how objects move across unsampled intervals. The key observation is that discarded motion can be recovered from the original video and re-encoded into sparse visual inputs.

MaP operates as follows. Given a full-frame-rate video, it first recovers dense point trajectories using a frozen point tracker (CoTracker3, tracking a 10×10 grid of query points, re-seeding every second). It then compensates for global camera motion using a global similarity transformation (translation, rotation, scale) fitted via linear least squares over visible points, yielding camera-compensated object velocities. From these velocities, it computes a scalar motion energy score per frame along three dimensions: speed, acceleration, and curvature (with the turn angle weighted by the smaller adjacent velocity to avoid spurious curvature from static points). These descriptors are normalized by their 95th percentile, equally weighted, smoothed with a length-3 moving average, and re-normalized to [0,1].

Motion-guided sampling then selects frames under a fixed budget B. It first uniformly samples na = max(2, ⌈αB⌉) anchors (α = 1/4) for global coverage. The remaining slots are filled by greedy non-maximum suppression over motion peaks, with a minimum temporal gap r = max(1, ⌊N/(2B)⌋) between selected frames. If NMS returns fewer than B frames, remaining slots are filled in descending order of motion energy without the gap constraint.

Inter-frame trajectory marking renders the query point motion between adjacent sampled frames onto the later frame. Only points with maximum displacement over the following ∆ frames exceeding τ = 0.03 of the frame width are retained, and at most K points are kept to control annotation density. Each trajectory segment is drawn as a single-colored polyline with a circular marker at its endpoint on the later sampled frame.

Experiments compare MaP against uniform sampling, AKS, FOCUS, SoM, and GoM on CLEVRER (five sub-tasks: object existence, moving direction, moving count, moving attribute, counterfactual inference) and SSv2 (four-way multiple-choice with object references abstracted to something), plus TempCompass for non-motion generalization. On CLEVRER, MaP achieves the highest average accuracy for both Qwen3-VL-2B (60.4% vs 57.6% uniform) and GPT-5.5 (79.1% vs 74.9% uniform). On SSv2, MaP achieves 53.2% on Qwen3-VL-2B (vs 51.8% uniform) and 80.0% on GPT-5.5 (vs 71.1% uniform, a gain of 8.9%). Semantic keyframe selectors (AKS, FOCUS) substantially degrade CLEVRER performance because critical evidence lies in continuous transitions with subtle frame-level semantic changes. SoM and GoM's dense static annotations obscure task-relevant visual evidence. On TempCompass, MaP's motion-guided sampling matches or slightly improves performance (67.6% vs 67.3% on Qwen3-VL-2B; 88.6% vs 88.3% on GPT-5.5), showing no degradation on non-motion tasks.

Frame-budget analysis on CLEVRER at 0.5, 1, 2 FPS shows that trajectory-marking gains increase with sampling rate (+0.40%, +1.80%, +4.41% for Qwen3-VL-2B). This is because marking effectiveness depends on trajectory fidelity: at 0.5 FPS, inter-frame trajectories collapse into coarse strokes and overlays compete with raw evidence, while denser budgets allow more faithful representation of curvature, direction changes, and speed variations.

Cumulative ablation on CLEVRER shows that trajectory marking is the primary source of improvement (+1.3% for GPT-5.5, +1.8% for Qwen3-VL-2B), with motion-guided sampling adding further gains (+1.0% for both). Timestamps are an input convention; the two contributed components are trajectory marking and motion-guided sampling.

Preprocessing cost analysis shows MaP has the lowest or comparable latency: 783 ms per video on CLEVRER, 609 ms on SSv2, and 1,920 ms on TempCompass, substantially faster than SoM and GoM, and competitive with AKS and FOCUS. Average GPU memory is 0.35 GB (due to the compact 98 MB tracker and window-based streaming), with peak memory slightly higher than selection-only methods but below SoM and GoM.

In conclusion, MaP is a training-free, plug-and-play visual-prompting framework that mitigates inter-frame motion loss by recovering inter-frame motion from full-frame-rate video, computing motion energy scores to guide keyframe sampling, and explicitly marking trajectories between adjacent sampled frames onto keyframes. It consistently improves MLLM performance on motion-reasoning tasks without degrading broader video reasoning, at low preprocessing cost.

Improvements for AI systems

Improvements to AI systems:

  1. Add an inter-frame motion encoder module that takes dense point trajectories (e.g., from CoTracker3) and converts them into compact visual overlays (polylines with endpoints) directly on sampled frames. This allows any frozen MLLM to see motion that would otherwise be lost between sparse frames, without retraining or modifying the model’s architecture.

  2. Implement motion-energy-based adaptive frame sampling instead of uniform sampling. The system computes per-frame speed, acceleration, and curvature scores from camera-compensated trajectories, then selects frames via greedy non-maximum suppression with a minimum temporal gap. This ensures keyframes capture high-motion transitions (collisions, direction changes, interactions) while preserving global coverage.

  3. Add a camera-motion compensation module that fits a global similarity transformation (translation, rotation, scale) to tracked points, isolating object motion from camera ego-motion. This improves accuracy in dynamic camera scenarios (e.g., autonomous navigation, egocentric video) by preventing false motion signals.

  4. Enable trajectory-fidelity-aware budget allocation: The system dynamically adjusts the number of sampled frames and the density of trajectory markings based on available frame budget. At higher sampling rates, it marks more detailed trajectories (curvature, acceleration), while at low budgets it prioritizes only the most informative motion strokes to avoid visual clutter.

  5. Integrate a preprocessing pipeline that is training-free and plug-and-play: The system can be inserted before any existing MLLM (e.g., GPT-5.5, Qwen3-VL) as a lightweight front-end, with low latency (under 2 seconds per video) and minimal GPU memory (0.35 GB), making it suitable for real-time interactive applications like robotic manipulation or autonomous driving.

What the improved AI system can do:

  • Reason about object motion, collisions, and causal interactions in videos with high accuracy, even when only a few frames are sampled (e.g., 1–2 FPS), by explicitly visualizing inter-frame displacement, direction changes, and speed variations.

  • Maintain or improve performance on non-motion tasks (e.g., object existence, attribute recognition, temporal comprehension) without trade-offs, as shown by TempCompass results.

  • Operate in real-time or near-real-time on edge devices or robots, given its low preprocessing cost and no need for model fine-tuning.

  • Generalize across different MLLMs and video domains (simulated physics, human actions, egocentric video) without per-domain adaptation, because it only modifies the input representation.

  • Provide interpretable visual prompts—the marked trajectories serve as explicit evidence for the model’s reasoning, aiding debugging and human oversight in safety-critical applications.

Sources

Related papers