Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo
The Hong Kong University of Science and Technology · Tencent
cs.CV, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
Project page: https://ddz16.github.io/cammotion.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
Terminology
Summary
Summary
This paper introduces a new formulation for camera-motion understanding called temporally grounded, compositional recognition,
which requires a model to localize motion-consistent intervals within a shot and identify every movement active within each interval. The authors argue that existing work typically assigns one or more labels to an entire clip, which overlooks that camera motion can change within a shot and that multiple movements can occur simultaneously.
To support this task, the authors introduce CAMCHOREO, a benchmark of 4,229 real single-shot YouTube clips with expert-annotated temporal segments. The annotations use a compact vocabulary of 20 direction-aware labels derived from 12 movement types, grouped into five families: rotation, translation, optical, subject-referenced, and stability. The benchmark contains 8,591 expert-annotated segments and 14,258 motion instances, with boundaries at 0.1-second resolution. Notably, 2,411 clips contain multiple segments, and 3,797 segments (44.2%) contain compound camera motion with multiple movement primitives active simultaneously. The clips span nine content domains, with aerial, documentary, and film footage contributing the largest shares.
The paper demonstrates that current multimodal large language models (MLLMs) perform poorly on this task. The strongest closed-source MLLM, Gemini-3.1-Pro, reaches only 43.3 frame-level micro F1. Scaling Qwen3-VL from 4B to 235B parameters increases micro F1 from 24.2 to only 33.4. The authors diagnose this as a representational gap: MLLM vision encoders are optimized for semantic alignment, while camera motion depends on cross-frame geometry including parallax, perspective change, and horizon rotation.
The authors propose CAMDISTILL, which distills geometric knowledge from a frozen 3D foundation model (VGGT or VGGT-Ω) into lightweight camera tokens during training, then removes the 3D model at inference. A lightweight Geometry-aware Camera Token Extractor (GCTE) predicts one camera token per frame from intermediate frozen vision features. GCTE uses alternating attention blocks: frame-wise cross-attention extracts camera-relevant evidence from each frame, while global camera self-attention compares evidence across time. A distillation loss aligns the student tokens with the teacher's camera tokens using cosine distance, combined with the standard next-token SFT loss.
The authors compare CAMDISTILL against a direct-injection baseline called CAMINJECT, which runs the 3D model at inference and concatenates teacher camera tokens with visual tokens. Results show CAMDISTILL matches the accuracy of direct feature injection without running the 3D teacher at inference. With the 4B backbone, both reach 67.5 micro F1; with the 8B backbone, CAMDISTILL achieves 67.8 versus CAMINJECT's 68.3. CAMDISTILL adds only 0.1 seconds latency and 1.8 GB memory over the base model, compared to CAMINJECT's 5.9 seconds and 4.8 GB overhead.
Ablation studies show that early-to-middle vision features perform best for the GCTE, four alternating blocks achieve the optimal trade-off, placing camera tokens before visual tokens improves micro F1 by 2.8 points, and VGGT-Ω as teacher improves micro F1 by 1.7 points over VGGT. The distillation weight peaks at λcam = 0.05. The method also transfers to external benchmarks: CAMDISTILL-8B improves over Qwen3-VL-8B by 20.7 mAP on CameraBench and 16.6 accuracy points on CMVQA.
The paper's contributions are threefold: (1) formulating temporally grounded, compositional camera-motion recognition and introducing CAMCHOREO as the first real-video benchmark combining variable-length segments with direction-aware multi-label annotations; (2) empirically showing that within-shot transitions and simultaneous movements are common, and that generic MLLMs and geometry-only pose rules remain inadequate; (3) proposing CAMDISTILL, which matches direct feature injection while removing the teacher and its cost at inference.
Improvements for AI systems
Improvements to AI Systems:
-
Temporally Grounded, Compositional Video Understanding – AI systems can now parse a single video shot into multiple, overlapping temporal segments, each labeled with a set of simultaneous camera movements (e.g.,
pan right + tilt up
during one interval, thenzoom in + roll
during the next). This enables fine-grained video editing, automated cinematography analysis, and precise shot-level metadata generation for search and retrieval. -
Geometry-Aware Vision Encoding for MLLMs – By distilling 3D geometric knowledge (parallax, perspective change, horizon rotation) into lightweight per-frame camera tokens, AI systems gain a representational layer that captures cross-frame motion geometry without needing a heavy 3D model at inference. This improves performance on camera-motion tasks from 33% to 68% micro F1 on the CAMCHOREO benchmark, while adding only 0.1s latency and 1.8GB memory.
-
Inference-Efficient Distillation of 3D Teachers – AI systems can now achieve the accuracy of running a full 3D foundation model (VGGT) at inference time, but without the 5.9s latency and 4.8GB memory overhead. The distilled camera tokens are generated from frozen intermediate vision features, making the system deployable on edge devices or real-time pipelines.
-
Compositional Multi-Label Temporal Segmentation – AI systems can now handle compound motions (44.2% of segments in the benchmark) and variable-length intervals, enabling richer video understanding than single-label-per-clip approaches. This supports applications like autonomous drone navigation (detecting simultaneous yaw + forward motion), sports broadcasting (analyzing camera work), and film studies (automated shot composition analysis).
-
Transferable Camera-Motion Reasoning – The distilled representation improves performance on external benchmarks (CameraBench +20.7 mAP, CMVQA +16.6 accuracy points), meaning AI systems can generalize camera-motion understanding across different video domains and tasks, such as visual question answering about camera movement or camera-based action recognition.
-
Early-to-Middle Vision Feature Utilization – AI systems can now leverage intermediate vision features (rather than final semantic embeddings) for geometric tasks, suggesting a general architectural principle: for motion/geometry understanding, tapping into earlier layers preserves spatial and temporal cues that are lost in deep semantic alignment. This can be applied to other spatio-temporal tasks like optical flow estimation or depth-from-video.
Sources
- Layer Normalization
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Geometry-Guided Camera Motion Understanding in VideoLLMs
- Distilling the Knowledge in a Neural Network
- ViPE: Video Pose Engine for 3D Geometric Perception
- LLaVA-OneVision: Easy Visual Task Transfer
- DROID-SLAM in the Wild
- Can video generation replace cinematographers? Research on the cinematic language of generated video
- OpenAI GPT-5 System Card
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
- VGGT-$\Omega$
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
- Cambrian-P: Pose-Grounded Video Understanding
- On the Generalization Capacities of MLLMs for Spatial Intelligence
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models