TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

arXiv:2608.13495 · cs.CV, cs.LG · Submitted 2026-08-13 · Read on arXiv

Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu, Jin Yao, David I. Inouye, Jing Gao, Danhua Guo, Burhan Yaman

Uber AV Labs · Purdue University

cs.CV, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval This paper addresses the challenge of efficiently retrieving relevant clips from large-scale driving logs, which is

Terminology

Summary

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

This paper addresses the challenge of efficiently retrieving relevant clips from large-scale driving logs, which is essential for data curation, model development, and safety analysis. The authors note that structured and rule-based retrieval systems can explicitly target driving events but require expert-defined rules, auxiliary data, and multi-stage perception pipelines. Multimodal embedding models offer a simpler alternative by representing each video with a single searchable vector, but general-purpose models often rely on shortcuts from static scene context and struggle to distinguish motion-centric events such as turning left versus right or accelerating versus decelerating.

The paper proposes TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization (GRPO). The authors first fine-tune Qwen3-VL-Embedding on paired clips and reasoning traces from nuReasoning using an InfoNCE objective, which substantially improves overall retrieval but remains insufficient for fine-grained motion understanding. TraVEL then uses ego-trajectory similarity as privileged training supervision; trajectories serve only during training, and retrieval still operates on single-vector video embeddings without ego poses, expert rules, or auxiliary perception outputs.

The authors construct a driving-video retrieval benchmark from nuReasoning and report that TraVEL improves motion-centric retrieval across model scales: relative to supervised fine-tuning (SFT), it raises longitudinal and lateral mAP by 9.8 and 4.7 points at 2B, with corresponding gains of 7.2 and 1.5 points at 8B.

The paper's contributions are threefold: (1) presenting TraVEL, a simple and effective framework for adapting general-purpose multimodal embedding models to driving-video retrieval; (2) introducing a trajectory-derived reward and a GRPO-based fine-tuning stage that make video embeddings more sensitive to ego motion; and (3) introducing a driving-video retrieval benchmark derived from nuReasoning and demonstrating the potential of embedding-based retrieval for motion-centric driving queries.

The authors diagnose the pretrained embedding space and find a substantial domain gap: zero-shot retrieval varies widely across actions and is particularly weak for directional turns, lane changes, and subtle within-lane motion. They observe that caption supervision alone does not provide a sufficiently dense signal for learning fine-grained vehicle motion, motivating the trajectory-aware GRPO stage.

In the TraVEL framework, each clip is represented by an ego-pose sequence resampled to T steps, from which scalar motion channels are extracted, specifically ego-frame lateral displacement and speed. The signals are warped using a logarithmic function to emphasize changes near zero, and dynamic time warping (DTW) is used to compare motion signals with tolerance to temporal offsets. The trajectory reward is computed by normalizing channel distances and combining them with non-negative weights, yielding values in [0, R] with larger values for more similar ego motion.

During optimization, text-to-video groups contain the paired video and high-scoring hard negatives with binary instance rewards, while video-to-video groups use graded ego-trajectory similarity rewards. GRPO standardizes rewards within each group and updates the shared embedding model, preserving instance-level retrieval while organizing the video space by motion.

In experiments, all off-the-shelf models exhibit a substantial domain gap on driving-video retrieval, with R@1 values at most 2.4 despite ranging from 151M to 8B parameters. Cosmos-Embed1 is the strongest off-the-shelf baseline, reaching 2.4, 9.2, and 14.8 at R@1, R@5, and R@10, respectively. Fine-tuning with TraVEL substantially improves instance retrieval: at 2B, it increases R@1, R@5, and R@10 from 1.3, 4.5, and 7.4 to 8.7, 24.5, and 35.3, respectively, while reducing the median rank from 266 to 23. At 8B, it raises the corresponding recall values from 1.6, 5.2, and 7.8 to 10.1, 28.3, and 39.8, and reduces the median rank from 264 to 17.

For motion-centric retrieval, SFT at 2B improves longitudinal and lateral mAP from 24.7 and 17.7 to 45.9 and 31.0, respectively. TraVEL further raises them to 55.7 and 35.7, corresponding to gains of 9.8 and 4.7 points over SFT. At 8B, SFT reaches 50.3 longitudinal and 38.4 lateral mAP, while TraVEL improves them to 57.5 and 39.9, for additional gains of 7.2 and 1.5 points. The largest improvements appear in stopping and turning behaviors, while smaller gains on subtle within-lane shifts suggest these fine-grained lateral distinctions remain challenging.

A qualitative example shows that Pretrained fails to recover the requested longitudinal behavior and returns a nighttime clip in which the ego vehicle accelerates, whereas TraVEL ranks the paired clip first, capturing both the gentle slowdown and the pedestrian interaction described by the query.

The authors conclude that TraVEL preserves the instance-level gains of SFT at both model scales and that readily available physical motion signals can complement language supervision to produce representations that better capture fine-grained driving behavior. They note that the current evaluation uses the available subset of nuReasoning with a retrieval pool of 1,715 videos, and that evaluating on larger pools, additional datasets, and broader driving conditions will be important for establishing scalability and generalization. They also acknowledge that fine-grained lane changes and within-lane shifts remain challenging, suggesting that ego motion alone does not capture every distinction relevant to driving-video retrieval, and propose future work incorporating the motion of surrounding agents and scene structure into the reward while retaining the efficiency of embedding-based search.

Improvements for AI systems

Improvements to AI systems:

  1. Motion-aware video retrieval with trajectory-guided fine-tuning:
  • Integrate ego-trajectory similarity (e.g., lateral displacement and speed signals, warped via log function and compared with dynamic time warping) as a reward in Group Relative Policy Optimization (GRPO) to fine-tune multimodal embedding models.

  • The improved system can retrieve driving clips based on fine-grained motion semantics (e.g., gentle slowdown while turning left vs. hard acceleration while turning right) without requiring ego poses, expert rules, or auxiliary perception at inference time—only a single video embedding.

  1. Dense motion supervision beyond language captions:
  • Use trajectory-derived rewards to supplement caption-based InfoNCE fine-tuning, addressing the sparse signal problem in text-only supervision.

  • The improved system can distinguish motion-centric events that static scene context cannot (e.g., lane changes, stopping, within-lane shifts) and rank the correct clip first even when the query describes subtle longitudinal or lateral behaviors.

  1. Scalable retrieval with graded similarity rewards:
  • Apply graded ego-trajectory similarity (not just binary instance labels) to organize the video embedding space by motion, while preserving instance-level retrieval via text-to-video binary rewards.

  • The improved system can retrieve videos with similar motion patterns even if they are not the exact paired clip, enabling better data curation for rare driving behaviors and safety analysis.

  1. Benchmark-driven evaluation for driving-video retrieval:
  • Use the proposed nuReasoning-derived benchmark (1,715-video pool) to systematically measure motion-centric mAP (longitudinal and lateral) and recall metrics.

  • The improved system can be validated across model scales (2B and 8B) and shows consistent gains (e.g., +9.8 longitudinal mAP at 2B, +7.2 at 8B over SFT), making it reliable for deployment in autonomous driving data pipelines.

  1. Future extension to multi-agent motion and scene structure:
  • Extend the trajectory reward to include surrounding agents' motion and scene geometry (e.g., relative positions, velocities) to capture distinctions that ego motion alone misses (e.g., subtle within-lane shifts).

  • The improved system can retrieve clips based on interactive driving scenarios (e.g., yielding to a pedestrian while decelerating) and maintain efficiency of embedding-based search without multi-stage perception.

Abstract

Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events, but typically require expert-defined rules, auxiliary data, and multi-stage perception pipelines. Multimodal embedding models offer a simpler and more efficient alternative by representing each video with a single searchable vector. However, general-purpose models often rely on shortcuts from static scene context and struggle to distinguish motion-centric events, such as turning left versus right or accelerating versus decelerating. In this work, we study how to adapt a general-purpose multimodal embedding model to driving-video retrieval. We first fine-tune Qwen3-VL-Embedding on paired clips and reasoning traces from nuReasoning using an InfoNCE objective. While this stage substantially improves overall retrieval, caption supervision alone remains insufficient for fine-grained motion understanding. We therefore introduce TraVEL (Trajectory-Guided Video Embedding Learning), a motion-aware fine-tuning framework that uses ego-trajectory similarity as a reward within Group Relative Policy Optimization. Trajectories serve only as privileged training supervision; retrieval still operates on single-vector video embeddings without ego poses, expert rules, or auxiliary perception outputs. We further construct a driving-video retrieval benchmark from nuReasoning. Experiments show that TraVEL improves motion-centric retrieval across model scales: relative to SFT, it raises longitudinal and lateral mAP by 9.8 and 4.7 points at 2B, with corresponding gains of 7.2 and 1.5 points at 8B. TraVEL thus combines physically grounded supervision with efficient embedding-based search.

Sources

Related papers