SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding".
Tom: Multimodal Large Language Models (MLLMs) have shown rapid progress in understanding single videos, but their ability to reason across multiple independent video streams remains poorly understood.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, we're looking at the paper titled "SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding," and I think the title tells us exactly what’s happening here: they're creating a controlled testing ground for when AI needs to look across multiple separate videos.
Jane: Exactly, and it’s not just about seeing if a model can watch two videos at once; it’s about the ability to reason across disjointed streams, which is way more complex.
Lu: The authors are from New York University, and they've structured their approach around isolating the core reasoning skills by simulating these different video interactions using three distinct simulator engines.
Meng: Isolating skills is smart; it helps you pinpoint exactly where the model is failing—is it a problem with time, space, or just pulling an object out of context?
Lalam: That isolation is key because when a model messes up in a real-world scenario, we often don't know if the issue lies in its core reasoning or just some noise in the training data.
The paper's summary: Tom: What they’re summarizing is how they moved away from relying on human annotation of real-world footage, which was creating precision issues with things like exact three dee distances and sub-second temporal offsets.
Jane: They created a dataset with eight thousand one hundred sixty-three question-answer pairs across nine thousand six hundred fifty unique videos where the ground truth is programmatically verified using simulation states.
Lu: The paper explains that they generate these distinct multi-video subsets by independently leveraging Habitat for navigation, Kubric for synchronization, and CLEVRER for sequential ordering to construct tasks with controlled variables like camera viewpoints and object identities.
Meng: So the core summary is that they’ve built a framework to test four specific diagnostic pillars: Temporal Alignment, Spatial Tracking, Comparative Reasoning, and Holistic Synthesis.
Lalam: That decomposition into those four pillars really makes the complex task of cross-video understanding much more manageable for researchers trying to diagnose where the model is struggling.
The paper's improvements: Tom: One big improvement they highlight is their methodology, which moves away from relying on human labeling, instead using a three-step generation pipeline involving simulator state extraction and synchronized rendering.
Jane: That programmatic grounding means the ground truth is precise, consistent across temporal and spatial variables, and available for every single task.
Lu: They specifically detail how they use Habitat to extract object IDs from frame-wise matrices for object re-identification, Kubric to infer temporal offsets between asynchronous videos by sampling crops, and CLEVRER to test sequential ordering of shuffled clips.
Meng: From an engineering standpoint, the improvement is that it turns the messy human labeling problem into a repeatable process that generates high-precision data at scale.
Lalam: And looking at their results, they show a substantial gap between current models and humans—the best model only achieved fifty-two point five percent average accuracy compared to an eighty-nine point five percent human baseline on zero-shot evaluation of leading MLLMs.
Conclusion: Tom: So, to wrap up, the SYNCR paper shows that we can isolate and test four key reasoning skills—Temporal Alignment, Spatial Tracking, Comparative Reasoning, and Holistic Synthesis—using a synthetic framework with programmatically verified data.
Jane: The main implication is that this controlled environment allows researchers to diagnose exactly where MLLMs fall short when it comes to integrating information across different video streams.
Lu: It suggests that simply scaling up models isn't enough; we need specialized training objectives and architectures specifically designed to handle the fine-grained physical and spatial reasoning challenges exposed by this benchmark.
Meng: I think for practical application, this framework points toward needing more sophisticated methods for cross-video grounding so that autonomous systems can reliably track objects when they move between different camera views.
Lalam: Ultimately, SYNCR provides a controlled testbed to guide the development of future MLLMs toward better holistic synthesis capabilities in real-world scenarios.
New York University
cs.CV
Submitted: 2026-05-08
Updated: 2026-09-30
Comments: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026) Workshop: BabyVLM: Toward Developmentally Plausible Multimodal Systems
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: Multimodal Large Language Models (MLLMs) have shown rapid progress in understanding single videos, but their ability to reason across multiple independent video streams remains poorly understood.
Key concepts
- Cross-Video Reasoning
- The ability of an AI system to align, compare, and integrate information from several separate video streams at the same time. This moves beyond understanding a single video and tests if a model can connect events happening in different videos to form a complete picture.
- Synthetic Grounding
- Creating highly precise ground truth for testing by generating data through simulation rather than relying on human annotation of real footage. This ensures that the answers provided by the benchmark are exact, consistent across time and space, and programmatically verifiable.
- Diagnostic Pillars
- Four specific areas tested in SYNCR: Temporal Alignment (sequencing events), Spatial Tracking (maintaining object location across views), Comparative Reasoning (analyzing differences between videos), and Holistic Synthesis (combining all observations into a global scene representation).
- Habitat, Kubric, CLEVRER
- Three distinct simulation engines used to generate varied video subsets. Habitat handles navigation and object tracking; Kubric focuses on synchronizing different camera viewpoints; and CLEVRER tests sequential ordering and kinematic comparisons by shuffling clips.
Terminology
Summary
Multimodal Large Language Models (MLLMs) have shown rapid progress in understanding single videos, but their ability to reason across multiple independent video streams remains poorly understood. Existing multi-video benchmarks rely heavily on human annotation of real-world footage, which limits the precision of ground truth and makes it difficult to diagnose model failures. This paper introduces SYNCR, a controlled synthetic benchmark designed for cross-video reasoning with programmatically verified grounding, addressing this gap by isolating core reasoning skills through simulation.
The Problem and Motivation
The rapid evolution of MLLMs has led to substantial progress in single-video understanding, but human perception and many real-world applications rarely involve a single isolated perspective. Intelligent systems must synthesize fragmented spatiotemporal information from multiple independent sources, motivating a shift toward cross-video reasoning: the ability to align, compare, and integrate information across disjointed video streams.
Current benchmarks like MVUEval and CVBench are valuable for broad semantic understanding but rely on human annotation which limits the precision of variables such as exact 3D distances, sub-second temporal offsets, camera-aligned object trajectories,
making it hard to attribute model errors.
The SYNCR Framework and Diagnostic Pillars
SYNCR is a controlled simulation-based framework built using three complementary engines: Habitat [21], Kubric [8], and CLEVRER [31]. It generates distinct multi-video subsets by independently leveraging
these engines to construct isolated tasks with controlled temporal offsets, camera viewpoints, object identities, physical trajectories, semantic instances, and topological routes.
This design decomposes cross-video understanding into four diagnostic pillars:
-
Temporal Alignment: For
synchronizing and sequencing disjointed events.
-
Spatial Tracking: For
maintaining object permanence and resolving cross-view geometry.
-
Comparative Reasoning: For
comparing kinematic or structural differences across videos.
-
Holistic Synthesis: For
aggregating fragmented observations into a global scene-level representation.
Dataset Generation with Synthetic Grounding
The dataset comprises 8,163 question-answer pairs grounded in 9,650 unique videos, ensuring that the ground truth is precise, consistent, and available across temporal, spatial, physical, and topological variables.
This grounding is achieved via a three-step generation pipeline: (1) simulator state extraction to gather precise, frame-aware variables
; (2) synchronized video rendering and sequence cropping; and (3) programmatic QA formulation. Specific task implementations include:
** Habitat:**
-
Object Re-identification: Ground truth is extracted
programmatically from frame-wise object ID matrices
to track exact timestamps of a target instance's appearance in the reference video. -
Object Counting: Ground-truth counts are derived by extracting
the exact number of unique semantic instance IDs for a queried object category across all generated trajectories,
with constraints like requiring the target instance to occupy at least 5% of the frame pixels. -
Route Planning: Ground truth is extracted from NavMesh connectivity logs, requiring a minimum path length of three nodes to infer
navigable connectivity.
** Kubric:**
-
Multi-Angle Synchronization: Requires inferring
temporal offsets between three asynchronous videos
by sampling crops and asking the model to determine starting timestamps relative to a reference video. -
Spatial Measurement: Requires identifying the object with the
minimum simulator-derived 3D Euclidean distance to the target object at the event frame.
** CLEVRER:**
-
Sequential Ordering: Involves reconstructing a continuous physical event from shuffled clips by splitting simulations into segments and asking for the
correct chronological order of these video segments.
-
Kinematic Comparison: Evaluates whether a model can compare continuous motion across videos by identifying
the object with the highest peak velocity.
Experimental Results and Conclusions
The zero-shot evaluation reveals a substantial gap between current models and humans,
with the best model achieving only 52.5% average accuracy compared to an 89.5% human baseline. Models perform relatively well on temporal ordering but struggle with precise physical and spatial reasoning,
particularly in Kinematic Comparison, where the best open-weight result is only 26.0%. While parameter scaling and reasoning-specialized post-training improve temporal alignment capabilities, they do not reliably address fine-grained physical tracking or global spatial synthesis.
Exploratory sim-to-real correlation analysis suggests that several SYNCR tasks track model trends on real-world benchmarks, while others expose reasoning capabilities underrepresented by existing evaluations,
positioning SYNCR as a controlled testbed for diagnosing cross-video reasoning failures.
Limitations
The framework has limitations, including the inherent constraints of the simulation engines regarding visual diversity and photorealism,
reliance on strict exact-match Question Answering (QA),
exclusion of audio streams, and a sample size bounded by recent architectural emergence.
Improvements for AI systems
Based on the SYNCR paper, here are specific improvements that can be made to existing Multimodal Large Language Models (MLLMs) and a description of what those improved systems could achieve:
-
Improve MLLM capabilities in cross-video reasoning by introducing a structured synthetic benchmark framework called SYNCR.
-
Develop MLLMs capable of performing four distinct diagnostic pillars:
-
Temporal Alignment (synchronizing and sequencing disjointed events).
-
Spatial Tracking (maintaining object permanence and resolving cross-view geometry).
-
Comparative Reasoning (comparing kinematic, structural, or numerical differences across videos).
-
Holistic Synthesis (aggregating fragmented observations into a global scene-level representation).
-
Implement reasoning-specialized post-training techniques (e.g., the
Thinking
checkpoint approach) specifically targeting temporal alignment to improve chronological ordering accuracy without sacrificing other reasoning skills like spatial tracking or kinematic comparison. -
Enhance model architectures and training objectives to resolve persistent bottlenecks in fine-grained physical and global spatial reasoning, particularly in Kinematic Comparison and Holistic Synthesis, which currently show near-chance performance across various scales.
-
Integrate simulation engines (Habitat for semantic navigation/topological synthesis, Kubric for multi-camera spatial/temporal reasoning, and CLEVRER for kinematic/collision reasoning) into a unified data generation pipeline to create high-precision ground truth grounded in programmatic physics and geometry, eliminating human annotation noise.
-
Develop robust cross-video grounding mechanisms that allow models to reliably track object identity across drastically different viewpoints (Spatial Tracking) and accurately estimate 3D relative distances between objects in dynamic scenes (Spatial Measurement).
-
The improved AI system could achieve the following:
-
Accurately synchronize events captured from multiple, unaligned surveillance feeds or autonomous vehicle cameras, allowing for precise event reconstruction across time offsets.
-
Maintain persistent object tracking and identity across continuous motion when an object moves from a single camera view to another (e.g., tracking a vehicle as it passes behind a building).
-
Perform complex comparative analysis of physical interactions between objects in different videos, such as identifying which object exhibits the highest peak velocity during an interaction or quantifying the difference in collision counts between scenes.
-
Synthesize fragmented visual evidence from multiple camera angles into a single, coherent global map or scene description, enabling tasks like autonomous navigation path planning based on inferred topological connectivity.
-
Achieve a performance level that approaches human expertise (89.5% baseline) on complex cross-video reasoning tasks by specifically addressing the current scaling plateaus and architectural limitations exposed by the SYNCR benchmark.
Abstract
Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorly understood. Existing multi-video benchmarks rely largely on human-annotated real-world footage, limiting the precision of spatial, temporal, and physical ground truth and making it difficult to diagnose model failures. We introduce SYNCR, a controlled synthetic benchmark for cross-video reasoning with programmatically verified grounding. Built using Habitat, Kubric, and CLEVRER simulator engines, SYNCR contains 4,000 multi-video question-answer pairs grounded in 4,827 unique videos. It evaluates MLLMs across eight tasks spanning four diagnostic pillars: Temporal Alignment, Spatial Tracking, Comparative Reasoning, and Holistic Synthesis. Our zero-shot evaluation of leading open- and closed-weight MLLMs reveals a substantial gap between current models and humans: the best model achieves only 64.5% average accuracy, compared to an 89.5% human baseline. Models perform relatively well on temporal ordering but struggle with precise physical and spatial reasoning, with the best model reaching only 29.8% accuracy on Kinematic Comparison. We further find that parameter scaling and reasoning-specialized post-training improve temporal alignment capabilities, but do not reliably address fine-grained physical tracking or global spatial synthesis. Finally, a sim-to-real correlation analysis suggests that SYNCR tracks model-level trends on a real-world multi-video benchmark.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- EMMA: Efficient Visual Alignment in Multi-Modal LLMs
- Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
- On the Binding Problem in Artificial Neural Networks
- World Models
- LLaVA-OneVision: Easy Visual Task Transfer
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- OpenAI GPT-5 System Card
- Gemini: A Family of Highly Capable Multimodal Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- MLVU: Benchmarking Multi-task Long Video Understanding
- CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models