A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding".
Jane: This paper introduces MVX-Bench, a new benchmark designed to rigorously evaluate Multimodal Large Language Models (MLLMs) on complex multi-video understanding tasks that go beyond simple event-level comparison.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: What's interesting about the title "A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding" is that it clearly states they are giving both a method and a test for this new way of thinking about video analysis.
Jane: They’re trying to solve the problem where current models struggle with things like matching people across videos or figuring out what happened sequentially across several clips simultaneously, which is much harder than just describing one scene.
Lu: The authors are bringing together a benchmark, MVXBench, and a framework called SAMA to systematically evaluate these capabilities using eleven classical computer vision tasks adapted for this multi-video setting.
Meng: So they aren't just testing if the model can watch two videos at once; they are testing different kinds of complex understanding, from low-level perception up to high-level reasoning across those streams.
Lalam: That structure sounds very methodical; having a defined benchmark like MVXBench gives developers a clear roadmap for how to build better AI that can handle these complex cross-video scenarios reliably.
The paper's summary: Tom: Basically, the paper summarizes the existing issue: current methods usually just smash videos together, but that causes problems because of training mismatches and information loss from compression.
Jane: They propose SAMA as a solution to this structural problem by giving the language model an external toolkit of visual tools and guiding it with specific skills.
Lu: SAMA is organized into an Agent Core, which has a text-only LLM Planner, and an external Verification Layer that checks the results from those tools.
Meng: The core idea here is that instead of letting the LLM do everything blindly, you give it specific skills and tools so it knows exactly when to ask for evidence and how to check if that evidence makes sense.
Lalam: It sounds like they are essentially giving the model an iterative process—it plans a step, uses a tool, checks the result against another source, and keeps going until it's certain about its findings.
The paper's improvements: Tom: The authors highlight that they address several shortcomings in previous work by focusing on tasks that go beyond just simple event-level comparison, like identity-level matching and fine-grained discrimination.
Jane: They suggest using the SAMA framework to overcome the lack of explicit coordination in current agentic approaches, which often leave the LLM to figure out its own tool use strategies implicitly.
Lu: A key improvement they focus on is making sure that when tools give conflicting information, there's a conflict-aware verification module that forces an adaptive reading loop until those inconsistencies are resolved.
Meng: That conflict detection mechanism is what I find interesting from a practical standpoint; it means the system doesn't just output an answer and stop, but it actually verifies its own work before concluding anything.
Lalam: By integrating task-specific skills, like prioritizing visual similarity tools for appearance questions, they are making the reasoning process much more structured and less random when tackling complex multi-video queries.
Conclusion: Tom: So to wrap up, this paper introduces MVXBench to test MLLMs across diverse video understanding dimensions and SAMA as a framework that uses tools and conflict detection for robust, iterative reasoning.
Jane: The implication is that we are moving toward AI systems capable of performing much deeper cross-video analysis, something that goes far beyond simple scene description.
Lu: The impact could be seeing MLLMs become genuinely useful for tasks requiring complex temporal and spatial coordination across multiple video streams in real-world scenarios.
Meng: From an engineering view, this provides a solid blueprint for building more reliable systems because it shows exactly how to handle the integration of diverse visual tools in a structured way.
Lalam: I feel really optimistic about this because if we can build models that can reliably perform identity matching and forensic discrimination using SAMA, the applications for security and complex visual verification open up significantly.
Yue Zhang, Liqiang Jing, Jia Li Yapeng Tian, Xinya Du Yunhui Guo, Vibhav Giridhar Gogate
University of Texas at Dallas
cs.CV
Submitted: 2026-03-16
Updated: 2026-09-29
Comments: EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: This paper introduces MVX-Bench, a new benchmark designed to rigorously evaluate Multimodal Large Language Models (MLLMs) on complex multi-video understanding tasks that go beyond simple event-level
Key concepts
- MVXBench
- A new benchmark designed to rigorously evaluate Multimodal Large Language Models (MLLMs) on complex multi-video understanding tasks. It tests capabilities beyond simple event comparison, focusing on difficult tasks like matching people across videos or understanding sequential events across multiple clips.
- SAMA Framework
- A solution proposed to address structural problems in video analysis. It gives the language model an external toolkit of visual tools and guides it with specific skills. It includes an Agent Core with a text-only LLM Planner and an external Verification Layer to check tool results.
- Agent Core
- The component of the SAMA framework that contains a text-only Large Language Model Planner. This planner is responsible for planning the steps in the reasoning process, deciding which tools to use, and guiding the agent's actions based on its instructions.
- Conflict-aware verification module
- A key improvement in SAMA that checks results from different tools. If tools provide conflicting information, this module forces an adaptive reading loop until those inconsistencies are resolved. This ensures the system verifies its own findings before concluding.
Terminology
Summary
This paper introduces MVX-Bench, a new benchmark designed to rigorously evaluate Multimodal Large Language Models (MLLMs) on complex multi-video understanding tasks that go beyond simple event-level comparison. It addresses the limitations of existing methods by proposing SAMA, a Skill-Augmented Agentic Framework, which enables iterative and structured reasoning across multiple video streams through the integration of visual tools, task-specific skills, and a conflict-aware verification mechanism.
MVXBench Benchmark
The authors introduce MVXBench, a MultiVideo CrossDimension Benchmark,
which reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework. This benchmark spans lowlevel perception, mid-level cross-video matching, and high-level reasoning.
The benchmark contains 1,442 questions over 4,255 videos from diverse real-world datasets,
covering indoor household scenes as well as outdoor environments. Key features include introducing tasks such as identity-level matching,
finegrained discrimination,
and structurally complex reasoning scenarios such as counterfactual inference and defeasible entailment.
SAMA Framework Overview
SAMA is proposed as a framework to overcome the structural challenges of direct multi-video concatenation, which suffer from a mismatch between single-video training and multi-video inference
and information loss caused by aggressive frame compression.
SAMA is organized into two main layers: an Agent Core and an external Verification Layer. The Agent Core consists of:
-
A text-only LLM Planner that performs task analysis and decision-making.
-
A multimodal toolkit of Tools for extracting evidence from videos (e.g., Video Reader, Scene Graph, Object Tracker).
-
Reusable Skills that provide
structured workflows and domain guidance to support tool use.
Conflict Detection and Adaptive Reasoning
The Verification Layer is crucial for enabling iterative and structured reasoning.
It includes a Conflict Detection module that identifies inconsistencies among tool outputs (e.g., when the Video Reader describes a traffic light as red while the Scene Graph detects green). When conflicts are detected, the system triggers an Adaptive Reading module
which guides the Planner to invoke additional tools for verification until the conflict is resolved. This process ensures that verified evidence is aggregated before further reasoning,
maintaining accuracy through a loop that continues until sufficient, validated evidence is collected.
Task-Specific Skills and Tool Orchestration
To guide the LLM planner in complex coordination, SAMA incorporates task-specific skills such as the Multi-video Comparison
skill. This skill enforces evidence-based decision-making
by instructing the planner to prioritize tools based on the question type—for instance, prioritizing Visual Similarity
for appearance questions or relying on Video Reader
for event alignment in similarity tasks. The framework also includes skills like the Temporal Grounding Skill,
which defines a protocol to propose candidate timestamps and mandates verification with the Video Reader before using them in reasoning.
Experimental Results
Experimental results on MVXBench demonstrate that SAMA "outperforms all evaluated baselines, including open-source model families such as InternVL3.5 (Wang et al., 2025b), Qwen2.5-VL (Bai et al., 2025), and Gemma-3 (Kamath et al., 2025), as well as proprietary models such as GPT. The study shows that the agentic framework substantially improves a vision model's performance, with SAMA achieving
52.7% overall accuracy," surpassing GPT-4o (47.7%) and other open-source end-to-end models. Ablation studies confirm the effectiveness of these components: removing skills reduces accuracy by 8.3 percentage points, and removing conflict detection leads to a performance decline, proving that resolving cross-tool inconsistencies is critical for reliable multi-video reasoning.
Contributions Summary
The main contributions are threefold: 1) Constructing MVXBench, a benchmark with 11 canonical computer vision tasks adapted to multivideo scenarios. 2) Presenting SAMA, the first agentic framework explicitly tailored to multi-video scenarios,
which integrates diverse visual tools and a conflict resolution mechanism for structured reasoning. 3) Demonstrating that SAMA effectively outperforms strong open-source baselines and GPT on MVXBench, with ablation studies validating the importance of skill design and conflict detection.
Limitations
The authors note a limitation: Our benchmark focuses on multiple-choice question-answering settings, which may not fully reflect open-ended generation scenarios.
Exploring free-form multi-video reasoning is identified as a direction for future work. The paper concludes by detailing the agent's execution process, including strict rules for thinking, tool calling, and iterative refinement to ensure accuracy.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the provided scientific paper (SAMA framework and MVX-Bench), along with what these improved systems can achieve:
The primary improvement lies in moving from single-video, direct inference models to sophisticated, iterative, and cross-modal agentic reasoning systems.
Here are the specific improvements:
-
Implementation of Skill-Augmented Agentic Frameworks (SAMA):
-
Integration of Diverse Visual Tools: Utilizing a rich toolkit including Perception (Scene Graph, Object Tracker), Detection (Scene Graph), and Other tools (Visual Similarity, Subtitle Retriever/Extractor).
-
Structured Procedural Guidance via Task-Specific Skills: Embedding explicit
how-to
logic into the LLM planner to dictate the correct sequence and invocation of tools based on the question type (e.g., prioritizing Visual Similarity for style questions vs. using Video Reader for event alignment). -
Conflict-Aware Verification Mechanism: Introducing a dedicated layer that automatically detects inconsistencies between heterogeneous tool outputs (e.g., Scene Graph counting vs. Video Reader description) and triggers an
Adaptive Re-reading
loop to iteratively refine evidence until contradictions are resolved. -
Structured Memory for Evidence Aggregation: Maintaining a memory dictionary to store validated evidence per video, using concise distilled representations for long sequences, ensuring that crucial information is not lost as the context window grows.
-
Multi-Step Reasoning and Iterative Planning (Agent Core): Employing a text-only LLM Planner that performs task analysis, selects appropriate tools/skills, reasons over intermediate evidence, and initiates new reasoning cycles when current evidence is insufficient—essentially simulating structured human problem-solving.
The resulting improved AI system (SAMA) can perform the following specific capabilities:
-
Identity-Level Matching: Accurately determine if two individuals appearing in different videos are the same person, even when clothing, viewpoint, and background vary (e.g., Re-Identification task).
-
Fine-Grained Discrimination (Forensic Detection): Distinguish between authentic and manipulated video content by cross-referencing visual artifacts with other tools (e.g., detecting inconsistencies between object counts reported by the Scene Graph and the Video Reader).
-
Complex Structural Reasoning: Execute high-level reasoning tasks such as counterfactual inference (predicting outcomes under hypothetical changes, like
What if they had held on to the handlebar?
) and defeasible entailment (dynamically revising a hypothesis based on new visual evidence). -
Quantitative Cross-Video Comparison: Perform precise numerical comparisons across videos, such as accurately counting specific objects and determining relationships between those counts (e.g.,
Which video has more labeled objects than the reference?
). -
Style and Aesthetic Analysis: Identify and compare abstract visual properties like art style, color palettes, line quality, or animation fluidity across different video sources without relying solely on content matching (e.g.,
Which candidate clip comes from the same show based solely on visual style?
). -
Robustness to Information Loss: Maintain high accuracy even when dealing with compressed or multi-modal inputs by using structured tools and verification loops instead of simple concatenation, mitigating the training–inference mismatch inherent in direct multi-video processing.
Abstract
Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direct inference, which introduces training-inference mismatch, information loss from frame compression, and a lack of explicit cross-video coordination. Meanwhile, current multi-video benchmarks primarily emphasize event-level comparison, leaving identity-level matching, fine-grained discrimination, and structured multi-step reasoning underexplored. To address these gaps, we introduce MVX-Bench, a Multi-Video Cross-Dimension Benchmark that reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework, comprising 1,442 questions over 4,255 videos from diverse real-world datasets. We further propose SAMA, a Skill-Augmented Agentic Framework for Multi-Video Understanding, which integrates visual tools, task-specific skills, and a conflict-aware verification mechanism to enable iterative and structured reasoning. Experimental results show that SAMA outperforms strong open-source baselines and GPT on MVX-Bench, and ablations validate the effectiveness of skill design and conflict resolution.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Pixtral 12B
- Open-World Object Counting in Videos
- Qwen2.5-VL Technical Report
- A Short Note on the Kinetics-700 Human Action Dataset
- CI-VID: A Coherent Interleaved Text-Video Dataset
- Gemma 3 Technical Report
- Building and better understanding vision-language models: insights and future directions
- LLaVA-OneVision: Easy Visual Task Transfer
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- Visual Instruction Tuning
- Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos
- VCA: Video Curious Agent for Long Video Understanding
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models