EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence".
Tom: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA),
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re talking about a paper called EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence, and it’s authored by Linpeng Huang, Weixing Chen, Zexin Chen, Yang Liu, and Liang Lin. What does that title actually tell us about what this research is focused on?
Jane: It tells us immediately that they are focusing on a benchmark for Video Question Answering where the goal isn't just correctness; it’s about grounding those answers in verifiable temporal evidence. They are aiming to create a standard for judging how well models can connect their thoughts to specific moments in a video.
Lu: The authors clearly identified a gap where existing VideoQA benchmarks, like those mentioned earlier, were only looking at the final answer accuracy and ignoring the necessary step of linking that answer back to the actual visual input.
Meng: It seems they’ve formalized this idea by proposing a new way to evaluate these models based on this evidence-grounded reasoning approach. That structured evaluation framework is important because it gives us a clearer metric for what "good" video understanding actually looks like.
Lalam: This paper suggests that the future of reliable AI in video won't just be about generating fluent text; it’s going to be about generating fluent text backed by undeniable proof from the source material.
The paper's summary: Tom: So, diving into what they actually summarized in this EG-VQA paper, the core idea is that current models often have a big gap between getting an answer right and actually showing us which part of the video led to that conclusion. This is why they introduced EG-Reasoner, a model trained with explicit supervision to do both tasks simultaneously.
Jane: That’s because they found that even very strong proprietary models struggle when it comes to faithfully locating the supporting temporal segments for their predictions, which is a fundamental issue in current VideoQA systems.
Lu: The summary highlights that they are using this joint generation of answers and evidence localization to achieve state-of-the-art performance among open-source models, which is interesting because it suggests that training methods can directly fix the grounding problem rather than just patching the evaluation system later.
Meng: I see they also introduced a new metric called EG-F1, which is designed to measure both temporal alignment and semantic consistency at the same time using optimal bipartite matching, which sounds like a very rigorous way to assess quality compared to older methods.
Lalam: This entire approach points toward an AI culture where we are not just looking for outputs, but for outputs that carry verifiable evidence attached, which really raises the bar for what we expect from our multimodal systems.
The paper's improvements: Tom: Now let’s talk about the specific improvements they suggest to fix this problem. They propose shifting from answer prediction to evidence-grounded reasoning, where the model has to produce both an answer and supporting temporal evidence alongside it, which promotes what they call "interpretable, verifiable, and accountable reasoning."
Jane: The key improvement here is moving away from simple answer prediction toward a joint generation task that forces the model to create localized evidence. This means when an AI makes a mistake, we can point directly to the specific time segment it got wrong.
Lu: They are also using Group Relative Policy Optimization, or GRPO, which uses a composite reward function that specifically rewards adherence to output format, answer correctness, and evidence alignment using terms like Rformat and eRevi. That explicit supervision is what seems to drive their success in making the model learn how to ground things properly.
Meng: From an engineering standpoint, seeing them use such a detailed reward structure that includes evidence alignment suggests a very sophisticated way to guide the model during training, ensuring it’s not just guessing but actually learning the relationship between text and video frames.
Lalam: This methodology provides a clear blueprint for building next-generation AI where we can design the training process to enforce this level of structured reasoning from the very start, rather than hoping for good results later.
Conclusion: Tom: So, wrapping up this discussion on EG-VQA: they’ve shown that even without massive scaling, structured evidence supervision is essential for making video understanding reliable and interpretable. They also presented their EG-Reasoner model which shows state-of-the-art performance compared to other open-source models.
Jane: Essentially, the main message is that we need more than just bigger models; we need explicit ways to supervise the reasoning process so the AI knows how to connect its claims directly to visual data. This is a really important step for building trustworthy systems.
Lu: The implication here for future research is that we should focus heavily on developing these types of supervision techniques—the evidence-grounded ones—because scaling alone isn't enough to solve the grounding challenge in video tasks.
Meng: For practical application, this means we can start designing evaluation protocols that test for this level of grounded reasoning, which will help us select the right models for high-stakes scenarios where a verifiable explanation is necessary.
Lalam: The work on EG-VQA gives us a clear direction: focus on building systems where evidence isn't an afterthought but an integral, supervised component of the entire reasoning process.
Tom: That’s everything we have for today on EG-VQA, focusing on how requiring models to generate localized evidence fundamentally improves accountability in video understanding. We’ll take a quick break and get right back to you after this.
Linpeng Huang, *Weixing Chen*, *Zexin Chen*, Yang Liu†, Liang Lin
Sun Yat-sen University · Peng Cheng Laboratory · Shenzhen University
cs.CV, cs.AI
Submitted: 2026-06-23
Updated: 2026-10-04
Project page: https://hcplab-sysu.github.io/EG-VQA
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA), but existing benchmarks predominantly evaluate answer correctness
Key concepts
- EG-VQA
- An open-ended evaluation protocol that demands models produce both an answer and temporally localized evidence from a video. This forces models to show their reasoning by providing verifiable clips, moving beyond simple answer prediction.
- Evidence-Grounded F1 (EG-F1)
- The primary metric used to evaluate the benchmark. It measures how well the model's generated evidence aligns temporally and semantically with the ground-truth evidence using bipartite matching based on temporal overlap and semantic similarity.
- EG-Reasoner
- A proposed model trained with explicit supervision. It learns to generate answers and localize supporting temporal segments simultaneously, using a composite reward function that enforces correctness, format adherence, and evidence alignment during training.
Terminology
Summary
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA), but existing benchmarks predominantly evaluate answer correctness without grounding predictions in relevant video evidence. This disconnect between answer generation and evidence understanding motivates the construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), an open-ended evaluation protocol that requires models to generate temporally localized evidence alongside answers, enabling verifiable and interpretable reasoning.
How it works
The core of the proposed framework involves shifting from answer prediction to evidence-grounded reasoning by requiring models to produce both an answer and supporting temporal evidence. This paradigm shift is illustrated in Figure 1(b), where the model must generate temporally localized evidence alongside its prediction, promoting interpretable, verifiable, and accountable reasoning.
The resulting dataset comprises 2,067 videos and 11,838 QA pairs annotated with fine-grained temporal evidence.
The Benchmark Construction
The EG-VQA dataset is constructed from three existing video sources: ActivityNet Captions [Caba Heilbron et al., 2015], HiREST [Zala et al., 2023], and YouCook2 [Zhou et al., 2018]. These raw annotations undergo a rigorous multi-stage filtering protocol to ensure reliability. The filtering process includes:
-
Removing corrupted, incomplete, or inaccessible video samples.
-
Refining temporal boundaries by having human annotators correct misaligned segment boundaries and descriptions.
-
Discarding videos with low event diversity that would yield trivial questions, ensuring the dataset supports
non-trivial reasoning.
Reasoning Categories and Evaluation Metric
The benchmark targets four levels of reasoning: descriptive, temporal, causal, and counterfactual. The evaluation protocol is unified by the Evidence-Grounded F1 (EG-F1) metric. This metric assesses both temporal alignment and semantic consistency against ground-truth evidence through optimal bipartite matching. Formally, it computes precision and recall based on the number of matched pairs derived from pairwise temporal overlap (IoU) and semantic similarity. The soft variant, EG-F1soft, is introduced to mitigate reward sparsity during training by allowing all predicted-ground-truth pairs to contribute proportionally to their similarity.
The Proposed Model: EG-Reasoner
To bridge the discrepancy between high answer accuracy and poor evidence grounding observed in existing models, EG-Reasoner is proposed. This model is trained with explicit evidence supervision, jointly learning to generate answers and localize supporting temporal segments. The training framework utilizes Group Relative Policy Optimization (GRPO), which compares multiple candidate responses sampled from the same input to estimate relative quality using a composite reward function: R(y) = λfRformat(y) + λeRevi(y) + λaRans(y). This reward design enforces adherence to output schema, answer correctness, and evidence alignment.
Experimental Findings
Experimental evaluation reveals that even strong proprietary models struggle to accurately ground their predictions, exposing a fundamental discrepancy between answer correctness and faithful evidence localization.
EG-Reasoner achieves state-of-the-art performance among open-source models, reaching 26.88% strict accuracy and 42.71% relaxed accuracy on the overall task. Notably, EG-Reasoner surpasses average human performance on several evidence localization metrics, suggesting that explicit evidence supervision is essential for reliable and interpretable video understanding.
The study concludes that scaling alone is insufficient for robust video understanding,
emphasizing the necessity of structured evidence supervision to move beyond superficial correlations.
Conclusion
The paper introduces EG-VQA, a benchmark requiring models to ground answers in explicit temporal evidence, and proposes EG-Reasoner, an evidence-grounded reasoning model trained with explicit supervision. The findings demonstrate that incorporating structured evidence supervision is essential for reliable and interpretable video understanding, positioning EG-VQA as a useful testbed for future research on evidence-aware multimodal reasoning.
The gist
The construction of the Evidence-Grounded Video Question Answering Benchmark (EG-VQA), which requires models to generate temporally localized evidence alongside answers, reveals a fundamental discrepancy between answer correctness and reasoning faithfulness in existing VideoQA systems, necessitating the proposal of EG-Reasoner trained with explicit evidence supervision to achieve state-of-the-art performance.
Improvements for AI systems
Based on the provided scientific paper, EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence,
here are specific, actionable improvements for AI systems and what those improved systems will be capable of:
) 1. Move from Answer-Only Prediction to Explicit Evidence Localization (The EG-VQA Paradigm Shift)
Current Video QA models often produce correct answers without verifiable reasoning. The paper advocates for a paradigm shift where the model must generate temporally localized evidence alongside the answer.
-
An improved system will be able to distinguish between
plausible
answers andverifiable
answers by forcing joint generation of reasoning and evidence localization. -
This capability allows for auditing: if a model makes an error, researchers can immediately pinpoint which specific 30-second clip was misinterpreted or ignored.
) 2. Implement Structured Evidence Supervision via Reinforcement Learning (EG-Reasoner Training)
The system is improved by training models like EG-Reasoner using a composite reward function that jointly evaluates format compliance, answer correctness, and evidence alignment (EG-F1).
-
An improved system will exhibit significantly higher performance on complex reasoning tasks like counterfactual questions because it is explicitly rewarded for generating evidence that aligns semantically and temporally with the answer.
-
This moves beyond superficial correlations; the model learns to generate dense, high-quality grounding signals rather than just predicting a label.
) 3. Develop a Unified, Multi-Criteria Evaluation Metric (EG-F1)
The introduction of the Evidence-Grounded F1 (EG-F1) metric unifies temporal alignment and semantic consistency using optimal bipartite matching (Hungarian algorithm).
-
An improved system will be evaluated on a single, faithful metric that simultaneously ensures the answer is correct AND the supporting evidence is temporally and semantically accurate.
-
This provides a rigorous, interpretable measure of reasoning quality that existing metrics lack.
) 4. Enhance Reasoning Capabilities in Complex Tasks (Causal and Counterfactual QA)
The training objective (GRPO) specifically targets causal and counterfactual questions, showing pronounced gains in these areas compared to simple descriptive tasks.
-
An improved system will excel at multi-step reasoning, such as
How is X related to Y?
orWhat if Z had not happened?
by synthesizing information across non-adjacent temporal segments. -
This capability allows the AI to perform deeper analysis of dynamic video events rather than just surface-level observation.
) 5. Improve Temporal Precision Across Scales (Fine-Grained Grounding)
The evaluation protocol uses multiple IoU thresholds (0.1, 0.3, 0.5, 0.7), and the EG-Reasoner consistently outperforms baselines across this range, especially at stricter thresholds.
-
An improved system will achieve superior fine-grained temporal localization—it won't just find a segment near the event; it will precisely identify the start and end times of the relevant action within a long video context.
-
This is crucial for procedural tasks where identifying intermediate steps is as important as identifying the final result.
) 6. Increase Robustness Against Scaling Limitations (Supervision Over Scale)
The paper demonstrates that scaling model size alone is insufficient; explicit evidence supervision (the reward mechanism) is necessary to resolve grounding failures, even with large proprietary models like GPT-4o and Gemini-2.5-Flash.
-
An improved system will be less prone to
hallucinated evidence
or producing high answer accuracy based on weak reasoning, regardless of its underlying parameter count. -
This makes the AI more reliable for high-stakes applications where verifiability is paramount.
) 7. Enable Explainable Reasoning Trajectories (The Block)
The structured generation output forces the model to produce an intermediate reasoning trajectory
(block).
-
An improved system will provide a transparent, step-by-step explanation of its decision process, showing exactly how it linked specific evidence segments to form a final answer.
-
This capability makes the AI's logic interpretable and debuggable by human researchers.
) Summary of Improved AI System Capabilities:
The improved system will be a state-of-the-art Video QA engine that is not only highly accurate but also inherently trustworthy, verifiable, and explainable. It will be capable of performing complex reasoning (causal/counterfactual), precisely pinpointing the exact temporal segments supporting its claims, and providing a transparent justification for every inference made.
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Reinforcing Video Reasoning with Focused Thinking
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GPT-4o System Card
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- CausalVLR: A Toolbox and Benchmark for Visual-Linguistic Causal Reasoning
- CinePile: A Long Video Question Answering Dataset and Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- STAR: A Benchmark for Situated Reasoning in Real-World Videos
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models