EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
summary
The gist
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA), but existing benchmarks predominantly evaluate answer correctness
In short
The EG-VQA benchmark requires video models to generate both answers and supporting temporal evidence for questions, addressing a gap where existing VideoQA systems lack evidence grounding. The proposed EG-Reasoner model uses explicit supervision to learn this joint task, showing that structured evidence supervision is crucial for reliable and interpretable video reasoning.
Key concepts
- EG-VQA
- An open-ended evaluation protocol that demands models produce both an answer and temporally localized evidence from a video. This forces models to show their reasoning by providing verifiable clips, moving beyond simple answer prediction.
- Evidence-Grounded F1 (EG-F1)
- The primary metric used to evaluate the benchmark. It measures how well the model's generated evidence aligns temporally and semantically with the ground-truth evidence using bipartite matching based on temporal overlap and semantic similarity.
- EG-Reasoner
- A proposed model trained with explicit supervision. It learns to generate answers and localize supporting temporal segments simultaneously, using a composite reward function that enforces correctness, format adherence, and evidence alignment during training.
Terminology used across episodes
This episode discusses
- EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence · Paper Radio
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Reinforcing Video Reasoning with Focused Thinking
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GPT-4o System Card
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- CausalVLR: A Toolbox and Benchmark for Visual-Linguistic Causal Reasoning
- CinePile: A Long Video Question Answering Dataset and Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- STAR: A Benchmark for Situated Reasoning in Real-World Videos
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
The paper
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence · Read on arXiv
Linpeng Huang, *Weixing Chen*, *Zexin Chen*, Yang Liu†, Liang Lin
Sun Yat-sen University · Peng Cheng Laboratory · Shenzhen University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence".
Tom: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA),
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re talking about a paper called EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence, and it’s authored by Linpeng Huang, Weixing Chen, Zexin Chen, Yang Liu, and Liang Lin. What does that title actually tell us about what this research is focused on?
Jane: It tells us immediately that they are focusing on a benchmark for Video Question Answering where the goal isn't just correctness; it’s about grounding those answers in verifiable temporal evidence. They are aiming to create a standard for judging how well models can connect their thoughts to specific moments in a video.
Lu: The authors clearly identified a gap where existing VideoQA benchmarks, like those mentioned earlier, were only looking at the final answer accuracy and ignoring the necessary step of linking that answer back to the actual visual input.
Meng: It seems they’ve formalized this idea by proposing a new way to evaluate these models based on this evidence-grounded reasoning approach. That structured evaluation framework is important because it gives us a clearer metric for what "good" video understanding actually looks like.
Lalam: This paper suggests that the future of reliable AI in video won't just be about generating fluent text; it’s going to be about generating fluent text backed by undeniable proof from the source material.
The paper's summary: Tom: So, diving into what they actually summarized in this EG-VQA paper, the core idea is that current models often have a big gap between getting an answer right and actually showing us which part of the video led to that conclusion. This is why they introduced EG-Reasoner, a model trained with explicit supervision to do both tasks simultaneously.
Jane: That’s because they found that even very strong proprietary models struggle when it comes to faithfully locating the supporting temporal segments for their predictions, which is a fundamental issue in current VideoQA systems.
Lu: The summary highlights that they are using this joint generation of answers and evidence localization to achieve state-of-the-art performance among open-source models, which is interesting because it suggests that training methods can directly fix the grounding problem rather than just patching the evaluation system later.
Meng: I see they also introduced a new metric called EG-F1, which is designed to measure both temporal alignment and semantic consistency at the same time using optimal bipartite matching, which sounds like a very rigorous way to assess quality compared to older methods.
Lalam: This entire approach points toward an AI culture where we are not just looking for outputs, but for outputs that carry verifiable evidence attached, which really raises the bar for what we expect from our multimodal systems.
The paper's improvements: Tom: Now let’s talk about the specific improvements they suggest to fix this problem. They propose shifting from answer prediction to evidence-grounded reasoning, where the model has to produce both an answer and supporting temporal evidence alongside it, which promotes what they call "interpretable, verifiable, and accountable reasoning."
Jane: The key improvement here is moving away from simple answer prediction toward a joint generation task that forces the model to create localized evidence. This means when an AI makes a mistake, we can point directly to the specific time segment it got wrong.
Lu: They are also using Group Relative Policy Optimization, or GRPO, which uses a composite reward function that specifically rewards adherence to output format, answer correctness, and evidence alignment using terms like Rformat and eRevi. That explicit supervision is what seems to drive their success in making the model learn how to ground things properly.
Meng: From an engineering standpoint, seeing them use such a detailed reward structure that includes evidence alignment suggests a very sophisticated way to guide the model during training, ensuring it’s not just guessing but actually learning the relationship between text and video frames.
Lalam: This methodology provides a clear blueprint for building next-generation AI where we can design the training process to enforce this level of structured reasoning from the very start, rather than hoping for good results later.
Conclusion: Tom: So, wrapping up this discussion on EG-VQA: they’ve shown that even without massive scaling, structured evidence supervision is essential for making video understanding reliable and interpretable. They also presented their EG-Reasoner model which shows state-of-the-art performance compared to other open-source models.
Jane: Essentially, the main message is that we need more than just bigger models; we need explicit ways to supervise the reasoning process so the AI knows how to connect its claims directly to visual data. This is a really important step for building trustworthy systems.
Lu: The implication here for future research is that we should focus heavily on developing these types of supervision techniques—the evidence-grounded ones—because scaling alone isn't enough to solve the grounding challenge in video tasks.
Meng: For practical application, this means we can start designing evaluation protocols that test for this level of grounded reasoning, which will help us select the right models for high-stakes scenarios where a verifiable explanation is necessary.
Lalam: The work on EG-VQA gives us a clear direction: focus on building systems where evidence isn't an afterthought but an integral, supervised component of the entire reasoning process.
Tom: That’s everything we have for today on EG-VQA, focusing on how requiring models to generate localized evidence fundamentally improves accountability in video understanding. We’ll take a quick break and get right back to you after this.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck