Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

arXiv:2605.06185 · cs.AI, cs.CV · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios".

Jane: The paper was written by Yu Zhao, Peizheng Yan, Liang Xie, Mingming Wang, Juntong Qi et al. from Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen) and Tianjin Key Lab of Intelligent Unmanned Swarm Tech & System, Tianjin University and Tianjin Artificial Intelligence Innovation Center and Defense Innovation Institute and Academy of Military Sciences and School of Future Technology, Shanghai University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, Jane, can you break down the summary of what this paper actually proposes? We've heard the problem, but what's the solution in simple terms?

Jane: The authors propose replacing fixed-length clip memory with something that captures meaning and causality. Instead of just chopping up a video every five or ten seconds—which is often arbitrary—they are segmenting videos into semantically coherent events.

Meng: That’s the biggest shift, moving from a temporal window to an event-based unit of thought. The paper calls these units State-Event-State, or SES triplets.

Tom: And these aren't just floating clips; they are structured data points that capture the state before an action and after it. It’s like giving the AI a storyboard with causal arrows between each panel.

Lu: Exactly, and instead of dumping all those into one massive attention window, they build a global Event Knowledge Graph or EKG out of these triplets. This structure is stored in a dual-store memory system—a Vector Database for semantic matching and a Graph Database for topological mapping.

Jane: It's brilliant because the retrieval process isn't just looking for similar pictures anymore; it’s looking for causal chains. We aren't just asking, "Does this look like that?" We are asking, "How did event A lead to event B?"

Tom: Which is exactly what makes the difference when comparing this approach to standard RAG methods. It’s not about finding a visual match; it's about reconstructing the logic of the sequence.

Improvements/Methodology Deep Dive: Tom: Let's talk methodology, because I think that’s where the engineering genius really shines. How does Event-Causal RAG actually find these event anchors in a continuous stream?

Meng: They use something called Dual-Sentinel Event Segmentation. They use a visual encoder like SigLIP to track similarity between frames and define triggers where those similarities drop below a high threshold, tau evt.

Jane: That’s how they detect the start of an event—the moment the visual state changes significantly. But it doesn' not stop there; they use a "center-out" expansion to confirm that the event window is stable enough, using thresholds like tau bg.

Tom: And if audio is involved, they integrate an ASR tool to ensure that's captured too, making sure the segmentation isn't just visual but also captures the spoken context.

Lu: Then comes the retrieval phase, which is where things get clever. They use a bidirectional retrieval strategy. Instead of blindly checking every single chunk in the graph, they start with an entry anchor and then perform a maximum of two hops within N=two hops to find relevant events.

Jane: This limited traversal combined with semantic refinement is key for efficiency. The paper introduces a deduplication mechanism where if any retrieved node's text is highly similar to something already seen, they prune it.

Tom: That semantic state collapse—compressing long, unchanging scenes into a single entry—is a huge win for reducing context volume and drastically increasing information density.

Meng: And I noticed that the architecture is remarkably robust; the twenty-four-hour infinite stream test showed it could handle continuous data without memory failure, which is critical for deployment.

Conclusion: Tom: We've seen how Event-Causal RAG tackles the problem of long video reasoning, but let's look at the results. Does this approach actually deliver what it promises?

Jane: The experimental results are very encouraging. On benchmarks like NExT-QA and EventBench, it consistently outperforms both standard clip-based retrieval baselines and models that utilize large context windows.

Lu: And this isn't just about short clips either benefiting from the SES abstraction; the improvements are consistent across all evaluated open-source backbones on EventBench, showing steady gains.

Tom: That’s especially impressive in the Action Reasoning task where it achieved a score of forty-six point six seven percent in Video-MME Long. It shows that when you go to hour-long videos, this structured memory actually helps the models see what's happening dynamically.

Meng: The ability to handle long-horizon causal reasoning is a real win, too. When questions require integrating evidence distributed across multiple events, the dual-store merging recovers that contextual information much better than simple RAG methods do.

Lalam: It seems like Event-Causal RAG provides a practical mechanism for making AI truly capable of understanding narrative flow over long spans of time.

Final Summary & Goodbye: Tom: So, to wrap up the whole discussion on Event-Causal RAG: we've seen that this framework overcomes the limitations of fixed clip memory by using a State-Event-State graph.

Jane: It provides a way to capture both the physical actions and their surrounding context. The combination of dual-sentinel segmentation, global event knowledge graphs, and bidirectional retrieval allows the AI to reconstruct causal chains without being overwhelmed by massive context size.

Tom: And we've seen that this approach is highly efficient; even running a twenty-four-hour surveillance stream on a single RTX five thousand ninety without crashing is incredible.

Lu: It really proves that structured, event-level memory can act as the long-term coherence the AI needs to transition from basic scene recognition to sophisticated causal understanding.

Meng: It’s a grounded, scalable solution that bypasses the OOM issue inherent in traditional methods.

Lalam: I think this is a huge leap toward building an AI that understands stories, not just frames.

Tom: We hope to see more evaluation datasets built for these ultra-long video scenarios to verify the applicability of Event-Causal RAG even further.

Jane: It’s been fascinating to discuss how the authors have created this comprehensive framework.

Tom: Alright, we've spent today talking about Event-Causal RAG; let's see what other groundbreaking research is waiting for us next!

Peizheng Yan, Yu Zhao, Liang Xie, Juntong Qi, Mingming Wang, Erwei Yin

Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen) · Tianjin Key Lab of Intelligent Unmanned Swarm Tech & System, Tianjin University · Tianjin Artificial Intelligence Innovation Center · Defense Innovation Institute · Academy of Military Sciences · School of Future Technology, Shanghai University

cs.AI, cs.CV

Submitted: 2026-08-19

Updated: 2026-08-20

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 87/100

The gist: Event-Causal RAG (EC-RAG) is a retrieval-augmented generation framework designed for "infinite long-video reasoning," addressing the limitations of existing large vision-language models (LVLMs) and

Key concepts

Event-Causal RAG
A retrieval-augmented generation framework designed for long video reasoning. It moves beyond simple clip matching by structuring video information into causal chains, allowing the AI to understand how one event leads to another.
State-Event-State (SES) triplets
Structured data points used by the framework that capture the state of a scene before an action occurs and the state after it. This provides a causal structure, like a storyboard with arrows, rather than just isolated video clips.
Global Event Knowledge Graph (EKG)
A structured memory system built from SES triplets. It stores information in dual-store memory (Vector and Graph Databases) to enable retrieval based on causal relationships and topological mapping.
Dual-Sentinel Event Segmentation
The methodology used to detect the start of an event by tracking visual similarity between frames. It uses specific thresholds ($ au_{evt}$) when similarities drop significantly, confirming a change in the visual state.

Terminology

Summary

Event-Causal RAG (EC-RAG) is a retrieval-augmented generation framework designed for infinite long-video reasoning, addressing the limitations of existing large vision-language models (LVLMs) and traditional Retrieval-Augmented Generation (RAG) approaches in handling ultra-long or infinite video sequences.

Problem Statement and Limitations:

While current LVLMs perform well on short and medium videos, they are inadequate for ultra-long or even infinite video reasoning, where models must maintain "coherent memory over extended durations and infer causal dependencies across temporally distant events. Existing end-to-end video understanding methods are fundamentally limited by the O(n squared) complexity of self-attention. Traditional RAG methods suffer from three critical limitations: (1) They fragment context by constructing memory over fixed-duration clips rather than semantically complete events; (2) they ignore explicit modeling of temporal dependency and causal relations between events"; and (3) they are costly in both storage and online processing.

The EC-RAG Solution:

EC-RAG replaces clip-level memory with a State-Event-State (SES) graph memory, representing each event as a structured graph that captures the event together with its surrounding state transitions. These graphs are unified into a global Event Knowledge Graph and stored in a dual-store memory.

Methodology: The framework operates through three core modules:

  1. Dual-Sentinel Event Segmentation (3.1): This module determines where an event begins and ends using both visual and audio cues.
  • Visual Sentinel: It uses a pre-trained SigLIP visual encoder to calculate the similarity between adjacent frames, S(t) = (Ev(f t), Ev(f t+1)). An event is anchored at t if S(t) falls below a high sensitivity threshold (tau evt = 0.97, S(t) < S(t-1) S(t) < S(t+1)). The system then expands this anchor until a stable background threshold (tau bg = 0.99) is exceeded, defining an event window T.

  • Audio Sentinel: It uses an ASR tool to extract speech, ensuring the temporal partition captures a strictly seamless temporal partition of Dynamic Events and Static Background.

  1. SES Event Memory Constructor (3.2): This module abstracts the events into structured knowledge.
  • The VLM extracts (State-Event-State) SES triplets using a two-step Chain-of-Thought (CoT) guidance, resulting in the structure: PreState to Event to PostState.

  • Dual-Store Memory: The system utilizes a Vector DB (for semantic routing) and a Graph DB (for topological mapping). When new triplets are merged, they are fused if their entity similarity exceeds gamma ent = 0.85. Causal chains are established if the PostState of Event i and the PreState of Event j exhibit semantic similarity (gamma evt = 0.85), creating a [:TEMPORAL NEXT] edge.

  1. Bidirectional Retrieval RAG (3.3): This module queries the memory to provide relevant context to the backbone VLM.
  • It employs a three-step process: Entry Anchoring (querying the VectorDB), Bidirectional Walk Selection (collectively finding nodes within N=2 hops in the GraphDB), and Semantic Refinement.

  • The refinement step uses an embedding model (e) to eliminate redundancy. If a candidate node n has a maximum cosine similarity S max(n) with previously seen nodes exceeding the deduplication threshold tau dup, it is pruned, which allows the system to compress prolonged, unchanging scenes into a single script entry.

Evaluation and Results:

The framework was tested across multiple benchmarks: NExT-QA (5–180s), EventBench (60–1800s), Video-MME Long (1800-3600s), and 24-hour+ streaming surveillance video.

  • Long Video Performance: On the extreme length challenge of Video-MME Long, EC-RAG achieves substantial improvements (+12.50%) in capturing dynamic features like action recognition through semantic dimensionality reduction and graph abstraction.

  • General Benchmarks: On EventBench, EC-RAG consistently outperforms strong clip-based retrieval baselines and long-context video models.

  • Streaming Capability: In a 24-hour industrial surveillance stream, EC-RAG achieved a strict information extraction accuracy of 90.57%, demonstrating its ability to maintain stable event recognition over extensive spatiotemporal spans.

  • Efficiency: The architecture decouples VRAM consumption from the total video duration, allowing the system to process infinite streams with zero memory degradation.

Ablation Study Findings: A study on EventBench showed that removing any core component—such as lacking the Elastic bidirectional diffusion sentinel or failing to perform dual-store topological merging—led to a significant decline in performance, confirming that each component is indispensable.

Improvements for AI systems

The following improvements leverage the core architectural innovations of Event-Causal RAG (EC-RAG) to address fundamental limitations in current large vision-language models (LVLMs) and retrieval-augmented generation (RAG) systems.


Concept: Replace fixed, arbitrary temporal windows with a State-Event-State (SES) Graph structure for modeling sequential data.

Mechanism: Instead of indexing raw video clips, the system identifies semantically coherent events (using the Dual-Sentinel Event Segmentation module). Each event is then represented as an explicit directed edge connecting a PreState (S pre) to a PostState (S post), capturing the physical transformation.

Impact on AI System:

  • What it does: Transforms continuous, unstructured input (video, sensor logs, or even long conversation transcripts) into a discrete, actionable causal graph.

  • Capability: Enables the AI system to perform causal reasoning and multi-event integration. It can answer complex questions like How did the initial state of [Object A] influence the final outcome after [Event B, C]? This overcomes the lost-in-the-middle problem inherent in linear context processing.

Concept: Implement a Dual-Store Memory system that combines high-dimensional semantic matching with explicit topological mapping.

Mechanism: Store data simultaneously in a Vector Database (for semantic similarity) and a Graph Database (for causal structure).

Impact on AI System:

  • What it does: Allows the retrieval mechanism to move beyond simple keyword or visual similarity searches. It retrieves not just what looks similar, but how events are related.

  • Capability: The system can execute multi-hop retrieval. For a query like Find all actions that occurred after the red truck appeared and before the smoke cleared, it navigates the Graph DB to find causal chains, rather than just matching individual clips. This is essential for complex, temporally distributed queries in surveillance or industrial monitoring.

Concept: Apply Semantic Refinement via embedding-based deduplication and graph pruning to reduce context volume without losing critical information.

Mechanism: When querying the retrieved event nodes, the system uses an embedding model to detect if newly found events are semantically identical to already stored events (Smax > tau dup). If redundant, it is pruned.

Impact on AI System:

  • What it does: Drastically compresses long-duration inputs (e.g., 24-hour video streams) into a highly dense, non-redundant set of critical events.

  • Capability: Enables infinite stream processing. The system can maintain high fidelity over massive time spans without suffering from the memory or attention bottlenecks that plague traditional end-to-end models, making it deployable on standard hardware for real-time industrial monitoring.

Concept: Use Sentinel Detection (Visual and Audio) to segment data into semantically coherent chunks rather than fixed time intervals.

Mechanism: Utilize a visual encoder (like SigLIP) to detect sharp fluctuations in frame similarity (tau evt) as event anchors, expanding the window until a stable background threshold (tau bg) is reached. Simultaneously, integrate ASR output into these segments.

Impact on AI System:

  • What it does: Guarante that the input data fed to the LLM is always a complete, physically coherent event, regardless of how many seconds it takes or whether a fixed clip boundary would have split the action.

  • Capability: Provides robust temporal grounding. The system ensures that when answering questions about events, it is not hallucinating based on fragmented data but is referencing a verified, whole causal unit.

The improved AI system can:

  1. Handle Ultra-Long Context: Process continuous streams of data (e.g., 24 hours of surveillance) without memory degradation or O(n 2) computational explosion.

  2. Perform Causal Inference: Accurately determine why an event occurred and how it relates to previous actions, even if those events are separated by a significant time gap.

  3. Improve Retrieval Accuracy: Deliver highly precise answers by combining semantic relevance (what looks like the target) with causal necessity (what links the target to its preceding/subsequent states).

  4. Optimize Resource Use: Achieve high-precision reasoning using compact, event-level memory structures, drastically reducing VRAM and inference costs compared to brute-force full-context processing.

Abstract

Large vision-language models perform well on short- and medium-length video understanding but still struggle to maintain coherent event memory and recover long-range relationships in ultra-long videos. End-to-end methods are limited by visual-token growth and context length, while fixed-segment retrieval often fragments complete events and weakens state-transition modeling.We propose Event-Causal RAG (EC-RAG), a lightweight retrieval-augmented framework for ultra-long and streaming video reasoning. A dual visual-audio sentinel mechanism segments video streams into semantically complete events, represented as State-Event-State (SES) structures that organize observable pre-event states, central events, and post-event states as event-local causal transitions. These transitions are stored in dual vector-graph memory and temporally connected through entity-consistent trajectories. During question answering, bidirectional graph retrieval recovers relevant predecessor and successor events, and answers are generated using both structured memory and the corresponding video evidence.We further introduce ECV-1H, an hour-scale long-video QA benchmark dedicated to directed event-causal reasoning, with all source videos exceeding one hour. It covers over 150 hours of untrimmed video and contains 1,251 fully human-annotated QA pairs. EC-RAG improves overall accuracy by 4.96%--11.67% across three open-source video foundation models and achieves consistent gains across public datasets. On a single RTX 5090 GPU with 32 GB of memory, EC-RAG can continuously process videos while maintaining controlled streaming memory usage.

Sources

Related papers