LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

arXiv:2609.39938 · cs.CL, cs.AI, cs.CV · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception".

Tom: Detailed Research Summary of LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception Paper Title: LEAP (Learned Block-wise Evidence Retrieval for Long Audio-Video Perception) Core Concept:

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Alright everyone, we've been deep into the technical details of the LEAP framework, and now it's time for us to wrap up what this whole paper actually means for us as listeners. We’re talking about how they tackled the challenge of long audio and video perception.

Jane: Exactly; it really boils down to finding specific pieces of evidence in hours of video or audio without having to listen or watch everything at once, which is a big deal when you're dealing with massive media files. The authors put a lot of work into making that retrieval process incredibly efficient so we don't have to worry about the context getting overwhelmed.

Lu: From a theoretical standpoint, I think what they’ve achieved is really about managing complexity over massive temporal scales; they’re using the block partitioning to keep the working memory footprint constant.

Meng: For engineering, I see that constant footprint translating into something practical for deployment; if we can handle huge recordings with predictable resource usage, that opens up a lot of doors for real-time applications.

Lalam: For me, this is huge because it means our models can finally grasp the long-term narrative and subtle cues in complex media more effectively, which could fundamentally improve how we understand human culture across vast datasets.

Tom: So, to sum up the core idea of LEAP from "Learned Block-wise Evidence Retrieval for Long Audio-Video Perception," it’s about learning how to grab exactly what you need from a massive file without having to process the entire thing every time.

Jane: And the authors of this paper are really clever because they developed these new techniques and training methods, like those LoRA adapters, to make that retrieval smarter and more accurate.

Lu: Their methodology is interesting because they decoupled the search phase from the reasoning phase using that transcript channel; it’s a very smart architectural choice for handling long-form data.

Meng: From an engineering viewpoint, I see the real win here being that we don't need to scale our compute linearly with the length of the video when performing these kinds of detailed searches.

Paper summary: Lalam: I think this has profound implications for how we build systems that understand context; imagine analyzing decades of historical footage with this kind of targeted retrieval capability built in.

Tom: It’s clear that LEAP is pushing the boundaries on how much information we can extract from long media files while keeping the computational load manageable.

Jane: So, while we've covered the mechanics, what are some of the bigger impacts this technology might have on our daily lives or future applications?

Lu: I think because they decoupled evidence localization from reasoning, it lets us search over pre-computed transcripts without actually decoding those media frames during the localization pass.

Meng: That decoupling sounds promising for efficiency; if we can search the transcript index first, that saves a lot of processing power before we even look at the actual video data.

Lalam: And that capability to handle historical retrieval as the stream arrives means systems could potentially understand events in real-time as they unfold, rather than waiting for everything to finish.

Tom: That idea of handling causal queries over a streaming input is compelling; it moves us closer to systems that can keep up with the flow of live data.

Jane: It really simplifies the whole workflow because we aren't stuck trying to cram an entire hour of video into a single context window for every question.

Lu: The authors specifically designed this so that the peak context and working memory requirements remain independent of the total recording duration, which is a significant achievement in managing temporal scale.

Meng: I want to circle back to implementation: even with these clever techniques, we still need robust systems for efficiently partitioning those long media files into those fixed-duration blocks effectively.

Lalam: That partitioning step is where the real practical work happens; if that initial split isn't done well, the retrieval won't be effective regardless of how clever the subsequent passes are.

Paper summary: Tom: So, we’ve established that LEAP provides a method for retrieving relevant evidence without loading the whole recording into memory, and now we’re looking at how this affects future development.

Jane: It really boils down to finding specific pieces of evidence in hours of video or audio without having to listen or watch everything at once, which is a big deal when you're dealing with massive media files.

Lu: The way they handle the search phase by allowing localization passes to score candidate windows over either the transcript or the media is a very smart architectural choice for handling long-form data.

Meng: I think that ability to manage resource usage predictably regardless of duration is what really matters for scaling these kinds of applications in real-world scenarios.

Lalam: This has profound implications for how we build systems that understand context; imagine analyzing decades of historical footage with this kind of targeted retrieval capability built in.

Tom: It’s clear that LEAP is pushing the boundaries on how much information we can extract from long media files while keeping the computational load manageable.

Jane: So, while we've covered the mechanics, what are some of the bigger impacts this technology might have on our daily lives or future applications?

Lu: The authors address existing approaches that rely on either selection-based methods which keep frames or clips before reasoning, or compression-based methods which reduce visual tokens to preserve coverage.

Meng: Their method seems to bypass those trade-offs by focusing purely on bounded retrieval rather than trying to select the best subset upfront before processing.

Lalam: For me, this means we can finally build tools that can accurately answer complex questions about very long media files, which is something many current omni-modal models struggle with due to their finite context windows.

Tom: Exactly! So, to sum up the core idea of LEAP from "Learned Block-wise Evidence Retrieval for Long Audio-Video Perception," it’s about learning how to grab exactly what you need from a massive file without having to process the entire thing every time.

Conclusion: Tom: We’ve just seen how LEAP works in action, but let's bring it back to the basics for a second and talk about what that title means: "Learned Block-wise Evidence Retrieval for Long Audio-Video Perception."

Jane: That title tells us exactly what the paper is aiming for: taking long audio and video, breaking them down into blocks, and then learning how to find the right evidence from those specific chunks.

Lu: What’s interesting about that title is the "block-wise" part; it suggests a structured way of looking at time that isn't just processing everything as one giant stream.

Meng: The authors are Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi1 and Xinru Jiang1 and Yanzhi Wang1; it’s a solid team of researchers working on this specific problem.

Lalam: For me, the implication is that we can finally tackle those really long media files that current models just can't hold in their memory when answering complex questions.

Tom: Exactly! So, if you boil it down for the listeners, LEAP is basically a smart system designed to hunt down specific clues in hours of video without having to watch every single second of it.

Jane: It tackles the context problem by dividing the media into manageable blocks and then using a focused pass to pull out only what’s actually relevant before generating an answer.

Lu: The core idea is keeping the memory usage constant, no matter how long those videos are, which is a significant theoretical achievement in managing temporal scales.

Meng: From my side, the practical implication is that we can deploy this on systems that need to handle massive media inputs because the computational cost stays within predictable bounds.

Lalam: This capability could fundamentally change how we analyze historical records or long-form events, making deep cultural understanding much more accessible to everyone.

Tom: It’s clear that LEAP is pushing the boundaries on how much information we can extract from long media files while keeping the computational load manageable, and next up, we're going to look at exactly how they did it technically.

Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi1, Xinru Jiang

Northeastern University

cs.CL, cs.AI, cs.CV

Submitted: 2026-09-30

Updated: 2026-09-30

Comments: 39 pages, 16 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: Paper Title: LEAP (Learned Block-wise Evidence Retrieval for Long Audio-Video Perception) Core Concept: LEAP is a novel framework designed to retrieve relevant evidence from long audio and video

Key concepts

Block Partitioning
The long audio or video file is split into smaller, fixed-duration chunks called blocks. These blocks overlap slightly to ensure smooth transitions when searching for evidence across the entire duration of the recording.
Lightweight Localization
A dedicated process that scans only one block at a time to score short candidate windows based on the query. This ensures that during the search phase, only a small, constant amount of context is needed in memory, keeping complexity low.
Transcript as a Second Channel
The pre-computed text transcript serves as an extra way to scan for information. This decouples finding evidence from reasoning; the system can search the text index quickly before diving into the raw audio/video for detailed verification.
Bounded Cost
The computational cost of LEAP is strictly limited by a fixed set of constants, not by how long the original recording is. This means memory usage and processing time remain constant ($O(1)$) even if you are analyzing a very long video.

Terminology

Summary

Paper Title: LEAP (Learned Block-wise Evidence Retrieval for Long Audio-Video Perception)

Core Concept: LEAP is a novel framework designed to retrieve relevant evidence from long audio and video recordings without requiring the entire recording to be processed in a single context. Its fundamental innovation lies in achieving constant, O(1) complexity with respect to recording duration while maintaining high retrieval accuracy.

LEAP operates by intelligently dividing long media files into fixed-duration blocks and employing a two-stage process: evidence localization followed by answer generation.

1. Evidence Retrieval Strategy (Localization Pass):

  • Block Partitioning: The recording (V) is first partitioned into overlapping, fixed-duration blocks (x 1,, x M).

  • Lightweight Localization: A dedicated localization pass (LOCALIZATIONPASS) is applied to each block. This pass scores short candidate windows within the block based on the query (q).

  • Evidence Selection: The highest-ranked windows from all blocks are pooled to form a set of selected evidence windows (U), which is bounded by a constant number of windows (BW=9 in one specific procedure).

  • Context Efficiency: Crucially, this localization pass reads only one block at a time, ensuring that the peak context and working memory required for this stage remain O(1) with respect to the total duration (T) of the recording.

2. Reasoning Strategy (Answer Pass):

  • Transcript as a Second Channel: LEAP decouples evidence localization from reasoning by utilizing a pre-computed transcript as a secondary scanning channel. This allows the same block grid used for localization to be searched on transcripts without decoding media frames for every search step.

  • Re-reading Selected Evidence: The final answer pass (ANSWER) re-reads the selected windows (U) from the raw audio and video streams, using a single, bounded pass that incorporates the outline (omega), question (q), and context derived from these specific windows.

3. Training Methodology (LoRA Adapters):

The framework is trained using two specialized LoRA adapters:

  • Localization LoRA: Improves the performance of the localization pass in selecting relevant windows. The objective for this training is defined within a single block, ensuring it never reads more than one block at a time.

  • Answer LoRA: Improves the quality of the answers generated from the selected evidence windows.

LEAP introduces several paradigm-shifting advantages:

  1. Evidence Retrieval Without Whole-Recording Read: The primary contribution is narrowing down full-length recordings to a bounded set of evidence-rich windows, achieving O(1) context and working memory complexity regardless of the recording's duration.

  2. The Transcript as a Second Scanning Channel: This decouples localization from reasoning. It allows the block grid to be searched on pre-computed transcripts, while the answer pass can selectively re-read raw audio/video streams for fine-grained visual and non-speech evidence that a transcript might miss.

  3. Historical Retrieval as the Stream Arrives (Causal Queries): The block grid natively supports causal queries. This enables streaming inference without requiring streaming-specific training, as the localization pass can select windows over media blocks or an online transcript index up to query time, and the answer pass re-reads them from raw media.

  4. Bounded Cost: The framework's computational cost is strictly bounded by a function of configured constants: n ans n q + n omega + BW times n delta + n sep, where every term on the right is a constant, ensuring peak context and memory per pass are O(1) in duration.

LEAP was evaluated across four rigorous benchmarks: TraceAV-Bench, LVOmniBench, VideoOdyssey-AV (videos exceeding 60 minutes), and MMOU.

  • Superior Performance: Across several AVQA benchmarks, LEAP consistently outperformed the Qwen3-Omni-30B-A3B baseline by a significant margin of 4.5% to 16.8%. Furthermore, when transferred to a second omni-modal backbone (MiniCPM-o 4.5), LEAP surpassed its published results by 3.1% to 13.0%.

  • Robustness Across Duration: The framework demonstrates no systematic collapse with respect to recording duration and leads significantly in every quintile on MMOU when duration is a factor, surpassing baselines that read the whole clip or sample uniformly/frame-matched.

Improvements for AI systems

As a fastidious researcher, I have analyzed the LEAP framework and its performance across various benchmarks. Based on this scientific paper, here are specific improvements for AI systems and what those improved systems can achieve:


The core improvement is the introduction of an evidence-retrieval framework that decouples evidence localization from reasoning, allowing models to effectively handle hour-scale audio-visual data without exhausting context limits.

Here are the specific improvements and capabilities derived from LEAP:

  1. Improve handling of long-form AVQA by achieving a peak context and working memory that is independent of recording duration (O(1) in duration).

  2. Enable efficient, bounded reasoning over hour-scale audio-visual recordings without requiring the entire recording to be processed in one context window.

  3. Preserve fine-grained acoustic and visual evidence by routing the final answering pass over raw audio-visual streams instead of relying solely on transcript outlines.

Specific improvements achievable with LEAP:

  1. Improve accuracy on long-form AVQA benchmarks (e.g., TraceAV, LVOmniBench, VideoOdyssey) by achieving significant gains ranging from 4.5% to 16.8% over state-of-the-art baselines (Qwen3-Omni).

  2. Enhance evidence coverage on complex benchmarks like VideoOdyssey and MMOU, where LEAP demonstrates superior performance compared to baseline methods, especially when using a trained localization adapter that learns to select evidence spans based on annotated events.

  3. Achieve better performance in causal access scenarios (e.g., StreamArena) by utilizing the block grid structure for historical retrieval, without requiring specific streaming-only training protocols.

  4. Optimize inference latency by reducing query-to-answer time significantly, particularly when leveraging a transcript index as an alternative retrieval channel instead of decoding media frames repeatedly.

  5. Improve robustness to video length variations (duration), as LEAP shows no systematic accuracy collapse with longer videos and even rises on benchmarks like MMOU with duration.

Specific capabilities of the improved AI system:

  1. A model can accurately answer complex questions about long videos (hours-long) by intelligently retrieving only the few relevant minutes or seconds of evidence, rather than processing irrelevant content from the entire recording.

  2. The system can pinpoint exact temporal windows for specific events (e.g., Find the moment X happened) with high precision, even when that event is scattered across a long video, by using a two-stage process: first localizing candidate windows via a lightweight pass and then re-encoding only those selected windows for the final answer generation.

  3. The system can maintain high fidelity for fine-grained visual details (e.g., subtle movements or non-speech audio cues) by routing the final answering pass over raw audio-visual streams, ensuring no evidence is lost during the reasoning phase, unlike systems that rely solely on pre-computed transcripts.

  4. The system can operate effectively in streaming environments by using a causal query protocol to only read the relevant preceding context from a pre-computed transcript index, enabling near real-time answering based on what has already arrived.

Abstract

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence without placing the whole recording in one context. LEAP divides a recording into fixed-duration blocks, applying a lightweight localization pass to each block to score short candidate windows. The highest-ranked windows are pooled and re-encoded in a single bounded answer pass. Consequently, the answer input and peak context remain independent of the recording duration. By decoupling evidence localization from reasoning, our framework can localize candidate temporal windows over pre-computed transcripts without decoding media frames, while preserving fine-grained visual and non-speech evidence by routing the final answering pass over raw audio-visual streams. LEAP trains both stages: a localization LoRA improves the selected windows, and an answer LoRA improves the answers read from the same windows. The block grid natively supports causal queries, enabling LEAP to support streaming inference without streaming-specific training. Across several AVQA benchmarks, LEAP improves over the Qwen3-Omni-30B-A3B baseline by 4.5-16.8%, and transfers to a second omni-modal backbone, MiniCPM-o 4.5, surpassing its published results by 3.1-13.0%.

Sources

Related papers