AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

arXiv:2606.24286 · cs.CL, cs.CV · Submitted 2026-06-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression".

Jane: Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we’re looking at this paper called AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression. It sounds like they are tackling that big problem of understanding long audio and video clips when you have limited context windows or way too much redundant information.

Jane: That’s right, and the core idea is to change how the AI handles compressing those massive amounts of data by treating it like a retrieval task instead of just cutting stuff out blindly.

Lu: Essentially, they are reframing multimodal token compression as a top-K retrieval problem, meaning instead of just taking the first K tokens, you have a big pool and you pick the best ones that answer the user's question.

Meng: So it’s moving away from rigid compression designs where one modality dictates how much of another gets cut, which is a common issue in current methods.

Lalam: It’s about finding the compact subset of tokens that actually supports answering what the user is asking, using three principles they borrowed from information retrieval.

Tom: Exactly. And those three principles are relevance, importance, and diversity—they are the criteria they use to select those best tokens for a fixed budget.

Jane: The paper explains how they implement each of those criteria with specific mechanisms to make sure the selection process is thoughtful rather than random or just based on one type of signal.

Lu: For instance, relevance is handled by text-guided cross-attention that conditions per-token scores on the user query, which makes sure we grab tokens that are relevant to what’s being asked.

Meng: That sounds like a specific way to weigh the textual input against the multimodal data, ensuring we prioritize what matters for the query.

Lalam: And for importance, they use bidirectional video-audio cross-attention within each temporal block, which gives them a signal that doesn't rely solely on the text query when things are sparse.

Tom: That’s smart because it means even if the text prompt is vague, they can still pick tokens that are intrinsically informative from both the video and audio parts of the clip.

Jane: Then there’s diversity, which they handle through something called Temporal-Aware Maximal Marginal Relevance, which tries to penalize similar tokens close together in time to stop them from just picking the same thing over and over.

Lu: That’s a clever way to suppress redundant adjacent tokens while still keeping important events that happen further apart in time.

Meng: So they aren't just looking for the most relevant stuff; they are actively trying to make sure the selected tokens cover different parts of the hour-long content.

Tom: And when you put it all together, this AVOC framework is designed to condense that interleaved video-audio sequence into a compact subset before feeding it to the large language model.

Jane: The big result here is that they achieved state-of-the-art performance on long audio and video benchmarks, specifically surpassing the second-best model by four point nine points on OmniVideoBench and five point five points on LVOmniBench.

Title and authors: Lu: They also maintained robust performance even when testing it with the Audio-Video Needle-in-a-Haystack task for durations as long as one hour.

Meng: So, it seems they solved a problem where previous methods tended to collapse or lose information when dealing with those very long sequences.

Tom: Right, and this really shows how using retrieval principles can give you a better way to manage context in these massive multimodal models. But we have to talk about what the numbers actually mean for us here.

Jane: The compression module itself is computationally lightweight; even when it keeps most of the information, it only adds about one point eight three four seconds to the time it takes for a model like MiniCPM-o four point five to get its first token.

Lu: That’s pretty efficient for a compression method that is doing this kind of detailed scoring and selection process.

Meng: But they also showed that reducing the retention ratio—how much information you keep—can actually speed up prefilling quite a bit, dropping latency nearly nine times when they set it at a retention ratio of zero point one.

Tom: So the practical benefit is that we can trade some context for much faster processing, which is huge for real-world applications where you need quick responses.

Jane: But the authors do mention a few things they couldn't fully address yet in this initial version of AVOC. They point out that because the compression module runs offline, it isn't directly applicable to scenarios where you’re streaming data live.

Lu: And they also noted that their experiments were done at one specific parameter scale, and the way they split the token budget between video and audio—a two:one ratio—is currently a fixed setting rather than something that adjusts based on what content you’re actually looking at <ref:2606.24286#pg0>.

Meng: So for deployment, it means we still have to think about how to make this system work in real-time streaming environments, which is a practical hurdle.

Tom: That leaves us wondering how flexible this retrieval approach can get when you move from one fixed budget split to something that adapts dynamically based on the video and audio content itself.

Jane: And that brings us right up to where we need to wrap up our discussion on AVOC, Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression.

Lu: Overall, the main thing here is successfully instantiating those three criteria—text relevance, bidirectional importance, and temporal diversity—to guide the selection of a compact subset under a tight budget.

Meng: It gives us a solid blueprint for how to structure information retrieval within these multimodal models without just relying on brute-force context window extensions.

Tom: It’s definitely a step toward making Omni-Modal LLMs capable of reasoning over that hour-long content that we see everywhere in the real world.

Jane: We’ll keep an eye on how they expand this framework for more dynamic budget allocation and streaming integration, but AVOC is clearly making serious progress in handling long sequences.

The paper's summary: Tom: So, to wrap up that first bit, AVOC is basically this system that takes those massive hour-long audio and video files, and instead of just chopping them up randomly, it uses a retrieval strategy to pick the most useful pieces.

Jane: Right. The core idea they’re pushing here is treating token compression like searching for things in a giant database where you have to find the best subset that answers your question using three specific rules.

Lu: They break those rules down into relevance, which means matching what you ask for, importance, which is about how much information a piece has on its own across both video and audio, and diversity, which stops the system from just picking the same sound effect over and over.

Meng: So it’s not just about keeping everything; it’s a carefully curated selection process guided by those three criteria to make sure you get the best possible answer from a huge chunk of data.

Tom: Exactly. And the results they're showing are pretty solid, actually beating the second-best models on major long-form benchmarks like OmniVideoBench and LVOmniBench by about five points on average.

Jane: That kind of gain, especially when you’re dealing with hour-long content, is significant because it means the model can actually see and understand what's going on across that entire duration without just getting lost in the noise.

Lu: The Needle-in-a-Haystack test is where this really shines, because they show high accuracy even when you’re looking for a tiny piece of information buried deep in an hour of video or audio.

Tom: And what I find really interesting from the methodology is how they handle the scoring—they combine text relevance and importance using a Z-score normalization to get that final per-token score.

Jane: That combination method is clever because it lets them balance what’s relevant to the query with what’s intrinsically informative across both modalities, even when one signal is weak.

Meng: From an engineering standpoint, the training setup using a differentiable top-K selection strategy with Gumbel-Softmax means they can actually train this whole compression module end-to-end with the large language model itself.

Lu: That makes it much more integrated than just adding a separate post-processing step; it’s baked into the learning process.

Tom: But we have to be real about what this doesn't do yet, which is that this entire compression module runs offline, so right now it's not set up for streaming data in real-time applications.

Jane: That’s a big caveat; it’s fantastic for understanding long clips you can process beforehand, but deploying it live as data flows would require some changes.

Meng: And they also have to deal with that fixed two-to-one split between video and audio budgets, which means the system isn't automatically adjusting how much focus it puts on one modality versus the other based on what’s actually in the content.

Lu: That fixed ratio is definitely a limitation; I think making that budget allocation dynamic based on context would open up a lot more possibilities for true omni-modal reasoning.

Tom: So, we’ve seen how they get these impressive accuracy numbers and the clever retrieval logic behind them, but now we need to think about how to make this framework flexible enough for real-world, live applications.

Jane: That’s exactly where we need to take our next look—how do you make that budget allocation adapt on the fly?

The paper's improvements: Tom: So we’ve seen how AVOC works—how it uses retrieval principles to pick the best parts of long audio and video clips for an AI model—and now we’re looking at what they suggest to do next.

Jane: They aren't just stopping there; they are pointing out a few things that could make this framework much more powerful for future AI work.

Lu: They mentioned that the current setup is somewhat rigid, especially with how the budget is split between video and audio—a fixed ratio of two to one—and they think that needs to be content-aware.

Meng: That means instead of a set rule, the system should look at what's actually in the clip and decide dynamically how much space to give each modality.

Tom: Right, so we’re moving from a fixed budget split to something that adapts based on the actual video and audio content being analyzed.

Jane: And they also touched on how this whole compression module is currently offline, which means it’s not ready for things like real-time streaming where data comes in continuously.

Lu: That’s a practical limitation, but I think the creative potential there is huge; imagine an AI that can dynamically adjust its context window based on the complexity of the video scene versus the dialogue.

Tom: It sounds wild, Lu, but for a system to be truly useful in many real-world scenarios—like live monitoring or interactive applications—it needs that ability to scale and adapt its resource usage.

Jane: And they also hinted at needing more research into how this compression fits into larger systems, not just as a standalone tool, but as part of a bigger pipeline.

Meng: I’m thinking about the engineering side here; if we can get that dynamic allocation working, it would change how we design video processing pipelines entirely.

Lu: It opens up possibilities for creating truly flexible multimodal agents that aren't locked into a specific way of handling their input data.

Tom: So, the paper is suggesting a move away from these fixed parameters toward something more intelligent and context-driven in how it manages its information budget.

Jane: It’s about moving from a set design to an adaptive one that can intelligently manage the trade-off between video detail and audio detail depending on the task at hand.

Meng: That adaptability is what makes it move from a strong benchmark result to something that actually gets deployed in complex, messy situations.

Lu: It’s about giving the AI more autonomy over its own context management rather than having a hard-coded set of rules guiding every token selection decision.

Conclusion: Tom: So we’re wrapping up this chat on AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression by summarizing what they did and where it goes from here.

Jane: Basically, they figured out that instead of just cutting up long audio and video clips randomly, you can treat the compression process like a smart search engine picking the most useful tokens based on relevance, importance, and diversity.

Lu: It’s a very structured way to handle context; it’s taking those complex multimodal sequences and turning them into a highly curated set that directly supports the AI's query.

Meng: The implication for practical systems is that we can get much better performance on long-form content without having to dramatically increase the context window size, which saves a ton of computational resources.

Tom: Right, and those numbers they showed—surpassing the second-best models by about five points on major benchmarks—show that this retrieval approach actually yields tangible gains in how well the AI understands long sequences.

Jane: It changes how we think about context management; it’s less about brute force and more about intelligent selection guided by those three specific criteria.

Lu: I think the real future here is making that budget allocation dynamic, as we talked about earlier, because right now it’s a fixed ratio of two to one which limits its flexibility.

Meng: If we can get that content-aware adjustment working, it would fundamentally change how we design video processing pipelines for any long-form AI task.

Tom: So, the paper is a solid piece of work because it proves that you can use established information retrieval principles to tackle the massive problem of hour-long audio and video understanding.

Jane: It gives us a concrete framework for making those large multimodal models actually handle the richness of real-world, long-form content.

Lalam: For me, this means our language model can better grasp complex narratives across extended media, which is how we’ll eventually improve the cultural context and storytelling capabilities of AI.

Lu: I’m really excited to see researchers take this compression logic and apply it to other challenging areas where context management is the bottleneck.

Meng: From an engineering standpoint, I think this moves us closer to building more efficient, scalable tools for video processing that don't require astronomical amounts of memory just to keep track of everything.

Tom: Well said. We’ve seen how AVOC uses retrieval-inspired token compression to boost understanding on long audio and video benchmarks.

Jane: It’s a big step toward making AI models capable of truly comprehending the rich, extended content we encounter every day.

Lu: Keep an eye on that dynamic budget allocation idea; it's where the next level of innovation could happen.

Meng: We’ll be watching how they integrate this module into real-time processing frameworks next, because that’s where we need to see it applied practically.

Yijing Chen, Wenhui Tan, Xiaoyi Yu, Yuyue Wang, Xin Cheng, Kaisi Guan

Gaoling School of Artificial Intelligence, Renmin University of China

cs.CL, cs.CV

Submitted: 2026-06-23

Updated: 2026-10-05

Comments: Accepted at NeurIPS

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 88/100

The gist: Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and

Key concepts

Top-K Retrieval Problem
Instead of compressing all tokens, AVOC treats token selection as finding the best subset (size K) from a large pool to answer a query. This mimics information retrieval, ensuring the chosen tokens are highly relevant to the user's question while maintaining diversity.
Relevance Scoring
This measures how well a token supports the user's text query. It uses text embeddings and cross-attention mechanisms to score each token based on its direct relevance to what the user is asking, guiding the selection process.
Importance Scoring
This assesses the intrinsic informativeness of a token within a temporal block by using bidirectional cross-attention between video and audio. This provides a query-agnostic signal that helps select tokens crucial for understanding the scene or event, even if the text query is vague.

Terminology

Summary

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. This paper introduces AVOC, a framework that addresses these bottlenecks by reframing multimodal token compression as a top-K retrieval problem to enable hour-level audio-video understanding in OmniModal LLMs.

The gist

AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively, while maintaining robust performance on Audio-Video Needle-in-a-Haystack task at durations up to one hour.<ref:2606.24286#pg11>

How it works

The core of AVOC is a learnable token compression module placed between the modality encoders and the LLM backbone, which redefines multimodal token compression as a top-K retrieval problem: given a fixed context budget and a large pool of candidate tokens, the model must retrieve a compact subset which best supports answering the user query.<ref:2606.24286#pg11> This approach leverages classical Information Retrieval (IR) principles, specifically three criteria: query-conditioned relevance, query-agnostic importance, and result diversity.<ref:2606.24286#pg11>

The module instantiates these three criteria via tailored mechanisms:

  1. Relevance is computed via text-guided cross-attention that conditions per-token scores on the user query.<ref:2606.24286#pg11>

  2. Importance is computed via bidirectional video-audio cross-attention within each temporal block, providing a query-agnostic signal that complements relevance when the textual query is sparse.<ref:2606.24286#pg11>

  3. Diversity is enforced through Temporal-Aware Maximal Marginal Relevance, which penalizes similarities within a local temporal window, suppressing redundant adjacent tokens while preserving recurring events that are temporally distant.<ref:2606.24286#pg11>

Methodology and Components

The process involves several stages to condense the interleaved multimodal token sequence X into a compact subset S of size K under a fixed budget K.<ref:2606.24286#pg11>

** Relevance scoring uses text embeddings projected into query and key spaces, followed by scaled dot-product attention to compute scorerel(xi) = 1/Ntext X j A rel j,i.<ref:2606.24286#pg11>**

** Importance scoring utilizes bidirectional cross-attention within temporal blocks between the video and audio modalities to estimate intrinsic informativeness, yielding scoreimp(xi) = 1/Nm¯ i X j (Am¯ imi)j,i.<ref:2606.24286#pg11>**

** A combined per-token score is defined by Z-score normalization of the relevance and importance scores: score(xi) = 1/2 score′ rel(xi) + score′ imp(xi).<ref:2606.24286#pg11>**

** Diversity selection employs Temporal-Aware Maximal Marginal Relevance (TA-MMR), which penalizes similarity within a local temporal window defined by Window(τi) = [τi − W, τi + W] defines the local temporal scope centered at τi with radius W, and restricts similarity to tokens of the same modality.<ref:2606.24286#pg11>**

** The budget K is split into a modality-aware budget (Kvideo, Kaudio) with a ratio of 2:1, and TA-MMR is performed independently within each modality to ensure a balanced cross-modal representation.<ref:2606.24286#pg11>**

Experimental Validation

AVOC was trained on 40k samples from datasets including AVSD [1], How2 [28], FineVideo [12], ChronusAV [5], and LongVILA sft [7]. The training involved a two-stage process where the compression module was initially disabled, followed by joint fine-tuning with the LLM, utilizing a differentiable top-K selection strategy based on Gumbel-Softmax to enable end-to-end gradient propagation.<ref:2606.24286#pg11>

Performance evaluation was conducted on WorldSense [13], OmniVideoBench [17], and LVOmniBench [36]. AVOC consistently achieved state-of-the-art results, with absolute gains of 1.7–7.2 points over the second-best results across the benchmarks.<ref:2606.24286#pg11>

Furthermore, in the Audio-Video Needle-in-a-Haystack task, AVOC maintained high retrieval accuracy across the entire duration and depth grid for both modalities, demonstrating its capability in ultra-long audio-video context modeling compared to baselines like OmniZip which exhibit a clear duration-induced collapse.<ref:2606.24286#pg11>

Efficiency and Limitations

The compression module is computationally lightweight; even at full retention, it adds only merely 1.834 s to the time to first token for MiniCPM-o 4.5.<ref:2606.24286#pg11> Reducing the retention ratio yields substantial prefilling speedups, with latency dropping nearly nine times at a retention ratio of 0.1.<ref:2606.24286#pg11>

The paper identifies three primary limitations: first, the compression module operates in an offline manner, preventing direct application to streaming scenarios; second, experiments are conducted at a single parameter scale, and third, the modality token budget allocation ratio Kvideo: Kaudio is currently set as a fixed hyperparameter rather than being content-adaptive.<ref:2606.24286#pg11>

The framework successfully instantiates three complementary criteria—text-guided relevance, bidirectional video-audio importance, and Temporal-Aware Maximal Marginal Relevance—to guide the selection of a compact, informative token subset under a tight context budget. AVOC offers a step toward Omni-Modal LLMs capable of reasoning over the rich, hour-long multimodal content that pervades real-world applications.

REFERENCES

[1] Huda AlAmri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Peter Anderson, Stefan Lee, and Devi Parikh. Audio visual scene-aware dialog. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2019), Long Beach, CA, USA (June 16-20, 2019), pages 7558–7567.<ref:2606.24286#pg14>

[2] Jaime Carbonell and Jade Goldstein. The use of mmr, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval (SIGIR 1998), pages 335–336, 1998.<ref:2606.24286#pg11>

[3] Jianghan Chao, Jianzhang Gao, Wenhui Tan, Yuchong Sun, Ruihua Song, and Liyun Ru. Jointavbench: A benchmark for joint audio-visual reasoning evaluation. arXiv preprint arXiv:2512.12772 (2025).<ref:2606.24286#pg14>

[4] Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models. In Ales Leonardis et al., Computer Vision - ECCV 2024 (September 29-October 4, 2024), Proceedings, Part LXXXI, Lecture Notes in Computer Science, pages 19–35.<ref:2606.24286#pg14>

[5] Yijing Chen, Yihan Wu, Kaisi Guan, Yuchen Ren, Yuyue Wang, Ruihua Song, and Liyun Ru. Chronusomni: Improving time awareness of omni large language models. arXiv preprint arXiv:2512.09841 (2025).<ref:2606.24286#pg14>

[6] Yukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu, Hanrong Ye, Ligeng Zhu, Zhijian Liu, Pavlo Molchanov, Jan Kautz, Xiaojuan Qi, Sifei Liu, Hongxu Yin, Yao Lu. Scaling rl to long videos. In Advances in Neural Information Processing Systems (NeurIPS 2025).<ref:2606.

Improvements for AI systems

  1. The AVOC framework can be applied to long-form audio-video understanding in Omni-modal Large Language Models (OLLMs) by reframing multimodal token compression as a top-K retrieval problem. This allows the system to select a compact subset which best supports answering the user query from a large candidate pool, leveraging classical Information Retrieval criteria: query-conditioned relevance, query-agnostic importance, and result diversity.

  2. The system can achieve state-of-the-art performance on long-form audio-video benchmarks by using three tailored mechanisms within the compression module: text-guided cross-attention for query-conditioned relevance, bidirectional video-audio cross-attention within each temporal block for query-agnostic importance, and Temporal Aware Maximal Marginal Relevance Selecting for local diversity.

  3. The improved system can perform fine-grained retrieval over ultra-long multimodal content, maintaining high accuracy on the Audio-Video Needle-in-a-Haystack task at durations up to one hour, as demonstrated by AVOC's performance in Figure 6 (h).

Abstract

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video understanding in Omni-modal Large Language Models. AVOC introduces a learnable token compression module between the modality encoders and the LLM backbone. We reframe multimodal token compression as a top- K retrieval problem: given a fixed context budget, the module must retrieve a compact subset of tokens that best supports answering the user query. We draw inspiration from three classical Information Retrieval criteria for selecting informative units from a large candidate pool: relevance, importance, and diversity. AVOC instantiates each criterion as a tailored mechanism for audio-video understanding, and integrates them into a unified retrieval-style compression pipeline. Experiments show that AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively. Moreover, AVOC maintains robust performance on Audio-Video Needle-in-a-Haystack task at durations up to one hour. Code and model are at github.com/YJCX330/AVOC.

Sources

Related papers