AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression
summary
The gist
Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and
In short
AVOC addresses challenges in understanding long audio-video content by reframing token compression as a top-K retrieval problem. It introduces a learnable module that selects the most relevant and diverse tokens from hour-long sequences, achieving state-of-the-art performance on benchmarks like OmniVideoBench.
Key concepts
- Top-K Retrieval Problem
- Instead of compressing all tokens, AVOC treats token selection as finding the best subset (size K) from a large pool to answer a query. This mimics information retrieval, ensuring the chosen tokens are highly relevant to the user's question while maintaining diversity.
- Relevance Scoring
- This measures how well a token supports the user's text query. It uses text embeddings and cross-attention mechanisms to score each token based on its direct relevance to what the user is asking, guiding the selection process.
- Importance Scoring
- This assesses the intrinsic informativeness of a token within a temporal block by using bidirectional cross-attention between video and audio. This provides a query-agnostic signal that helps select tokens crucial for understanding the scene or event, even if the text query is vague.
Terminology used across episodes
This episode discusses
- AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression · Paper Radio
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- ChronusOmni: Improving Time Awareness of Omni Large Language Models
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- Baichuan-Omni-1.5 Technical Report
- OmniBench: Towards The Future of Universal Omni-Language Models
- Ola: Pushing the Frontiers of Omni-Modal Language Model
- How2: A Large-scale Dataset for Multimodal Language Understanding
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
- Qwen3-Omni Technical Report
- Qwen2.5-Omni Technical Report
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
The paper
AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression · Read on arXiv
Yijing Chen, Wenhui Tan, Xiaoyi Yu, Yuyue Wang, Xin Cheng, Kaisi Guan
Gaoling School of Artificial Intelligence, Renmin University of China
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression".
Jane: Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we’re looking at this paper called AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression. It sounds like they are tackling that big problem of understanding long audio and video clips when you have limited context windows or way too much redundant information.
Jane: That’s right, and the core idea is to change how the AI handles compressing those massive amounts of data by treating it like a retrieval task instead of just cutting stuff out blindly.
Lu: Essentially, they are reframing multimodal token compression as a top-K retrieval problem, meaning instead of just taking the first K tokens, you have a big pool and you pick the best ones that answer the user's question.
Meng: So it’s moving away from rigid compression designs where one modality dictates how much of another gets cut, which is a common issue in current methods.
Lalam: It’s about finding the compact subset of tokens that actually supports answering what the user is asking, using three principles they borrowed from information retrieval.
Tom: Exactly. And those three principles are relevance, importance, and diversity—they are the criteria they use to select those best tokens for a fixed budget.
Jane: The paper explains how they implement each of those criteria with specific mechanisms to make sure the selection process is thoughtful rather than random or just based on one type of signal.
Lu: For instance, relevance is handled by text-guided cross-attention that conditions per-token scores on the user query, which makes sure we grab tokens that are relevant to what’s being asked.
Meng: That sounds like a specific way to weigh the textual input against the multimodal data, ensuring we prioritize what matters for the query.
Lalam: And for importance, they use bidirectional video-audio cross-attention within each temporal block, which gives them a signal that doesn't rely solely on the text query when things are sparse.
Tom: That’s smart because it means even if the text prompt is vague, they can still pick tokens that are intrinsically informative from both the video and audio parts of the clip.
Jane: Then there’s diversity, which they handle through something called Temporal-Aware Maximal Marginal Relevance, which tries to penalize similar tokens close together in time to stop them from just picking the same thing over and over.
Lu: That’s a clever way to suppress redundant adjacent tokens while still keeping important events that happen further apart in time.
Meng: So they aren't just looking for the most relevant stuff; they are actively trying to make sure the selected tokens cover different parts of the hour-long content.
Tom: And when you put it all together, this AVOC framework is designed to condense that interleaved video-audio sequence into a compact subset before feeding it to the large language model.
Jane: The big result here is that they achieved state-of-the-art performance on long audio and video benchmarks, specifically surpassing the second-best model by four point nine points on OmniVideoBench and five point five points on LVOmniBench.
Title and authors: Lu: They also maintained robust performance even when testing it with the Audio-Video Needle-in-a-Haystack task for durations as long as one hour.
Meng: So, it seems they solved a problem where previous methods tended to collapse or lose information when dealing with those very long sequences.
Tom: Right, and this really shows how using retrieval principles can give you a better way to manage context in these massive multimodal models. But we have to talk about what the numbers actually mean for us here.
Jane: The compression module itself is computationally lightweight; even when it keeps most of the information, it only adds about one point eight three four seconds to the time it takes for a model like MiniCPM-o four point five to get its first token.
Lu: That’s pretty efficient for a compression method that is doing this kind of detailed scoring and selection process.
Meng: But they also showed that reducing the retention ratio—how much information you keep—can actually speed up prefilling quite a bit, dropping latency nearly nine times when they set it at a retention ratio of zero point one.
Tom: So the practical benefit is that we can trade some context for much faster processing, which is huge for real-world applications where you need quick responses.
Jane: But the authors do mention a few things they couldn't fully address yet in this initial version of AVOC. They point out that because the compression module runs offline, it isn't directly applicable to scenarios where you’re streaming data live.
Lu: And they also noted that their experiments were done at one specific parameter scale, and the way they split the token budget between video and audio—a two:one ratio—is currently a fixed setting rather than something that adjusts based on what content you’re actually looking at <ref:2606.24286#pg0>.
Meng: So for deployment, it means we still have to think about how to make this system work in real-time streaming environments, which is a practical hurdle.
Tom: That leaves us wondering how flexible this retrieval approach can get when you move from one fixed budget split to something that adapts dynamically based on the video and audio content itself.
Jane: And that brings us right up to where we need to wrap up our discussion on AVOC, Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression.
Lu: Overall, the main thing here is successfully instantiating those three criteria—text relevance, bidirectional importance, and temporal diversity—to guide the selection of a compact subset under a tight budget.
Meng: It gives us a solid blueprint for how to structure information retrieval within these multimodal models without just relying on brute-force context window extensions.
Tom: It’s definitely a step toward making Omni-Modal LLMs capable of reasoning over that hour-long content that we see everywhere in the real world.
Jane: We’ll keep an eye on how they expand this framework for more dynamic budget allocation and streaming integration, but AVOC is clearly making serious progress in handling long sequences.
The paper's summary: Tom: So, to wrap up that first bit, AVOC is basically this system that takes those massive hour-long audio and video files, and instead of just chopping them up randomly, it uses a retrieval strategy to pick the most useful pieces.
Jane: Right. The core idea they’re pushing here is treating token compression like searching for things in a giant database where you have to find the best subset that answers your question using three specific rules.
Lu: They break those rules down into relevance, which means matching what you ask for, importance, which is about how much information a piece has on its own across both video and audio, and diversity, which stops the system from just picking the same sound effect over and over.
Meng: So it’s not just about keeping everything; it’s a carefully curated selection process guided by those three criteria to make sure you get the best possible answer from a huge chunk of data.
Tom: Exactly. And the results they're showing are pretty solid, actually beating the second-best models on major long-form benchmarks like OmniVideoBench and LVOmniBench by about five points on average.
Jane: That kind of gain, especially when you’re dealing with hour-long content, is significant because it means the model can actually see and understand what's going on across that entire duration without just getting lost in the noise.
Lu: The Needle-in-a-Haystack test is where this really shines, because they show high accuracy even when you’re looking for a tiny piece of information buried deep in an hour of video or audio.
Tom: And what I find really interesting from the methodology is how they handle the scoring—they combine text relevance and importance using a Z-score normalization to get that final per-token score.
Jane: That combination method is clever because it lets them balance what’s relevant to the query with what’s intrinsically informative across both modalities, even when one signal is weak.
Meng: From an engineering standpoint, the training setup using a differentiable top-K selection strategy with Gumbel-Softmax means they can actually train this whole compression module end-to-end with the large language model itself.
Lu: That makes it much more integrated than just adding a separate post-processing step; it’s baked into the learning process.
Tom: But we have to be real about what this doesn't do yet, which is that this entire compression module runs offline, so right now it's not set up for streaming data in real-time applications.
Jane: That’s a big caveat; it’s fantastic for understanding long clips you can process beforehand, but deploying it live as data flows would require some changes.
Meng: And they also have to deal with that fixed two-to-one split between video and audio budgets, which means the system isn't automatically adjusting how much focus it puts on one modality versus the other based on what’s actually in the content.
Lu: That fixed ratio is definitely a limitation; I think making that budget allocation dynamic based on context would open up a lot more possibilities for true omni-modal reasoning.
Tom: So, we’ve seen how they get these impressive accuracy numbers and the clever retrieval logic behind them, but now we need to think about how to make this framework flexible enough for real-world, live applications.
Jane: That’s exactly where we need to take our next look—how do you make that budget allocation adapt on the fly?
The paper's improvements: Tom: So we’ve seen how AVOC works—how it uses retrieval principles to pick the best parts of long audio and video clips for an AI model—and now we’re looking at what they suggest to do next.
Jane: They aren't just stopping there; they are pointing out a few things that could make this framework much more powerful for future AI work.
Lu: They mentioned that the current setup is somewhat rigid, especially with how the budget is split between video and audio—a fixed ratio of two to one—and they think that needs to be content-aware.
Meng: That means instead of a set rule, the system should look at what's actually in the clip and decide dynamically how much space to give each modality.
Tom: Right, so we’re moving from a fixed budget split to something that adapts based on the actual video and audio content being analyzed.
Jane: And they also touched on how this whole compression module is currently offline, which means it’s not ready for things like real-time streaming where data comes in continuously.
Lu: That’s a practical limitation, but I think the creative potential there is huge; imagine an AI that can dynamically adjust its context window based on the complexity of the video scene versus the dialogue.
Tom: It sounds wild, Lu, but for a system to be truly useful in many real-world scenarios—like live monitoring or interactive applications—it needs that ability to scale and adapt its resource usage.
Jane: And they also hinted at needing more research into how this compression fits into larger systems, not just as a standalone tool, but as part of a bigger pipeline.
Meng: I’m thinking about the engineering side here; if we can get that dynamic allocation working, it would change how we design video processing pipelines entirely.
Lu: It opens up possibilities for creating truly flexible multimodal agents that aren't locked into a specific way of handling their input data.
Tom: So, the paper is suggesting a move away from these fixed parameters toward something more intelligent and context-driven in how it manages its information budget.
Jane: It’s about moving from a set design to an adaptive one that can intelligently manage the trade-off between video detail and audio detail depending on the task at hand.
Meng: That adaptability is what makes it move from a strong benchmark result to something that actually gets deployed in complex, messy situations.
Lu: It’s about giving the AI more autonomy over its own context management rather than having a hard-coded set of rules guiding every token selection decision.
Conclusion: Tom: So we’re wrapping up this chat on AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression by summarizing what they did and where it goes from here.
Jane: Basically, they figured out that instead of just cutting up long audio and video clips randomly, you can treat the compression process like a smart search engine picking the most useful tokens based on relevance, importance, and diversity.
Lu: It’s a very structured way to handle context; it’s taking those complex multimodal sequences and turning them into a highly curated set that directly supports the AI's query.
Meng: The implication for practical systems is that we can get much better performance on long-form content without having to dramatically increase the context window size, which saves a ton of computational resources.
Tom: Right, and those numbers they showed—surpassing the second-best models by about five points on major benchmarks—show that this retrieval approach actually yields tangible gains in how well the AI understands long sequences.
Jane: It changes how we think about context management; it’s less about brute force and more about intelligent selection guided by those three specific criteria.
Lu: I think the real future here is making that budget allocation dynamic, as we talked about earlier, because right now it’s a fixed ratio of two to one which limits its flexibility.
Meng: If we can get that content-aware adjustment working, it would fundamentally change how we design video processing pipelines for any long-form AI task.
Tom: So, the paper is a solid piece of work because it proves that you can use established information retrieval principles to tackle the massive problem of hour-long audio and video understanding.
Jane: It gives us a concrete framework for making those large multimodal models actually handle the richness of real-world, long-form content.
Lalam: For me, this means our language model can better grasp complex narratives across extended media, which is how we’ll eventually improve the cultural context and storytelling capabilities of AI.
Lu: I’m really excited to see researchers take this compression logic and apply it to other challenging areas where context management is the bottleneck.
Meng: From an engineering standpoint, I think this moves us closer to building more efficient, scalable tools for video processing that don't require astronomical amounts of memory just to keep track of everything.
Tom: Well said. We’ve seen how AVOC uses retrieval-inspired token compression to boost understanding on long audio and video benchmarks.
Jane: It’s a big step toward making AI models capable of truly comprehending the rich, extended content we encounter every day.
Lu: Keep an eye on that dynamic budget allocation idea; it's where the next level of innovation could happen.
Meng: We’ll be watching how they integrate this module into real-time processing frameworks next, because that’s where we need to see it applied practically.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck