OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

arXiv:2609.39490 · cs.CV, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning".

Tom: As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding the paper "OmniReasoning:

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Well team, we're diving into this paper today. It's called "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning," and it looks like they’ve really zeroed in on how AI can actually handle understanding videos when you have sound involved.

Jane: That sounds intense, Tom; what’s the main idea behind this research?

Lu: Basically, the authors are addressing a gap where current models treat audio and vision separately instead of reasoning with them together.

Meng: So they're trying to build something that truly understands when you have both sound and sight at the same time.

Lalam: I'm really curious how this moves beyond just recognizing things in a picture or hearing sounds alone.

Tom: Exactly, Jane; they claim their work focuses on genuine audio-visual joint reasoning, which is a big step forward because most existing models handle those modalities independently.

Jane: So the core thesis seems to be that we need methods that explicitly model how audio and visual information interact during the reasoning process.

Lu: They set up a whole system to test this, starting with a dedicated benchmark called OmniReasoningBench, which demands joint audio and visual inputs for success.

Meng: A benchmark sounds like the first big hurdle; how rigorous are these tasks that require both modalities?

Tom: It’s quite rigorous; they say it tests seven hundred fifty reasoning steps over video content and also includes four hundred reasoning questions that go beyond what’s actually in the video itself.

Jane: That means we aren't just checking if the model can describe the scene visually or transcribe what it hears; they're testing complex understanding.

Lalam: From my perspective, this is huge because when an AI can reason across those different sensory inputs simultaneously, its ability to interact with the real world gets much richer.

Tom: And they support this benchmark with a data engine called OmniQA, which automatically builds evidence-grounded question pairs by explicitly linking audio and visual observations.

Lu: That’s clever because OmniQA generates training data like OmniReasoning-SFT-112K and OmniReasoning-RL-19K that includes timestamped clues and dependency chains.

Meng: So they aren't just feeding the model random stuff; they are giving it specific instructions on *why* certain visual or audio clues are important together.

Paper summary: Jane: That explicit supervision is what seems to be the backbone of their approach, guiding the learning process toward that joint understanding.

Lu: And then they propose ModalityFactored Self-Distillation, which is a novel method for training where it factors token-level credit assignment based on modality-specific evidence.

Tom: So instead of just judging a response overall, this method looks at whether the model's tokens are supported by audio clues, visual clues, or the joint audio-visual clues separately.

Lalam: That sounds like a really sophisticated way to reward interactions; it’s not just about getting the right answer overall, but understanding *how* the different pieces of evidence connect during that reasoning path.

Jane: It seems they are disentangling which modality contributed what to the final token prediction, which is a very granular form of feedback.

Meng: From an engineering standpoint, having this level of detail in how we train the model is something we’ve been chasing for a long time when dealing with complex inputs.

Tom: They claim that by using OmniQA and MFSD, their final model, OmniReasoning30B-A3B, achieved fifty point zero percent accuracy on OmniVideoBench and forty-two point five percent on OmniReasoningBench.

Lu: Those results show measurable improvements over the base model Qwen3-Omni-30B-A3BThinking, with gains of twelve point eight and nine point three percentage points respectively.

Jane: Those specific numbers give us a concrete idea of how much better this joint reasoning is compared to the previous versions we’ve seen in the research.

Meng: Improvements like those, especially on benchmarks that test long-video understanding like LVOmniBench, suggest that this isn't just a minor tweak; it shows the methodology actually helps with more complex tasks.

Lalam: If this technique can reliably improve how AI reasons across different sensory streams, I think it could significantly enhance how we build truly intuitive AI systems for daily life.

Tom: It really points to the fact that modeling these cross-modal evidence dependencies directly leads to better performance across the board, even on those long-video tasks.

Jane: So, the core message of "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning" is that we need more specific training signals to get models to truly connect audio and vision together.

Paper summary: Lu: The implication for future research is clear; this work sets a solid step for facilitating further research into omni-modal joint reasoning structures.

Meng: I think the practical impact here is in building AI that can handle real-world scenarios where context relies heavily on both what you see and what you hear simultaneously.

Lalam: For culture, if we can develop models that reason this deeply across modalities, it could lead to much more nuanced and empathetic interactions in applications, making the AI feel less like a tool and more like a collaborator.

Tom: So we’ve seen how they built the framework and what kind of performance they got with OmniReasoning30B-A3B, which opens up a lot of interesting avenues for future exploration.

Jane: It really highlights that the way we supervise these models matters profoundly when the task demands true cross-modal understanding.

Lu: We should keep an eye on how this factorization idea plays out in other reasoning tasks that involve multiple data types, because this approach seems adaptable.

Meng: I'll be watching how they translate this methodology into more practical deployment scenarios next, because getting these complex models running efficiently is where the real challenge lies.

Lalam: And from a cultural standpoint, I think the ability to handle these integrated reasoning tasks will push us toward building AI that can genuinely understand context in a richer way.

Tom: Absolutely; this paper gives us a clear path on how to structure our data and our learning algorithms to force these models into that joint reasoning space.

Jane: It’s exciting because it shows that targeted supervision based on modality-specific evidence is highly effective for complex multimodal tasks like those in "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning."

Lu: The authors laid out a solid foundation, providing both the benchmark and the learning mechanism needed to push those limits.

Meng: I think we’ll be looking closely at how they plan to extend OmniQA beyond these initial reasoning tasks in future work.

Lalam: It feels like this research is pushing the boundary of what we can expect from AI systems in terms of integrated perception and comprehension.

Conclusion: Tom: So we’ve been deep in the weeds with this paper on OmniReasoning, and now it’s time to zoom out a bit to talk about what this actually means for us all.

Jane: It really boils down to how these researchers tackled the challenge of getting AI to truly connect what it sees with what it hears simultaneously.

Lu: The title itself, "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning," tells you exactly that they’re aiming at a level of understanding we haven't really hit before.

Meng: I see why they focused on that joint reasoning aspect; from an engineering standpoint, forcing modalities to work together is where the real complexity usually hides in these systems.

Lalam: It suggests that if we can get AI to process the world through both sight and sound at once, the level of context awareness it has could really evolve beyond what we see now.

Tom: Exactly, Lalam; they’re not just making models that recognize a picture *and* a sound; they're building systems that reason about them as one cohesive experience.

Jane: And the authors achieved this by creating a specific testing environment and then designing a unique way for the AI to learn those connections directly from evidence.

Lu: Their methodology with OmniQA and ModalityFactored Self-Distillation shows a really smart way to provide targeted supervision precisely where the cross-modal gaps usually show up.

Meng: It’s interesting how they used that factorization approach; it sounds like a very surgical way to teach the model about those specific token interactions without just flooding it with general data.

Lalam: I think the real impact here is in how we build more intuitive interfaces and applications, where context isn't just visual but also auditory, which could lead to incredibly nuanced interactions.

Tom: Right, Lalam; it’s about moving past single-modality performance and building AI that operates with a richer sense of the world.

Jane: So when we look at the authors and their work on OmniReasoning, we see a clear effort to systematically bridge those sensory gaps in modern AI development.

Lu: Their approach sets a strong foundation for future research into how different types of evidence can be factored into a single reasoning process across various domains.

Meng: We'll need to see how practical these learned dependencies are when we deploy these models in real-world startup scenarios, but the theoretical grounding they provided is solid. **(Sound of upbeat, engaging radio outro music swells)**

Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei

School of Intelligence Science and Technology, Peking University

cs.CV, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/ByteDance-Seed/Seed2.0

Project page: https://pku-value-lab.github.io/OmniReasoning-Homepage

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding the paper "OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning." The goal is to

Key concepts

OmniReasoningBench
A specialized benchmark designed to require models to use both audio and video inputs simultaneously for successful reasoning. It tests complex tasks involving 750 steps over video content and 400 questions requiring knowledge beyond the visual scope, ensuring true joint understanding.
OmniQA
An automated data engine that generates high-quality training data by automatically linking audio and visual observations. It creates supervision signals by adding timestamped clues and dependency chains to explicitly show how audio and video evidence relate to each other.
ModalityFactored Self-Distillation (MFSD)
A novel learning method that factors credit assignment based on modality-specific evidence. It disentangles the contribution of different clues while explicitly modeling the cross-modal interactions between audio and visual tokens to reward correct dependencies during training.

Terminology

Summary

As a fastidious and diligent AI researcher, I have meticulously analyzed both provided texts regarding the paper OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning. The goal is to synthesize these two sources into a single, comprehensive, and highly detailed summary that accurately reflects the paper's contributions, methodology, and results.

Here is the combined summary:


The research presented in OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning addresses a critical gap in current multimodal AI by focusing on genuine audio-visual joint reasoning. Existing models often treat modalities—audio, vision, and language—independently, leading to insufficient evaluation and poor elicitation of complex cross-modal reasoning capabilities. This paper tackles this limitation through a tripartite approach: introducing a rigorous benchmark, developing an evidence-grounded data engine for supervision, and proposing a novel learning method that explicitly models modality interactions.

The work is structured around three interconnected components designed to advance audio-visual joint reasoning:

1. OmniReasoningBench (The Benchmark):

A primary contribution is the creation of OmniReasoningBench, a dedicated benchmark specifically designed to necessitate joint audio and visual inputs for successful reasoning. This benchmark moves beyond single-modality evaluation, demanding that models demonstrate genuine cross-modal understanding. It is structured to rigorously test complex reasoning tasks, encompassing both 750 reasoning steps over video content and 400 reasoning questions that extend beyond the scope of the video itself.

2. OmniQA (The Data Engine):

To provide high-quality, explicit supervision for joint reasoning, the authors developed OmniQA, an automated data engine. OmniQA's function is to automatically construct evidence-grounded question pairs by explicitly linking audio and visual observations. Crucially, this engine generates training data—such as OmniReasoning-SFT-112K and OmniReasoning-RL-19K—that includes timestamped clues and dependency chains, providing explicit supervision signals for cross-modal training. This systematic construction ensures that the resulting data directly targets the required joint reasoning capabilities.

3. ModalityFactored Self-Distillation (MFSD) (The Learning Method):

The learning mechanism proposed is ModalityFactored Self-Distillation (MFSD), an on-policy self-distillation method tailored for this cross-modal task. MFSD's core innovation lies in its ability to factor token-level credit assignment according to modality-specific evidence. This disentangles the contributions of individual clues and, more importantly, explicitly models the cross-modal interactions between audio and visual evidence at the token level. By utilizing a combination of joint-clue support and centered interaction, MFSD derives fine-grained token-level advantage weights that reward correct cross-modal dependencies. This method is shown to outperform existing reinforcement learning approaches (like GRPO and RLSD) by explicitly rewarding these crucial cross-modality interactions.

The synergy between these components forms the core of the research: OmniQA constructs both the questions for OmniReasoningBench and the training corpora, while its modality-specific clue annotations serve as the supervision signal for MFSD. MFSD then evaluates sampled responses under modality-specific clue contexts to derive fine-grained supervision. The overall process is designed to explicitly model and optimize for these cross-modal evidence dependencies.

The paper details the mechanics of OmniReasoningBench, the construction protocols for OmniQA, and the precise mathematical formulation of MFSD. The experimental setup includes training with the resulting data (OmniQA) and fine-tuning with MFSD to produce the final model, OmniReasoning30B-A3B.

The results demonstrate that explicitly constructing, annotating, and optimizing for cross-modal evidence dependencies yields substantial performance gains. The final model, OmniReasoning30B-A3B, achieved significant improvements:

  • 50.0% accuracy on OmniVideoBench.

  • 42.5% accuracy on OmniReasoningBench.

These results represent notable improvements over the baseline model (Qwen3-Omni-30B-A3BThinking), showing gains of 12.8 and 9.3 percentage points, respectively. Furthermore, the work shows that this methodology enhances general and long-video understanding benchmarks, including LVOmniBench, Video-MMMU, and VideoMME-v2.

In conclusion, the core finding is that advancing audio-visual joint reasoning requires explicitly modeling the dependencies between audio and visual evidence.

Improvements for AI systems

Here are specific improvements that can be made to existing AI systems by leveraging the concepts, data, and methodology presented in this scientific paper:


  1. The development of a unified benchmark for audio-visual joint reasoning (OmniReasoningBench). This system will enable researchers to rigorously test whether models truly possess cross-modal understanding rather than relying on single-modality shortcuts.

  2. The creation of an evidence-grounded data engine (OmniQA) that automatically constructs complex QA pairs with timestamped clues and dependency chains linking audio and visual events. This allows for the creation of high-quality, structured training data that explicitly teaches models how to connect disparate pieces of information across modalities.

  3. The implementation of Modality-Factored Self-Distillation (MFSD) as a reinforcement learning method. This method will provide fine-grained, token-level credit assignment by disentangling the contributions of individual modalities and their non-additive interactions, ensuring that models are rewarded for generating reasoning steps that genuinely rely on complementary evidence rather than exploiting one modality in isolation.

  4. The application of MFSD to large language models (e.g., OmniReasoning30B-A3B). This will result in models with substantially improved performance on complex audio-visual tasks, including long video understanding and general video reasoning, by explicitly optimizing for cross-modality interaction during the reinforcement learning phase.

  5. The creation of modality ablation studies (as demonstrated in Section A.4) to systematically isolate the impact of audio versus visual inputs. This capability will allow developers to precisely quantify the marginal benefit of combining modalities, guiding future architectural decisions for more efficient omni-modal architectures.

  6. The establishment of a pipeline for reasoning beyond video (Reasoning beyond Video). This will allow models to apply knowledge learned from a video context (e.g., physical rules or procedural steps) to novel, unseen scenarios or new figures, extending the utility of video understanding far beyond simple observation within the original footage.

  7. The integration of structured clue annotations into training data and RL processes. By retaining modality-specific clue annotations, the system can provide targeted supervision that goes beyond simple outcome-level guidance (like GRPO), leading to more robust and generalizable reasoning capabilities across diverse multimodal inputs.

Abstract

Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-visual joint reasoning, together with time-stamped clue chains that guide the annotation of thinking process. Besides our benchmark, this engine produces training data OmniReasoning-SFT-112K and OmniReasoning-RL-19K. Finally, we propose an on-policy self-distillation method Modality-Factored Self-Distillation (MFSD). It evaluates each sampled response under modality-specific clue contexts, disentangling the contributions of individual clues and their cross-modal interactions for token-level credit assignment. With our training data and learning method, our model OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points, respectively. Moreover, it delivers substantial gains on general and long-video benchmarks, including Video-MME-v2. We hope our work offers a solid step for facilitating future research in omni-modal joint reasoning.

Sources

Related papers