Realtime-Venus: A full-duplex interaction system with asynchronous delegation

arXiv:2609.13814 · cs.CV, eess.AS · Submitted 2026-09-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Realtime-Venus: A full-duplex interaction system with asynchronous delegation".

Jane: The paper was written by Venus Team and Ant Group from Ant Group.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper discussion segment 1: Tom: Moving on from the initial concept, let's look at what the actual core summary of "Realtime-Venus: A full-duplex interaction system with asynchronous delegation" tells us about how these systems actually function in everyday life.

Jane: Basically, this paper details a system engineered to make AI conversations feel like a real two-way street. Instead of the old way where the AI has to wait for you to stop speaking before it can reply, this system lets it listen and talk at the same time, which is what they call full-duplex interaction.

Lu: That simultaneous operation is really powerful because they’ve built two separate models—one dedicated to audio-visual interactions and another focused purely on spoken interaction—but they tie them together using a shared causal timeline. It’s a clever way to manage that complexity without everything running into a mess of conflicting signals.

Meng: I'm thinking about the practical side again when you talk about asynchronous delegation; are we talking about the AI pausing its main conversation just to run a big calculation in the background so it doesn't get stuck waiting for results? I’m curious how that practical setup looks for a user in action.

Lalam: It shouldn't feel like a stutter or a sudden stop because of this; that’s where this system gets its real power. If the AI can handle a complicated reasoning task quietly in the background while you’re talking, it means it can keep our main dialogue perfectly on track and only jump in when we actually need to respond to what we just said.

Tom: Exactly! It shifts the focus away from simple question-and-answer back and forth toward genuine collaboration where the AI is actively managing multiple streams of activity at once, instead of just waiting for its turn in line. We have to figure out exactly how they built this technical setup to make it work.

Jane: And what’s really brilliant is that they introduced this concept of asynchronous delegation so effectively. This means if a request comes in that needs external knowledge, like checking live traffic, the system can immediately hand that task off to a separate worker so it doesn't have to pause the main dialogue loop at all.

Lu: That separation between immediate control over what we say and background execution of tasks is what allows for such robust multitasking; it’s about managing multiple threads of activity concurrently in a way that keeps the front end feeling snappy. It’s really elegant in its design structure when you look at how they organize the timing.

Meng: From an engineering standpoint, I still wonder how they managed that handover efficiently across those two different model types—the Omni and Audio variants—without introducing unacceptable delays when switching between perception tasks and delegation requests. That transition speed is definitely what makes or breaks the user experience.

Lalam: They solved it by using a shared framework called Realtime-Venus-Harness; it’s designed to bind those tasks to the evidence available right at the moment of request, so the AI knows exactly what context to send off for that external job. It keeps things clean and predictable for the user experience.

Tom: So, we're talking about a system where you can have a deep, continuous dialogue going on while complicated work is happening silently in the background, and that handoff between those two worlds is completely seamless for the end user.

Jane: Right, and they showed this works incredibly well even when things get messy—like when someone jumps in with a correction or throws out a new request mid-sentence. They proved that the conversation stays smooth because of this unified structure, even when things get chaotic.

Lu: That ability to stop a stale response and immediately incorporate the new intent before continuing is huge; it shows true adaptability in real-time dialogue systems that moves way beyond just managing simple turn-taking flow.

Meng: I was looking at their performance metrics on tool use, and they reported an eighty-six point zero percent tool selection F1 score for Realtime-Venus-Omni; that’s really strong evidence about how well the AI decides *when* to delegate a task to a tool versus just answering based on its own knowledge. That metric speaks volumes about their internal decision-making process itself.

Lalam: And I think that success in delegation is what truly matters for how we see the future of AI culture; if the AI can reliably know when it needs external help versus when it can handle things internally, then we get assistants that feel genuinely capable of supporting complex human endeavors. It validates the idea of an assistant that knows its limits.

Tom: It sounds like the main message here is that this system doesn't just reply; it’s proactive and manages its own cognitive load by smartly deciding what to do next—whether that’s talking back or handing off a heavy lift when necessary.

Jane: And they demonstrated great results in conversational continuity, especially when handling tricky things like user backchannels or even interruptions, which means the conversation stays smooth even when things get chaotic. It keeps the overall user experience high because it doesn't break under pressure at all.

Lu: The way they structured the unified stream serialization across three different synchronized streams—user input, assistant output, and background work—is incredibly elegant and sets a really high bar for future multimodal research in this field. That’s a very sophisticated architectural choice we should all be paying attention to as we build things.

Meng: I was thinking about scaling this up: if we move to even more complex tasks that need specialized hardware, how do we make sure that asynchronous delegation stays efficient under really heavy load? That's definitely the practical test they’ll face when deploying something like this in a real-world setting. We need really solid scaling solutions for this kind of complexity.

Lalam: Even with those technical hurdles, the big implication is that we can start seeing assistants that function as true collaborative partners rather than just sequential responders; they're becoming genuine teammates who can handle both the conversation and the heavy lifting together.

Paper discussion segment 2: Tom: Now that we've covered the initial setup, let's pivot to what the research is actually saying about how this system functions in practice. What does this paper tell us about its practical application beyond just a theoretical concept?

Jane: Well, at its core, this paper describes a system engineered to make AI conversations feel like a real two-way street. The main idea is that instead of the old way where the AI has to wait for you to finish speaking before it can reply, this system lets it listen and talk at the same time through full-duplex interaction.

Lu: That simultaneous operation is really powerful because they’ve built two separate models—one dedicated to audio-visual interactions and another focused purely on spoken interaction—but they tie them together using a shared causal timeline. It’s a clever way to manage that complexity without everything running into a mess of conflicting signals.

Meng: I'm thinking about the practical side again when you talk about asynchronous delegation; are we talking about the AI pausing its main conversation just to run a big calculation in the background so it doesn't get stuck waiting for results? I’m curious how that practical setup looks for a user in action.

Lalam: It shouldn't feel like a stutter or an abrupt stop because of this; that’s where this system gets its real power. If the AI can handle a complicated reasoning task quietly in the background while we talk, it means it can keep our main dialogue perfectly on track and only jump in when we actually need to respond to what we just said.

Tom: Exactly! It shifts the focus from simple question-and-answer back and forth toward genuine collaboration where the AI is actively managing multiple streams of activity at once, instead of just waiting for its turn in line. We have to figure out exactly how they built this technical setup to make it all function correctly.

Jane: And what’s really brilliant is that they introduced asynchronous delegation so effectively; this means if a request comes in that needs external knowledge, like checking live traffic, the system can immediately hand that task off to a separate worker so it doesn't have to pause the main dialogue loop at all.

Lu: That separation between immediate control over what we say and background execution of tasks is what allows for such robust multitasking; it’s about managing multiple threads of activity concurrently in a way that keeps the front end feeling snappy. It’s really elegant in its design structure when you look at how they organize the timing.

Meng: From an engineering standpoint, I still wonder how they managed that handover efficiently across those two different model types—the Omni and Audio variants—without introducing unacceptable delays when switching between perception tasks and delegation requests. That transition speed is definitely what makes or breaks the user experience here.

Lalam: They solved it by using a shared framework called Realtime-Venus-Harness; it’s designed to bind those tasks to the evidence available right at that moment of request, so the AI knows exactly what context to send off for that external job. It keeps things clean and predictable for the user experience.

Tom: So, we're talking about a system where you can have a deep, continuous dialogue going on while complicated work is happening silently in the background, and that handoff between those two worlds is completely seamless for the end user experience to see.

Jane: Right, and they showed this works incredibly well even when things get messy—like when someone jumps in with a correction or throws out a new request mid-sentence. They proved that the conversation stays smooth because of this unified structure, even when things get chaotic. It really shows the robustness of the design.

Lu: That ability to stop a stale response and immediately incorporate the new intent before continuing is huge; it shows true adaptability in real-time dialogue systems that moves way beyond just managing simple turn-taking flow.

Meng: I was looking at their performance metrics on tool use, and they reported an eighty-six point zero percent tool selection F1 score for Realtime-Venus-Omni; that’s really strong evidence about how well the AI decides *when* to delegate a task to a tool versus just answering based on its own knowledge. That metric speaks volumes about their internal decision-making process itself.

Lalam: And I think that success in delegation is what truly matters for how we see the future of AI culture; if the AI can reliably know when it needs external help versus when it can handle things internally, then we get assistants that feel genuinely capable of supporting complex human endeavors. It validates the idea of an assistant that knows its limits.

Tom: It sounds like the main message here is that this system doesn't just reply; it’s proactive and manages its own mental load by smartly deciding what to do next—whether that’s talking back or handing off a heavy lift when necessary.

Jane: And they demonstrated great results in conversational continuity, especially when handling tricky things like user backchannels or even interruptions, which means the conversation stays smooth even when things get chaotic. It keeps the overall user experience high because it doesn't break under pressure at all.

Lu: The way they structured the unified stream serialization across three different synchronized streams—user input, assistant output, and background work—is incredibly elegant and sets a really high bar for future multimodal research in this field. That’s a very sophisticated architectural choice we should all be paying attention to as we build things.

Meng: I was thinking about scaling this up: if we move to even more complex tasks that need specialized hardware, how do we make sure that asynchronous delegation stays efficient under really heavy load? That's definitely the practical test they’ll face when deploying something like this in a real-world setting. We need really solid scaling solutions for this kind of complexity.

Lalam: Even with those technical hurdles, the big implication is that we can start seeing assistants that function as true collaborative partners rather than just sequential responders; they're becoming genuine teammates who can handle both the conversation and the heavy lifting together.

Paper discussion segment 3: Tom: Moving on from what we know about its core mechanics, let’s pivot to what this research is suggesting for taking this technology even further with "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Where are they pointing next for future development?

Jane: They are focusing heavily on expanding how long these AI can maintain a meaningful conversation, especially by tackling long-term video understanding with a training-free memory module that uses motion compensation to keep track of things across hours.

Lu: That memory augmentation is absolutely game changer for context retention; being able to retrieve relevant visual information from an hour ago means the AI isn't constantly forgetting the beginning of a long discussion, which solves a huge problem in current systems. It fundamentally changes how they model temporal dependencies in these models.

Meng: From my viewpoint, that’s fantastic for practical applications where we need assistants to remember context over very long sessions, like monitoring a complex process or reviewing hours of footage; I just want to see how much computational overhead that memory retrieval adds to the real-time interaction loop.

Lalam: It’s about enabling deeper cultural understanding; if the AI can recall nuanced visual context from an hour ago, it moves from being a reactive chatbot to something that can truly grasp the *history* of a shared experience, which is vital for sophisticated digital assistants. That depth of memory makes it feel like a real partner.

Tom: And they're also pushing for better control over the streaming chunks themselves and extending the overall context window so we can have these incredibly nuanced conversations without losing track of earlier details throughout the entire duration. They aren't just making it last; they’re making sure the conversation remains precise throughout its entire length.

Jane: So, the focus isn't just on keeping a conversation going; it’s on making that interaction deeply informed by a much longer memory and more precise control over every little speech decision they make. It's about quality over simple duration in terms of how deep the understanding goes.

Lu: That really points toward a new way of modeling temporal dependencies in AI systems—moving beyond short-term context windows to truly persistent understanding. It’s pushing the boundaries of what we thought was possible for conversational memory within these models.

Meng: I’m looking forward to seeing how those fine-grained controls work in production environments; the challenge will be ensuring that this enhanced memory doesn't slow down the actual real-time response generation when things get busy. Speed is still a big concern for me.

Lalam: Even with those technical hurdles, the ultimate implication is that we can start seeing assistants that are truly collaborative partners rather than just sequential responders, capable of remembering and reflecting on our entire interaction history. That level of memory makes the partnership feel much more meaningful to users.

Tom: So, to sum up these improvements on "Realtime-Venus: A full-duplex interaction system with asynchronous delegation," it’s moving toward an AI that’s not only responsive but also deeply knowledgeable across extended timeframes and incredibly precise about its interactions.

Jane: It’s a huge leap forward in how we design conversational flow, proving that true real time engagement is achievable even with the most complicated reasoning tasks running alongside it. We’re moving past just waiting for our turn now.

Lu: The methodology they used for training that couples proactive duplex data with delegation scenarios is what truly separates this from standard models; it builds agents that understand the *flow* of human intent better, which is a massive methodological win in how we train these systems.

Meng: It’s a great piece of work, and I'm eager to see how those asynchronous patterns translate into more efficient, production-ready systems in the near future; we need those practical deployment details to see if it hits the ground running efficiently.

Lalam: I think this paper sets a new standard for what we expect from advanced AI interactions—a standard that prioritizes deep engagement and reliable background support. It shows us the direction we should be heading for assistants that are truly collaborative partners, not just sequential responders.

Tom: We’ve got some seriously exciting findings today on "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Next time, we’ll be looking at how these models handle those tricky interruptions we talked about earlier.

Conclusion: Tom: Alright everyone, let's wrap up our deep dive into "Realtime-Venus: A full-duplex interaction system with asynchronous delegation," and what a significant piece of research this is for building interactive AI.

Jane: It really shows us a new way of thinking about how we design these assistants—moving away from simple turn-taking to something that's genuinely proactive and capable of managing multiple streams at once.

Lu: I think the core innovation here is that they’ve managed to unify perception, control, and delegation onto a single causal timeline; it’s like they built a master clock for the entire interaction session.

Meng: From what I've seen in the implementation details, managing that dual-loop runtime between the interaction loop and the capability loop sounds like a huge engineering feat to get running smoothly in practice.

Lalam: For me, it's amazing because this capability means we can build assistants that aren't just reactive tools but can maintain deep cultural context over long conversations while simultaneously working on complex tasks for us.

Tom: That’s the big picture, Lalam—an AI that’s truly engaged and capable of supporting our long-term needs without ever dropping the ball during a heavy computation.

Jane: It makes me feel really optimistic about what this could mean for everyday interaction; imagine an assistant that can listen to a meeting while simultaneously pulling up relevant data from several external sources in the background. That’s a huge shift in expectation for us all.

Lu: The way they structured the unified stream serialization across three synchronized streams—user, assistant, and background—is incredibly elegant and sets a high bar for future multimodal research. It’s a very sophisticated architectural choice that we should all be paying attention to moving forward.

Meng: I wonder how scalable this is when we move to even more complex tasks requiring specialized hardware; ensuring that asynchronous delegation remains efficient under heavy load is definitely the practical test they’ll face in real-world deployment scenarios. We need robust scaling solutions for this kind of complexity.

Lalam: Even with those technical hurdles, the ultimate implication is that we can start seeing assistants that are truly collaborative partners rather than just sequential responders; they're becoming true teammates who can handle both the conversation and heavy lifting together.

Tom: So, to sum up "Realtime-Venus: A full-duplex interaction system with asynchronous delegation" gives us a powerful blueprint for building AI agents that are both incredibly responsive and deeply capable of handling complex, multi-faceted demands simultaneously.

Jane: It’s a huge leap forward in how we design conversational flow, proving that true real time engagement is achievable even with the most complicated reasoning tasks running alongside it.

Lu: The methodology they used for training that couples proactive duplex data with delegation scenarios is what truly separates this from standard models; it builds agents that understand the *flow* of human intent better, which is a massive methodological win for future AI development.

Meng: It’s a great piece of work, and I'm eager to see how those asynchronous patterns translate into more efficient, production-ready systems in the near future; we need those practical deployment details to see if it hits the ground running efficiently.

Lalam: I think this paper sets a new standard for what we expect from advanced AI interactions—a standard that prioritizes deep engagement and reliable background support. It shows us the direction we should be heading for assistants that are truly collaborative partners, not just sequential responders.

Tom: We’ve got some seriously exciting findings today on "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Next time, we’ll be looking at how these models handle those tricky interruptions we talked about earlier.

Ant Group

cs.CV, eess.AS

Submitted: 2026-09-12

Updated: 2026-09-23

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 91/100

The gist: Realtime-Venus is a proactive full-duplex interaction system that combines native conversational modeling with asynchronous delegation, built on MiniCPM-o 4.5 (Cui et al., 2026).

Key concepts

Full-duplex interaction
This system allows an AI to listen and talk at the same time. Instead of waiting for a user to finish speaking before replying, it enables simultaneous operation, making conversations feel like a real two-way street.
Asynchronous delegation
This feature lets the AI pause its main conversation to run big calculations or handle external knowledge requests in the background. This prevents the AI from getting stuck waiting for results while keeping the main dialogue perfectly on track.
Realtime-Venus-Harness
This is a shared framework used by Realtime-Venus to bind tasks to available evidence at the moment a request is made. It helps the AI know exactly what context to send off for external jobs, ensuring clean and predictable user experiences.
Tool selection F1 score
This metric measures how well the AI decides when it needs to delegate a task to an external tool versus answering based on its own knowledge. A high score indicates strong internal decision-making about when to seek help.

Terminology

Summary

Realtime-Venus is a proactive full-duplex interaction system that combines native conversational modeling with asynchronous delegation, built on MiniCPM-o 4.5 (Cui et al., 2026). The system consists of two separately trained 9B models: Realtime-Venus-Omni for audio–visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events.

The system utilizes a dual-loop runtime: the Interaction loop (Realtime-Venus Frontend) continuously processes incoming media to update the session state and control when to listen or speak; the Capability loop (Realtime-Venus-Harness) handles private natural-language delegation requests asynchronously by executing tasks in the background and returning results for integration into the dialogue.

The unified streaming formulation aligns user observations, model outputs, private delegation requests, and background results on a shared causal timeline. The frontend jointly predicts interaction-control tokens, response text, and delegation requests. This shared policy supports maintaining a response during user backchannels (C RESUME), revising its unspoken continuation after a correction (C RESPOND), and initiating background work when a request requires external capabilities.

Realtime-Venus-Harness is a shared framework for asynchronous capability execution and result delivery. It binds tasks to evidence available at the request boundary, executes registered capabilities asynchronously, and returns results to the originating session while preserving frontend control over conversational responses. The framework tracks each work item through the lifecycle: Q UEUED → RUNNING → C OMPLETED → D ELIVERING → D ELIVERED.

The training requires trajectories that connect conversational events with interaction decisions and delegated execution. A unified data pipeline combines scenario planning, speech realization, and temporal alignment to construct these trajectories. Duplex scenarios distinguish backchannels and other-directed speech from interruptions that require stopping, repairing, or redirecting a response. Proactive trajectories supervise when to initiate a response and when to continue listening. Delegation scenarios connect private requests with background execution, returned information, and subsequent responses.

Realtime-Venus-Omni uses both audio–visual and audio-only data, whereas Realtime-Venus-Audio uses the audio-only subset. The models are trained on a common post-training corpus of over 2.8 million samples covering nine data categories: offline understanding (general AV understanding, general audio understanding, spoken question answering), proactive duplex interaction (visual-, multimodal-, speech-, and audio-driven interaction with response timing and interruption handling), and delegation (delegate–backend–restate workflow).

The architecture involves continuous perception, conversational control, and native speech generation within a single autoregressive interaction loop. Realtime-Venus-Omni encodes aligned visual and audio streams with SigLIP2 (Tschannen et al., 2025) and Whisper-Medium (Radford et al., 2023), while Realtime-Venus-Audio removes the ViT-based visual branch and processes only streaming audio.

The unified stream serialization involves three synchronized streams—the user stream, assistant stream, and background stream—on a shared conversational clock. The User stream is a causal sequence of time-aligned perceptual features (audio or audio/video). The Assistant stream combines foreground text, aligned S3 speech tokens (Du et al., 2024a,b), and optional text-only delegation instructions. The Background stream is an asynchronous text-only sequence connecting the models to a more capable backend agent for complex reasoning and tool-based tasks.

The model family inherits the Omni-Flow architecture of MiniCPM-o 4.5, where each one-second unit interleaves visual tokens from the current frame with temporally aligned audio features (for Realtime-Venus-Omni) or contains only audio features (for Realtime-Venus-Audio). At each unit, the language model predicts or to control perception and speech generation. The system distinguishes between pauses and background noise versus backchannels (C RESUME), interruptions that require stopping or revising the response (C RESPOND), and other-directed speech.

The training recipe uses a unified post-training recipe that mixes proactive duplex data, delegation data, and general understanding data. Supervision is sparse: loss is computed only on response spans, excluding system, user, and media-placeholder tokens. For full-duplex data, supervision covers the per-second / decision tokens together with the spoken text. The acoustic decoder remains fixed and is excluded from the training objective.

The evaluation assesses understanding (StreamingBench for video benchmarks; MMAU, MMSU for audio understanding), conversational continuity (Full-Duplex-Bench v1.5), and delegation decisions (internal delegate benchmark). Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, while Realtime-Venus-Audio achieves the highest scores in several audio understanding and spoken question answering comparisons. Full-duplex evaluations show high continuation rates under non-interruptive speech. Tool use evaluations show that Realtime-Venus-Omni achieves 86.0% tool selection F1, 53.1% argument accuracy, and 43.0% Pass@1, compared with 85.0%, 54.2%, and 48.0% for Realtime-Venus-Audio in Full-Duplex-Bench v3 (FDB-v3). The delegate benchmark shows that Realtime-Venus-Omni achieves an overall routing accuracy of 75.93%, exceeding Realtime-Venus-Audio by 7.04 percentage points, with the two variants exhibiting different strengths regarding delegation recall and non-delegation specificity.

The conclusion is that Realtime-Venus demonstrates competitive multimodal understanding and strong conversational continuity under non-interruptive speech, while its asynchronous design provides access to external capabilities while keeping interaction active. Memory augmentation further supports hour-scale video understanding, with improvements across all evaluated duration bins. The findings support coordinating immediate conversational responses with longer-running computation as a promising direction for assistants that remain engaged with users while handling tasks beyond the frontend’s own capabilities. Future work will explore finer-grained streaming chunks to better capture brief events and improve the timing of conversational responses, and extend the context window to support longer interactions.

Key contributions include:

(1) Proactive full-duplex interaction models:

"We develop Realtime-Venus-Omni and RealtimeVenus-Audio, two separately trained 9B models for proactive audio–visual interaction and full-duplex spoken dialogue, respectively. Both models integrate continuous perception, conversational control, native speech generation, and private delegation under a unified streaming formulation. To our knowledge, Realtime-Venus-Omni is the first full-duplex omni model to support asynchronous backend invocation for reasoning and tool execution while maintaining video interaction."

(2) Asynchronous capability execution with Realtime-Venus-Harness:

"We introduce a shared execution framework that binds tasks to evidence available at the request boundary, executes registered capabilities asynchronously, and returns results to the originating session while preserving frontend control over conversational responses."

(3) A coupled duplex and delegation data pipeline:

"We develop a pipeline combining scenario planning, speech realization, and temporal alignment to construct trajectories coupling conversational events with delegation requests, background results, and response continuations."

The system is designed to preserve stable context for background execution while allowing the frontend to adapt its response to subsequent changes in user intent. The example provided demonstrates the full delegate cycle:

The first response chunks acknowledge the spoken query and commit to looking it up; the span dispatches the task to the Harness and the turn closes with, so the frontend returns to listening while the backend executes.

The system architecture is summarized by Figure 3, illustrating a dual-loop runtime where both frontends share a delegation interface. The example interaction shows how a user query triggers delegation:

When the referee blows the whistle, please remind me.

Delegate (00:30) → Realtime-Venus-Omni executes the task and returns results in the background, while the frontend continues to listen or prepare subsequent responses.

The system's unified runtime abstraction defines session state as: "Let session σ be with frontend m ∈ [audio, omni], and let k index one-second chunks; session superscripts are omitted. The media inputs are xaudio = uk and xomni = (uk, vk), where uk contains causal audio k features and vk contains aligned visual features, with vk = ∅ when no frame is available." The output Ok = (ck, Yk, Dk, Sk) comprises interaction-control tokens, foreground text, delegation text, and speech tokens; only Yk and Sk are user-facing. The causal runtime transition is (sk+1, Ok) = Fm (sk, xm k, bk; ηk).

The system's long-term memory module is training-free and uses a visual memory gating mechanism based on motion-compensated prediction cost to archive relevant frames, and a retrieval mechanism inspired by the MaxSim operator to perform fine-grained matching between query tokens and visual tokens of stored historical frames. This allows the model to recover relevant historical information beyond its rolling context window without continuously retaining the complete input history.

The system's full-duplex control distinguishes between:

(1) Pauses and background noise:

Silence, hesitation, or acoustic activity not directed at the assistant produces, keeping perception continuously active without prematurely starting or terminating a response.

(2) Backchannels:

A short acknowledgment such as “yes” or “right” does not claim the conversational floor. The model preserves and continues the current response plan.

(3) Interruptions:

"When overlapping user speech takes the conversational floor with a correction, redirection, or new request, the model emits to terminate the stale response. It then incorporates the new intent into its conversational state before generating the next segment or delegate request, revising the unspoken continuation based on the new input while leaving already played audio unchanged."

The system's delegation target specifies computational handling: "Requests requiring additional perceptual evidence, broader context, or substantial replanning may invoke a registered capability; those involving external information or executable actions are routed to tools or specialized components. The event’s effect on user intent links both targets: A backchannel preserves speech and computation; other-directed and background speech create no new route. A stable-intent interruption may stop playback while retaining the semantic plan; an intent-changing correction cancels dependent work and triggers replanning."

The system's training data is organized into categories such as general audio understanding, spoken question answering, proactive duplex interaction, and delegation, with video data constituting approximately 70% of the corpus, consumed exclusively by Realtime-Venus-Omni. The training recipe mixes proactive duplex data, delegation data, and general understanding data.

The system's evaluation protocol assesses understanding (StreamingBench for video benchmarks; MMAU, MMSU for audio understanding), conversational continuity (Full-Duplex-Bench v1.5), and delegation decisions (internal delegate benchmark). Evaluation metrics include C RESPOND, C RESUME, UNCERTAIN, and UNKNOWN responses under Full-Duplex-Bench v1.5. Tool use evaluations in FDB-v3 measure tool selection F1, argument accuracy, and Pass@1. The delegate benchmark evaluates delegation accuracy (Overall), external capabilities (Delegation Recall), routine interaction (Non-delegation Specificity), and reasoning (Routing Accuracy).

The evaluation results show that Realtime-Venus-Omni achieves an overall routing accuracy of 75.93%, exceeding Realtime-Venus-Audio by 7.04 percentage points, with the two variants exhibiting different strengths regarding delegation recall and non-delegation specificity. The system's conclusion is that its asynchronous design provides access to external capabilities while keeping interaction active: background reasoning and tool execution proceed while the frontend continues receiving inputs and managing speech.

The example of a query is demonstrated in Figure 4:

What should each day look like take a break every two hours, plus lunch and an for breaks and meals?

The assistant response demonstrates full-duplex capability: Aim for 5–7 driving hours daily, Start after breakfast, drive two hours, then take a short break. Drive two more hours, stop for lunch, and continue one to three hours before dinner and overnight rest.

This example illustrates the system's ability to handle interruptions while maintaining conversational continuity during other overlapping speech. For instance:

User interruption (13s–15s) Audio Input (Realtime-Venus-Audio) Help me plan a three-day road trip with safe pacing.

"Assistant response (truncated) Aim for 5–7 driving hours daily, Start after breakfast, drive two hours, then take a short break. Drive two more hours, stop for lunch, and continue one to three hours before dinner and overnight rest."

The system's ability to handle delegation is demonstrated in Figure 4:

User: 请你帮我查询一下北京今天汽车限行尾号。 (Query external service)

Assistant response (truncated) 今天北京市限行尾号为 3 和 8,好的,我查询一下。 followed by the dispatch of a task to check the real-time traffic information.

The system's ability to handle complex, multi-step reasoning with tool use is demonstrated in Figure 4:

User: 请你帮我查询一下北京今天汽车限行尾号。 (Query external service)

Assistant response (truncated) 今天北京市限行尾号为 3 和 8,好的,我查询一下。 限行时间为 7:00–20:00. followed by the execution of a tool call to query external service.

The system's ability to handle spoken interaction and interruption is demonstrated in Figure 4:

User: Help me plan a three-day road trip with safe pacing.

"Assistant response (truncated) Aim for 5–7 driving hours daily, Start after breakfast, drive two hours, then take a short break. Drive two more hours, stop for lunch, and continue one to three hours before dinner and overnight rest."

The system's ability to handle interruption-continuation trade-off is demonstrated in Table 6: "Realtime-Venus-Audio achieves the highest continuation rates for user backchannels and background speech, and ranks second to Realtime-Venus-Omni for speech directed to others. However, its interruption-response rate is lower than those of JoyDuplex, GPT-4o, and Gemini 3.1 Live."

The system's ability to handle tool use is demonstrated in Table 7: Realtime-Venus-Omni achieves 86.0% tool selection F1, 53.1% argument accuracy, and 43.0% Pass@1, compared with 85.0%, 54.2%, and 48.0% for Realtime-Venus-Audio.

The system's ability to handle delegation decisions is demonstrated in Table 8: Realtime-Venus-Omni achieves an overall routing accuracy of 75.93%, exceeding Realtime-Venus-Audio by 7.04 percentage points. The benchmark identifies different delegation failure patterns: Realtime-Venus-Audio recognizes most requests requiring external capabilities but frequently delegates routine requests that should be handled locally.

The system's conclusion is that its asynchronous design provides access to external capabilities while keeping interaction active: background reasoning and tool execution proceed while the frontend continues receiving inputs and managing speech. The system's memory augmentation further supports hour-scale video understanding, with improvements across all evaluated duration bins. Together, these findings support coordinating immediate conversational responses with longer-running computation as a promising direction for assistants that remain engaged with users while handling tasks beyond the frontend’s own capabilities."

The example of a query is demonstrated in Figure 4:

User: What should each day look like take a break every two hours, plus lunch and an for breaks and meals?

The assistant response demonstrates full-duplex capability: Aim for 5–7 driving hours daily, Start after breakfast, drive two hours, then take a short break. Drive two more hours, stop for lunch, and continue one to three hours before dinner and overnight rest.

The system's ability to handle tool use is demonstrated in Table 7: "Realtime-Venus-Omni achieves 86.0% tool selection F1, 53.1% argument accuracy, and 4

Improvements for AI systems

Based on the provided paper, here are specific improvements for AI systems derived from the Realtime-Venus architecture, focusing on its core innovations:


  1. The implementation of a dual-loop runtime (Interaction Loop vs. Capability Loop) allows for high-fidelity real-time interaction without sacrificing complex reasoning.

  2. The system can perform concurrent perception (audio/visual) and background task execution (tool use/reasoning).


Improved AI System Capabilities:

  1. A full-duplex conversational agent that can simultaneously listen to a user, process visual context, and execute complex, external tool calls (like searching real-time data) without interrupting the primary conversation flow.

  2. The system maintains conversational coherence during long background computations by using an asynchronous delegation framework where the frontend manages live interaction while the harness executes tasks in the background.

  3. The agent can handle interruptions gracefully: it can stop a stale response, incorporate new user corrections or redirections into its ongoing plan, and resume speech with revised context without losing track of previous dialogue history.

  4. The system supports hour-scale video understanding through a training-free long-video memory module that uses motion compensation to store visual context and retrieval mechanisms (based on MaxSim/MMR) to reconstruct relevant historical information during conversation.

  5. It can exhibit superior performance in spoken question answering (e.g., achieving high scores on Llama Questions and Speech CMMLU) by leveraging a unified training recipe that combines offline understanding with proactive, duplex interaction data.

Abstract

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.

Sources

Related papers