FacePlex: Toward Natural Full-Duplex Conversational Avatars
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "FacePlex: Toward Natural Full-Duplex Conversational Avatars".
Jane: Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we’re checking out FacePlex today, which is this paper titled "FacePlex: Toward Natural Full-Duplex Conversational Avatars." Basically, they're tackling the tough problem of making natural face-to-face conversation that needs both real-time speech and synchronized facial movement happening at the same time.
Jane: That sounds really complex for something we hear on a radio show, Tom. Could you explain what this paper claims is the main gap it's trying to fill?
Tom: Well, FacePlex formalizes the task as a full-duplex joint speech-facial motion generation system, which means they are aiming to generate both speech and facial motion together at every single time step without having to wait for whole sentences. The authors claim this unified streaming framework addresses the limitations of existing systems that only handle one part or operate offline.
Lu: I think the core idea here is moving away from treating speech and face as separate problems, which makes a lot of sense for creating truly immersive interaction, Lu says.
Meng: From an engineering standpoint, it’s interesting how they formalize the task this way; it sounds like they're building a system where every step has to handle both modalities simultaneously under streaming constraints.
Lalam: I see this as a major step forward because if we can generate these coupled streams online, the cultural impact could be huge for how we interact with digital interfaces.
Tom: Exactly, and the paper points out two big challenges they are trying to solve: generating facial motion fast enough for real-time interaction, which means every eighty milliseconds for matching speech token rates, and aligning those speech tokens with the facial motion when they are streaming asynchronously.
Jane: Aligning them asynchronously sounds tricky because speech tokens and continuous movements don't have the same temporal structure, Tom. How does FacePlex actually manage that synchronization issue?
Tom: They introduce Rolling Flow Matching to adapt flow matching for online motion generation by committing new motion frames at each step, which helps with the speed challenge of generating facial motion. Plus, they bring in Rolling Cross-Attention to couple the streaming audio queue with the motion queue so they can condition each other as generation progresses.
Lu: That rolling structure sounds like a clever way to handle the temporal mismatch between those two streams without needing pre-packaged chunks, which is what really caught my eye.
Paper summary: Meng: I'm thinking about the practical side of things; how does this rolling mechanism translate into actual computational efficiency when running these models in real-time?
Lalam: If this framework works well, it could fundamentally improve the quality of digital communication interfaces across many applications, Lu muses.
Tom: And to tackle the alignment issue specifically, Rolling Cross-Attention allows each motion pair to attend to a bounded window of past and near-future speech as it is progressively denoised. They even analyze four different attention topologies: full, aligned, causal, and anti-causal.
Jane: Four different ways to manage that context flow sounds like a lot of design work before you even get into the actual generation part. Which one do they suggest is the best choice?
Tom: The paper analyzes how these choices affect synchronization and motion quality under sub-second emission, but it doesn't single out one as definitively superior; they examine their effects on those metrics.
Lu: The idea of controlling context flow through different attention topologies suggests a very nuanced approach to modeling how speech influences the face, which is really creative.
Meng: So, if we look at the results, what kind of improvements are they actually showing compared to previous work?
Tom: The evaluation shows that FacePlex achieves stronger lip-sync quality and motion fidelity than audio-driven facial motion baselines. The training data used was a mix of synthetic self-play and real dyadic interaction recordings from the Seamless Interaction dataset, totaling about one thousand one hundred thirty-eight hours of paired speech–motion streams.
Jane: That’s a substantial amount of paired data they used; having both synthetic and real data sounds like a smart way to train something robust.
Lalam: The fact that they achieved high ratings across criteria like Lip Synchronization and Conversational Interaction is really encouraging for how natural the resulting avatars could feel.
Tom: And the ablation studies confirm the importance of these components, showing that removing Rolling Flow Matching severely degrades temporal coherence, and removing Rolling Cross-Attention hurts lip-sync and motion fidelity by replacing cross-modal attention with single hidden-state conditioning.
Lu: It’s interesting how the ablation studies clearly isolate why each part is necessary for achieving that level of joint performance.
Meng: So, while they show the mechanism works, what’s the biggest practical limitation they admit? Where does this system stop working effectively?
Paper summary: Tom: The paper flags a few things; one limitation is that they are focusing on FLAME-parameter facial motion rather than photorealistic video, which means it doesn't account for appearance-level factors like lighting or identity preservation. Also, it currently only models speech-coupled facial motion and doesn't model full embodied behaviors such as gaze control or body gesture.
Jane: So they are focusing on the synchronization aspect of speech and face generation, but they aren't aiming to create a fully realistic video output with complex body language yet?
Tom: That’s right; the focus is on the coupled streams themselves, not necessarily photorealism or full embodiment. However, the potential impact remains large because it enables more natural and accessible real-time conversational interfaces for things like telepresence and remote education by treating speech and visual outputs as coupled streams instead of separate modules.
Lu: Thinking about the future, if we can get better at modeling these coupled streams, I see possibilities for creating much more intuitive virtual companions or tutors that truly feel present.
Meng: From an engineering viewpoint, the next step would probably be integrating those embodied behaviors you mentioned into this framework so it’s not just talking and moving its mouth but actually interacting physically in a simulated environment.
Lalam: If we can achieve this level of natural coupling, it opens up possibilities for conversational AI that feels genuinely human in a way that current systems just don't get yet.
Tom: So, to wrap up the main idea of FacePlex: it’s a unified streaming framework that jointly generates speech and facial motion online, using Rolling Flow Matching and Rolling Cross-Attention to achieve better lip-sync and motion fidelity than existing methods.
Jane: And the authors are Hah Min Lew, Jae-Ho Lee, Ha001211@korea.ac.kr, Ji-Su Kang2 jisu.kang@klleon.io, Gyeong-Moon Park1† gm-park@korea.ac.kr and others on the team at Korea University and Klleon AI Research Institute for their work on this paper?
Lu: Indeed, their work is quite innovative in how they structure the training objective to mirror that mixed-flow-time queue structure directly, which is a key design choice.
Meng: It’s solid research, but as an engineer, I'm curious if they have any plans to extend this beyond just conversational avatars into more complex real-time interaction scenarios soon?
Lalam: The potential impact on making digital communication feel more natural is significant; it moves the goal closer to truly fluid interaction.
Conclusion: Tom: So, we've looked at how FacePlex tackles generating speech and facial motion together in real-time. Jane, let's get back to those core elements of the paper—the title and who came up with this work.
Jane: Right, Tom, the title itself, "FacePlex," really tells you what they’re aiming for: a unified framework for full-duplex conversational avatars. And I think it points directly to how they structured their entire approach to solving the streaming problem.
Lu: Exactly! From a creative standpoint, seeing that name suggests an attempt to fuse two traditionally separate domains into one coherent system, which is where the real potential lies for next-generation digital presence.
Meng: I'm thinking about the authors; they’ve put together a pretty solid set of components—the PersonaPlex LLM backbone, the audio branch, and that motion branch—which shows a structured design effort to manage that complexity.
Lalam: I agree with Meng; it’s clear they didn't just throw some loose ideas together; there was a deliberate architectural plan to handle the joint generation challenge under streaming constraints.
Tom: It really comes down to how they managed those two big hurdles we talked about earlier—the speed of motion and the timing alignment between speech and face.
Jane: That’s the crucial point, Tom; they didn't just build a better speaker or a better face generator in isolation; they built a system that treats them as one continuous stream.
Lu: And by unifying them, Lu thinks we open up possibilities for truly fluid interaction that goes far beyond simple back-and-forth dialogue.
Meng: From an engineering viewpoint, the implication is that the next generation of interactive AI won't just be reactive; it will be actively presenting a cohesive and synchronized presence.
Lalam: That synchronization, Lu, is where the cultural impact really hits; if these avatars can sound and look perfectly matched in real-time, it could fundamentally change how we interact with digital learning and remote education.
Tom: So we’re seeing a unified system that handles the hard technical synchronization problems mentioned earlier. It’s a lot to take in, but the potential for making digital communication feel much more natural is huge.
Jane: That feeling of naturalness, Tom, is what makes this research so exciting; it moves us closer to conversational AI that doesn't feel like it's reading from a script or reacting with a noticeable delay.
Lu: And as we look ahead, I see this framework as a foundation for much more complex embodied interactions down the line, imagine avatars with full body language!
Meng: That’s where my practical curiosity kicks in; integrating those larger physical behaviors into this coupled stream model will be the next big engineering challenge they have to face.
Lalam: If we can achieve that level of natural coupling, Tom, it means AI companions could develop a presence that feels genuinely human in a way current systems simply can't replicate yet.
Korea University
cs.AI, cs.CV, cs.LG
Submitted: 2026-06-29
Updated: 2026-09-28
Comments: Project page: https://hahminlew.github.io/faceplex
Project page: https://hahminlew.github.io/faceplex
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 83/100
The gist: Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion.
Key concepts
- Full-Duplex Joint Speech-Facial Motion Generation
- This is the core task: generating both spoken audio and corresponding facial movements at the same time without waiting for a full sentence to finish. It requires the system to handle continuous, real-time input and output streams concurrently, which is necessary for natural conversation.
- Rolling Flow Matching (RFM)
- RFM adapts a technique called flow matching specifically for online motion generation. Instead of generating all motion at once, it commits new motion frames incrementally as the stream progresses. This allows the model to learn how to generate smooth, continuous facial movement step-by-step during streaming.
- Rolling Cross-Attention (RCA)
- RCA links the audio stream and the motion generation process together. It uses attention mechanisms to allow speech context and previous motion states to condition each other dynamically. This coupling ensures that the generated lip movements accurately reflect both what is being said and where they should be in time.
Terminology
Summary
Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. FacePlex proposes a unified streaming framework that jointly generates speech and facial motion online, addressing the limitations of existing systems that only handle one modality or operate offline.
The gist
FacePlex enables full-duplex joint speech-facial motion generation under online streaming constraints, while achieving stronger lip-sync quality and motion fidelity than audio-driven facial motion baselines.
Problem Formulation and Motivation
The paper formalizes the task as a full-duplex joint speech-facial motion generation system,
defined as simultaneously processing incoming user signals and generating outgoing responses at every time step without buffering complete utterances, which inherently requires streaming. The research identifies two coupled challenges: (C1) Generating high-quality facial motion fast enough for real-time interaction, where facial motion must be generated every 80 ms to match the speech token rate; and (C2) Aligning speech with facial motion when both stream asynchronously due to different temporal abstractions between speech tokens and continuous articulator trajectories.
FacePlex Framework Components
FacePlex is a unified streaming framework built on three jointly trained components: a PersonaPlex LLM backbone, an audio branch, and a motion branch. The system maintains three rolling queues that advance once per model step to manage the streaming process:
-
The audio queue (AT) stores generated audio chunks that have been predicted but not yet emitted so each chunk can be released together with its aligned facial motion.
-
The hidden-state queue (HT) stores recent PersonaPlex-backbone hidden states used to condition motion generation.
-
The motion queue (MT) stores the corresponding motion-pair states being progressively refined through denoising.
Rolling Flow Matching (RFM)
Rolling Flow Matching is introduced to adapt flow matching to online motion generation by committing new motion frames at each streaming step, addressing challenge (C1). It maintains a small motion queue with slots at different flow-time states, where the front slot is near-clean and ready for emission, and the back slot is newly initialized noise. The training objective mirrors this mixed-flow-time queue structure to allow the model to learn streaming generation directly rather than relying on offline chunk generation.
Rolling Cross-Attention (RCA)
Rolling Cross-Attention couples the streaming audio queue with the motion queue, allowing speech and facial motion to condition each other as generation progresses, addressing challenge (C2). RCA conditions the rolling motion queue on a parallel rolling hidden-state queue (HT) using a binary mask A. The paper analyzes four attention topologies: full, aligned, causal, and anti-causal. Full RCA combines both past and future speech context to allow each motion pair to attend to an approximately ±240 ms speech context throughout its denoising life cycle.
Training and Evaluation
The system is trained end-to-end using a combination of PersonaPlex speech-modeling losses and the RFM objective. Training data is constructed from both synthetic self-play (PersonaPlex in two-speaker mode) and real dyadic interaction recordings from the Seamless Interaction dataset, totaling approximately 1,138 hours of paired speech–motion streams. Evaluation involves chunk-wise streaming protocols for facial motion and metrics for full-duplex speech interaction (e.g., Pause TOR, Backchannel Frequency). User studies confirm that FacePlex receives the highest ratings across all criteria in Lip Synchronization, Natural & Coherence, Conversational Interaction, and Overall MOS compared to baselines.
Ablation Summary
Ablation studies validate the necessity of key components: removing RFM severely degrades metrics by destroying temporal coherence; removing RCA degrades lip-sync and motion fidelity by replacing cross-modal attention with single hidden-state conditioning; training on real data alone yields the worst lip-sync, while combining synthetic and real data yields the best overall trade-off. RCA masking strategies are also analyzed, showing that Full RCA (combining past, aligned, and future context) achieves the strongest motion fidelity and best PLRS scores.
Limitations and Impact
Limitations include focusing on FLAME-parameter facial motion rather than photorealistic video, which excludes appearance-level factors like lighting or identity preservation. Furthermore, the system currently focuses on speech-coupled facial motion and does not yet model full embodied behaviors such as gaze control or body gesture. However, the potential positive impact includes enabling more natural and accessible real-time conversational interfaces for telepresence and remote education by modeling speech and visual outputs as coupled streams rather than separate modules.
Potential Negative Impact
The same capabilities that enhance naturalness introduce risks, including the potential for high-quality speech-synchronized facial motion to be misused to create deceptive synthetic media or impersonate real individuals, increasing the risk of deepfakes and social engineering.
Improvements for AI systems
Based on the research presented in FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars,
here are specific, high-impact improvements that could be implemented in existing AI systems, and what the resulting improved system would be capable of.
The core innovation of FacePlex is moving from uni-modal (speech-only or audio-driven) generation to a unified, streaming, full-duplex architecture. The improvements focus on achieving seamless synchronization under real-time constraints.
Here are the specific improvements:
The improved AI system would be capable of:
Abstract
Natural human conversation is inherently a real-time interaction in which speech and facial behavior continuously evolve. Enabling such interaction requires a conversational avatar to jointly generate speech and facial motion in real time, prepare facial motion for upcoming speech before the corresponding audio is emitted, and produce non-verbal reactions that reflect the ongoing dialogue. However, existing conversational avatar systems cannot address such requirements: audio-driven methods rely on pre-given speech, while joint streaming generation alone does not ensure anticipatory articulation or semantically appropriate reactions. We propose FacePlex, a unified framework for full-duplex speech-facial motion generation and real-time avatar rendering. FacePlex jointly coordinates speech, facial motion, and Gaussian splatting rendering on a shared streaming timeline. To prepare facial motion for upcoming speech, FacePlex predicts a short speech continuation and uses its future hidden states through asymmetric conditioning and denoising, while continuously updating them as new user audio arrives without observing future user input. For dialogue-grounded non-verbal behavior, we construct SemReact, a dataset aligning dialogue context, reaction semantics, and facial motion, and introduce a semantic behavior router that guides continuous motion during both speaking and listening. Extensive experiments show improved audio-visual synchronization and facial articulation, natural dialogue-grounded non-verbal responses, and low-latency of 122 ms end-to-end avatar interaction.
Sources
- Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
- UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
- Moshi: a speech-text foundation model for real-time dialogue
- Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Rolling Diffusion Models
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection