BEACON: Behavior and Appearance Control for Subject-Specific Video Generation

arXiv:2609.13264 · cs.CV, cs.AI · Submitted 2026-09-07 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-07

Updated: 2026-09-07

License: http://creativecommons.org/licenses/by/4.0/

The gist: Generating human-centric videos that preserve both visual identity and person-specific expressive behavior remains a fundamental challenge.

Terminology

Abstract

Generating human-centric videos that preserve both visual identity and person-specific expressive behavior remains a fundamental challenge. In addition to reproducing appearance, a model must replicate the facial behaviors that characterize how a subject expresses emotion over time. However, most state-of-the-art methods condition generation on a single reference image, which contains no information about these temporal dynamics. As a result, they tend to preserve the subject's visual identity but often produce expressions with limited variation and weak subject specificity. To mitigate this issue, we introduce BEACON, a lightweight framework for person-specific video generation that produces more expressive videos by disentangling visual identity from expressive behavior. BEACON conditions generation on two complementary signals: a reference image encoding the identity and a reference video capturing subject-specific facial dynamics. By conditioning on these complementary signals, BEACON generates videos that better preserve both the subject's appearance and characteristic facial dynamics, while also supporting identity-expression transfer. Our experiments on the MEAD and RAVDESS datasets show that by fine-tuning on approximately 2,000 pairs and updating about 1% of the pretrained Wan video diffusion model, BEACON improves facial expressivity over state-of-the-art video generation methods while maintaining competitive identity preservation.

Related papers