Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
cs.CV, cs.AI
Submitted: 2026-01-04
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Human identity-preserving text-to-video generation remains challenging under large changes in viewpoint, facial expression, illumination, and motion.
Terminology
Abstract
Human identity-preserving text-to-video generation remains challenging under large changes in viewpoint, facial expression, illumination, and motion. Existing methods condition the generator on a single reference portrait, but a static image cannot capture how identity-bearing cues evolve across views and expressions, leading to face deformation, pose locking, identity drift, or over-smoothed faces. We observe that a short reference clip naturally provides richer temporal and multi-view identity cues than any single image, motivating a video-referential formulation. This richer signal, however, introduces a new challenge: identity evidence is distributed across many frames and must be distilled into a compact, stable representation under a limited token budget. To this end, we propose Slot-ID, a lightweight identity-conditioning framework built on a frozen text-to-video backbone. Slot-ID employs a slot-based temporal identity encoder with Sinkhorn-routed iterative reading to distill a compact, stable set of identity tokens from the reference clip, complemented by an image-anchor stream for dual-source conditioning. Extensive experiments demonstrate that Slot-ID outperforms state-of-the-art methods in identity preservation and visual naturalness while remaining competitive in prompt following, with particularly large gains under challenging pose, expression, and motion variations.
Sources
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Open-Sora Plan: Open-Source Large Video Generation Model
- IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
- OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
- I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- InstantID: Zero-shot Identity-Preserving Generation in Seconds
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models