BEACON: Behavior and Appearance Control for Subject-Specific Video Generation
| Source: arXiv AI
Tags: video-generation, diffusion-models, facial-expression, identity-preservation, Wan, human-video
BEACON generates person-specific videos that preserve both visual identity and characteristic facial dynamics by disentangling the two into separate conditioning signals — updating only ~1% of the Wan video diffusion model trained on 2,000 pairs, with improved expressivity over state-of-the-art methods.
Details
Pokrzywa et al. target a key limitation of current human-centric video generation: conditioning only on a single reference image preserves appearance but produces generic, low-variation facial expressions that don't reflect the subject's characteristic emotional style.\n\nBEACON separates two conditioning signals: a reference image encodes visual identity; a reference video captures the subject's specific facial dynamics over time. By conditioning generation on both, BEACON produces videos that preserve appearance while maintaining subject-specific expressivity. The framework also supports identity-expression transfer — applying one person's behavioral dynamics to another person's appearance.\n\nThe implementation builds on the Wan video diffusion model, updating approximately 1% of parameters via fine-tuning on roughly 2,000 image-video pairs. This is a notably lean training footprint. Evaluation on MEAD and RAVDESS benchmark datasets shows improved facial expressivity over current state-of-the-art methods while maintaining competitive identity preservation.\n\nNo public code release is mentioned in the abstract. The paper is a research contribution, with potential applications in entertainment, digital avatars, and personalized video content.