Pith. sign in

REVIEW 31 cited by

V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02511 v1 pith:OLFAK2RP submitted 2024-06-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords conditionsgenerationportraitsignalsaudiocontroleffectiveimage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the field of portrait video generation, the use of single images to generate portrait videos has become increasingly prevalent. A common approach involves leveraging generative models to enhance adapters for controlled generation. However, control signals (e.g., text, audio, reference image, pose, depth map, etc.) can vary in strength. Among these, weaker conditions often struggle to be effective due to interference from stronger conditions, posing a challenge in balancing these conditions. In our work on portrait video generation, we identified audio signals as particularly weak, often overshadowed by stronger signals such as facial pose and reference image. However, direct training with weak signals often leads to difficulties in convergence. To address this, we propose V-Express, a simple method that balances different control signals through the progressive training and the conditional dropout operation. Our method gradually enables effective control by weak conditions, thereby achieving generation capabilities that simultaneously take into account the facial pose, reference image, and audio. The experimental results demonstrate that our method can effectively generate portrait videos controlled by audio. Furthermore, a potential solution is provided for the simultaneous and effective use of conditions of varying strengths.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instant Expressive Gaussian Head Avatars at Over 100 FPS

    cs.CV 2025-12 conditional novelty 7.0 of 10

    A single-photo avatar encoder with per-Gaussian feature-space deformation animates faces at 107 FPS with expression quality competitive with diffusion models.

  2. Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.

  3. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0 of 10

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  4. OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    OmniHuman-1.5 combines MLLM-based planning with a multimodal diffusion transformer and pseudo-last-frame identity conditioning to generate context-aware avatar videos from audio and a reference image.

  5. TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 6.0 of 10

    TalkVid is a 1,244-hour, 7,729-speaker, 15-language talking-head video dataset with a stratified evaluation benchmark, and models trained on it generalize better across demographics.

  6. Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

    cs.CV 2025-08 reject novelty 6.0 of 10

    A text-only talking face generator maps text to visemes, hallucinates pseudo-audio, and renders lip-synced video; the paper's headline metrics are partially contradicted by its own tables.

  7. X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    X-NeMo trains a 1D identity-agnostic motion descriptor end-to-end with a diffusion model, enabling zero-shot portrait animation with improved identity and expression fidelity.

  8. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.

  9. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...

  10. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.

  11. OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A mixed-condition training recipe, combining text, audio, and pose signals, lets a single diffusion transformer animate photos into realistic talking or singing videos at unprecedented data scale.

  12. UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting Control

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UniAvatar integrates FLAME-based 3D motion rendering and SH-based illumination rendering into a diffusion talking-head model, enabling separate or combined control of motion and lighting in generated videos.

  13. FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FADA distills a diffusion-based talking avatar model into a 6-step student that mimics multi-condition classifier-free guidance with learnable tokens, achieving 4.17 to 12.5 times NFE speedup with comparable quality.

  14. VividFace: A Diffusion-Based Hybrid Framework for High-Fidelity Video Face Swapping

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion-based framework for video face swapping that hybrid-trains on images and videos, and reports gains in identity preservation and temporal consistency over frame-by-frame baselines.

  15. MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MEMO introduces memory-guided linear attention and emotion-aware multi-modal attention for audio-driven talking video generation, reporting state-of-the-art quality on self-collected test sets.

  16. INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A unified two-stage model uses dual-track audio and learnable memory banks to generate expressive head motions for an agent that freely switches between speaking and listening.

  17. Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

    cs.MM 2024-11 conditional novelty 6.0 of 10

    Sonic generates audio-driven portrait videos from a single image using only global audio cues, with a new time-aware shift fusion that improves long-video stability and lip sync.

  18. EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    EmotiveTalk generates talking-head videos by decoupling audio into lip and expression latents and conditioning a video diffusion model on the separate signals.

  19. InfinityHuman: Towards Long-Term Audio-Driven Human

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.

  20. StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A diffusion-based avatar generator that uses a timestep-aware audio adapter, audio-adaptive guidance, and weighted sliding-window fusion to produce audio-synced talking-head videos several minutes long with reduced id...

  21. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  22. OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-driven avatar generator that injects Wav2Vec2 audio features as additive latents into multiple DiT layers of a LoRA-fine-tuned Wan2.1 model, improving lip-sync and enabling prompt-controlled full-body animation.

  23. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.

  24. KeySync: A Robust Approach for Leakage-free Lip Synchronization in High Resolution

    cs.CV 2025-05 conditional novelty 5.0 of 10

    KeySync applies a keyframe-interpolated diffusion model with a lower-face mask to generate 512x512 lip-synced video with reduced expression leakage and SAM2-based occlusion handling.

  25. Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation

    cs.CV 2025-04 conditional novelty 5.0 of 10

    DICE-Talk improves emotional talking-head generation by combining an audio-visual Gaussian emotion prior, a vector-quantized emotion bank, and an auxiliary emotion classifier in a diffusion model.

  26. Joint Learning of Depth and Appearance for Portrait Image Animation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single diffusion model jointly generates portrait RGB images and aligned depth maps, and its fine-tuned variants can estimate depth, edit from depth, relight, and produce audio-driven talking heads with depth.

  27. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  28. FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control

    cs.CV 2025-07 reject novelty 4.0 of 10

    A single caption can drive pose generation, person-image synthesis, and relighting through a three-stage FashionPose pipeline, with reported text-to-pose gains on DF-PASS that are undermined by inconsistent tables.

  29. SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    SyncAnimation introduces a NeRF-based system that generates audio-synchronized upper-body and head animations with facial expressions in real time.

  30. Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling

    cs.CV 2024-12 conditional novelty 4.0 of 10

    Dialogue Director converts dialogue scripts into multi-view storyboards using GPT-4-based script analysis, multi-view diffusion, and cinematic layout planning, with mixed quantitative gains over baselines.

  31. Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook

    cs.CV 2024-11 conditional novelty 4.0 of 10

    The paper introduces BioDeepAV, a benchmark of real and fake talking-face videos, and reports that state-of-the-art deepfake detectors drop sharply when tested on deepfakes from unseen generators.

Pith tools