Pith. sign in

REVIEW 12 cited by

DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.09767 v3 pith:TDVTYJRX submitted 2023-12-15 cs.CV

classification cs.CV
keywords dreamtalkemotionsemotionalpersonalizedtalkingacrossconsistentlyconveniently
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotional talking head generation has attracted growing attention. Previous methods, which are mainly GAN-based, still struggle to consistently produce satisfactory results across diverse emotions and cannot conveniently specify personalized emotions. In this work, we leverage powerful diffusion models to address the issue and propose DreamTalk, a framework that employs meticulous design to unlock the potential of diffusion models in generating emotional talking heads. Specifically, DreamTalk consists of three crucial components: a denoising network, a style-aware lip expert, and a style predictor. The diffusion-based denoising network can consistently synthesize high-quality audio-driven face motions across diverse emotions. To enhance lip-motion accuracy and emotional fullness, we introduce a style-aware lip expert that can guide lip-sync while preserving emotion intensity. To more conveniently specify personalized emotions, a diffusion-based style predictor is utilized to predict the personalized emotion directly from the audio, eliminating the need for extra emotion reference. By this means, DreamTalk can consistently generate vivid talking faces across diverse emotions and conveniently specify personalized emotions. Extensive experiments validate DreamTalk's effectiveness and superiority. The code is available at https://github.com/ali-vilab/dreamtalk.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

    cs.MM 2025-10 conditional novelty 6.0 of 10

    Current omni-modal LLMs underperform on audio-visual emotional reasoning, and automatic scores diverge from human perceptual judgments; AV-EMO-Reasoning provides a benchmark to measure this.

  2. FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.

  3. Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new large dataset and the FSCD model improve automated quality scoring of AI-generated talking-head videos, beating 15 baselines in correlation with human ratings.

  4. Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MF-Talk, a mask-free and identity-reference-free three-stage pipeline, improves visual quality and identity preservation in talking-face generation while remaining competitive on lip-sync.

  5. Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A UK-based mixed-methods study finds a wide gap between public trust in biometrics and expert concern about deepfake spoofing, and proposes a tri-layer mitigation framework.

  6. Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark

    cs.MM 2025-02 conditional novelty 6.0 of 10

    AvaMERG is a new text-speech-vision avatar benchmark for empathetic response generation, and the Empatheia system is claimed to outperform baselines on both textual and multimodal empathy tasks.

  7. MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    MoDiT, a diffusion transformer conditioned on 3DMM coefficients and Wav2Lip references, produces talking-head videos with improved same-identity lip sync and more natural blinks in its reported benchmarks.

  8. MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MirrorMe adapts the LTX video diffusion transformer to generate real-time, high-fidelity audio-driven halfbody animations with identity preservation and hand pose control.

  9. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  10. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.

  11. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  12. NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results

    cs.CV 2025-06 conditional novelty 4.0 of 10

    All 19 valid entries in the NTIRE 2025 XGC quality assessment challenge outperformed their track baselines at predicting human quality scores for user-generated video, AI-generated video, and talking heads.

Pith tools