Pith. sign in

REVIEW 15 cited by

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08136 v2 pith:35LX3C7A submitted 2024-07-11 cs.CV

classification cs.CV
keywords echomimicaudiosfaciallandmarksmethodsportraitaudiodriven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The area of portrait image animation, propelled by audio input, has witnessed notable progress in the generation of lifelike and dynamic portraits. Conventional methods are limited to utilizing either audios or facial key points to drive images into videos, while they can yield satisfactory results, certain issues exist. For instance, methods driven solely by audios can be unstable at times due to the relatively weaker audio signal, while methods driven exclusively by facial key points, although more stable in driving, can result in unnatural outcomes due to the excessive control of key point information. In addressing the previously mentioned challenges, in this paper, we introduce a novel approach which we named EchoMimic. EchoMimic is concurrently trained using both audios and facial landmarks. Through the implementation of a novel training strategy, EchoMimic is capable of generating portrait videos not only by audios and facial landmarks individually, but also by a combination of both audios and selected facial landmarks. EchoMimic has been comprehensively compared with alternative algorithms across various public datasets and our collected dataset, showcasing superior performance in both quantitative and qualitative evaluations. Additional visualization and access to the source code can be located on the EchoMimic project page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-human Interactive Talking Dataset

    cs.CV 2025-08 conditional novelty 6.0 of 10

    The paper contributes a 12-hour multi-person conversational video dataset with pose and speaking annotations, plus a baseline model for generating full-body talking videos of two to four people.

  2. Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new large dataset and the FSCD model improve automated quality scoring of AI-generated talking-head videos, beating 15 baselines in correlation with human ratings.

  3. MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MagicAnime is a 400k-clip multimodal cartoon dataset with hierarchical annotations and benchmarks for image-to-video, pose-driven, face reenactment, and audio-driven animation generation.

  4. ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ARIG introduces a real-time, frame-wise autoregressive head generation framework with diffusion-based continuous motion prediction, improving interactive realism over clip-wise methods.

  5. Exploring Timeline Control for Facial Motion Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model generates natural facial motions from user-specified multi-track timelines, using TICC-based frame-level action interval annotation for training and evaluation.

  6. Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A motion-prior diffusion model with archived-frame memory improves identity, lip-sync, and head-motion consistency in long talking-face videos.

  7. Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

    cs.LG 2026-07 reject novelty 5.0 of 10

    A classifier trained on rPPG waveforms from face videos detects talking-face deepfakes with AUC 0.806, and detection difficulty varies by generator (AUC 0.690–0.985).

  8. EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    EDTalk++ disentangles talking-head video into four orthogonal motion banks (mouth, pose, eyes, expression) and drives them from either video or audio inputs.

  9. Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FRVD warps a source face toward driving poses with implicit keypoints, then repairs lost details inside Stable Video Diffusion's latent space, reporting gains over seven baselines on large-pose reenactment benchmarks.

  10. MoDA: Multi-modal Diffusion Architecture for Talking Head Generation

    cs.GR 2025-07 conditional novelty 5.0 of 10

    MoDA uses flow matching in a compact face-motion space with a progressively fused multi-modal transformer to generate expressive, lip-synced talking-head videos from a single image and audio.

  11. FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases

    cs.CV 2025-07 conditional novelty 5.0 of 10

    FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.

  12. GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GGTalker combines large-scale audio-to-expression and expression-to-texture priors with rapid per-identity fine-tuning to create high-quality 3D talking heads from a short video.

  13. SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SyncTalk++ synthesizes speech-driven talking-head videos via 3D Gaussian Splatting and reports state-of-the-art synchronization and quality at up to 101 FPS.

  14. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An audio-conditioned video diffusion transformer that animates portraits from image, video, text, and audio inputs with a sliding-window fusion for long videos.

  15. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

Pith tools