Pith. sign in

REVIEW 1 cited by

DiffTalker: Co-driven audio-image diffusion for talking faces via intermediate landmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07509 v1 pith:U7E6GFOL submitted 2023-09-14 cs.CV

classification cs.CV
keywords difftalkerfacesaudiotalkingdiffusionlandmarksimagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving. DiffTalker addresses the challenges associated with directly applying diffusion models to audio control, which are traditionally trained on text-image pairs. DiffTalker consists of two agent networks: a transformer-based landmarks completion network for geometric accuracy and a diffusion-based face generation network for texture details. Landmarks play a pivotal role in establishing a seamless connection between the audio and image domains, facilitating the incorporation of knowledge from pre-trained diffusion models. This innovative approach efficiently produces articulate-speaking faces. Experimental results showcase DiffTalker's superior performance in producing clear and geometrically accurate talking faces, all without the need for additional alignment between audio and image features.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Simple and Efficient Baseline for Zero-Shot Generative Classification

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GDC classifies images by fitting one Gaussian per class to DINOv2 embeddings of diffusion-generated reference images, reaching 71.4% on ImageNet at 0.03 seconds per image.

Pith tools