Pith. sign in

REVIEW 1 cited by

VectorTalker: SVG Talking Face Generation with Progressive Vectorisation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.11568 v1 pith:UV52CSHG submitted 2023-12-18 cs.CV cs.GR

classification cs.CVcs.GR
keywords imagevectoranimationaudio-drivengenerationreconstructiontalkingvectortalker
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

High-fidelity and efficient audio-driven talking head generation has been a key research topic in computer graphics and computer vision. In this work, we study vector image based audio-driven talking head generation. Compared with directly animating the raster image that most widely used in existing works, vector image enjoys its excellent scalability being used for many applications. There are two main challenges for vector image based talking head generation: the high-quality vector image reconstruction w.r.t. the source portrait image and the vivid animation w.r.t. the audio signal. To address these, we propose a novel scalable vector graphic reconstruction and animation method, dubbed VectorTalker. Specifically, for the highfidelity reconstruction, VectorTalker hierarchically reconstructs the vector image in a coarse-to-fine manner. For the vivid audio-driven facial animation, we propose to use facial landmarks as intermediate motion representation and propose an efficient landmark-driven vector image deformation module. Our approach can handle various styles of portrait images within a unified framework, including Japanese manga, cartoon, and photorealistic images. We conduct extensive quantitative and qualitative evaluations and the experimental results demonstrate the superiority of VectorTalker in both vector graphic reconstruction and audio-driven animation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FACEMUG fuses up to five input modalities in the StyleGAN latent space to perform local, incremental facial edits while preserving unedited regions.

Pith tools