Pith. sign in

REVIEW 1 cited by

Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15687 v2 pith:MGIYAXKZ submitted 2024-01-28 cs.CV cs.GR

Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

classification cs.CV cs.GR
keywords facialanimationco-speechgnpfamulti-modalitydatasetexpressionsgeneralized
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism and a lack of lexible conditioning. We address this challenge through a trilogy. We first introduce Generalized Neural Parametric Facial Asset (GNPFA), an efficient variational auto-encoder mapping facial geometry and images to a highly generalized expression latent space, decoupling expressions and identities. Then, we utilize GNPFA to extract high-quality expressions and accurate head poses from a large array of videos. This presents the M2F-D dataset, a large, diverse, and scan-level co-speech 3D facial animation dataset with well-annotated emotional and style labels. Finally, we propose Media2Face, a diffusion model in GNPFA latent space for co-speech facial animation generation, accepting rich multi-modality guidances from audio, text, and image. Extensive experiments demonstrate that our model not only achieves high fidelity in facial animation synthesis but also broadens the scope of expressiveness and style adaptability in 3D facial animation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

    cs.CV 2026-05 unverdicted novelty 7.0

    TokTalk trains a chunk-based conditional flow matching model on a new audio-token to 3D facial motion dataset to enable real-time expressive facial animation from Audio-LLM tokens with low overhead adaptation.