REVIEW 6 cited by
DiffusionTalker: Personalization and Acceleration for Speech-Driven 3D Face Diffuser
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Speech-driven 3D facial animation has been an attractive task in both academia and industry. Traditional methods mostly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the non-deterministic fact of speech-driven 3D face animation and employ the diffusion model for the task. However, personalizing facial animation and accelerating animation generation are still two major limitations of existing diffusion-based methods. To address the above limitations, we propose DiffusionTalker, a diffusion-based method that utilizes contrastive learning to personalize 3D facial animation and knowledge distillation to accelerate 3D animation generation. Specifically, to enable personalization, we introduce a learnable talking identity to aggregate knowledge in audio sequences. The proposed identity embeddings extract customized facial cues across different people in a contrastive learning manner. During inference, users can obtain personalized facial animation based on input audio, reflecting a specific talking style. With a trained diffusion model with hundreds of steps, we distill it into a lightweight model with 8 steps for acceleration. Extensive experiments are conducted to demonstrate that our method outperforms state-of-the-art methods. The code will be released.
Forward citations
Cited by 6 Pith papers
-
GraphAvatar: Compact Head Avatars with GNN-Generated 3D Gaussians
Head avatars are produced by graph-neural-network-generated 3D Gaussians, cutting model size to about 10 MB and improving reported image quality over prior Gaussian-splatting avatars.
-
GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian Splatting
GazeGaussian is a 3D Gaussian Splatting based gaze redirection method that separately models face and eyes and reports state-of-the-art redirection accuracy and image quality on ETH-XGaze, ColumbiaGaze, MPIIFaceGaze, ...
-
SMGDiff: Soccer Motion Generation using diffusion probabilistic models
SMGDiff generates real-time, user-controllable soccer animations with an autoregressive diffusion model plus a contact guidance module, trained on a new 1.08-million-frame soccer motion dataset.
-
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
A 3D facial animation framework that disentangles content and emotion and predicts frame-wise emotion intensity from audio plus text for dynamic expressions.
-
A Dynamic and High-Precision Method for Scenario-Based HRA Synthetic Data Collection in Multi-Agent Collaborative Environments Driven by LLMs
Fine-tuning Qwen2.5-7B on reactor-operator simulator data yields workload estimates that the authors report as more accurate than zero-shot commercial LLMs, but the evaluation lacks a demonstrated train/test split.
-
A Comprehensive Review of Human Error in Risk-Informed Decision Making: Integrating Human Reliability Assessment, Artificial Intelligence, and Human Performance Models
A review of human error research concluding that integrating AI and cognitive models into human reliability assessment can markedly improve predictive fidelity, but data scarcity and opacity remain barriers.
Discussion (0). Continue with ORCID to comment.