Pith. sign in

REVIEW 4 cited by

RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18284 v2 pith:BFFKAAIF submitted 2024-06-26 cs.CV

classification cs.CV
keywords facialalignmentaudio-drivenfacegenerationidentityreal-timesynchronization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and practical applications. The challenges are two-fold: 1) Preserving unique individual traits for achieving high-precision lip synchronization. 2) Generating high-quality facial renderings in real-time performance. In this paper, we propose a novel generalized audio-driven framework RealTalk, which consists of an audio-to-expression transformer and a high-fidelity expression-to-face renderer. In the first component, we consider both identity and intra-personal variation features related to speaking lip movements. By incorporating cross-modal attention on the enriched facial priors, we can effectively align lip movements with audio, thus attaining greater precision in expression prediction. In the second component, we design a lightweight facial identity alignment (FIA) module which includes a lip-shape control structure and a face texture reference structure. This novel design allows us to generate fine details in real-time, without depending on sophisticated and inefficient feature alignment modules. Our experimental results, both quantitative and qualitative, on public datasets demonstrate the clear advantages of our method in terms of lip-speech synchronization and generation quality. Furthermore, our method is efficient and requires fewer computational resources, making it well-suited to meet the needs of practical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automated Synthesis of Facial Mechanisms for Conversational Animatronic Robots

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A parametric linkage face template plus hierarchical collision-driven optimization synthesizes manufacturable facial mechanisms from 2D portraits and runs them with dual-identity conversational motion.

  2. JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync

    cs.CV 2025-07 conditional novelty 6.0 of 10

    JOLT3D jointly trains a 3DMM reconstruction network with a talking head generator, then uses FACS mouth blendshapes from a diffusion model to lip-sync videos while preserving the original chin contour.

  3. FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A frequency-energy router that blends LoRA experts according to the latent's bandwise energy improves diffusion fine-tuning quality and style consistency across multiple backbones.

  4. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

Pith tools