Pith. sign in

REVIEW 2 cited by

One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.15126 v3 pith:4QN3XPQA submitted 2020-11-30 cs.CV

classification cs.CV
keywords videoconferencingkeypointmodelrepresentationsynthesistalking-headmotion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a neural talking-head video synthesis model and demonstrate its application to video conferencing. Our model learns to synthesize a talking-head video using a source image containing the target person's appearance and a driving video that dictates the motion in the output. Our motion is encoded based on a novel keypoint representation, where the identity-specific and motion-related information is decomposed unsupervisedly. Extensive experimental validation shows that our model outperforms competing methods on benchmark datasets. Moreover, our compact keypoint representation enables a video conferencing system that achieves the same visual quality as the commercial H.264 standard while only using one-tenth of the bandwidth. Besides, we show our keypoint representation allows the user to rotate the head during synthesis, which is useful for simulating face-to-face video conferencing experiences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instant Expressive Gaussian Head Avatars at Over 100 FPS

    cs.CV 2025-12 conditional novelty 7.0 of 10

    A single-photo avatar encoder with per-Gaussian feature-space deformation animates faces at 107 FPS with expression quality competitive with diffusion models.

  2. Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Total-Editing is a unified 3D head avatar framework that separately controls appearance, motion, and lighting through an intrinsically decomposed neural radiance field, and reports stronger identity, expression, pose,...

Pith tools