Pith. sign in

REVIEW 3 cited by

Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13570 v2 pith:SJFIWFAR submitted 2024-03-20 cs.CV

classification cs.CV
keywords headmulti-viewlearnsynthesizervideospseudoavatarmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose a novel learning approach for feed-forward one-shot 4D head avatar synthesis. Different from existing methods that often learn from reconstructing monocular videos guided by 3DMM, we employ pseudo multi-view videos to learn a 4D head synthesizer in a data-driven manner, avoiding reliance on inaccurate 3DMM reconstruction that could be detrimental to the synthesis performance. The key idea is to first learn a 3D head synthesizer using synthetic multi-view images to convert monocular real videos into multi-view ones, and then utilize the pseudo multi-view videos to learn a 4D head synthesizer via cross-view self-reenactment. By leveraging a simple vision transformer backbone with motion-aware cross-attentions, our method exhibits superior performance compared to previous methods in terms of reconstruction fidelity, geometry consistency, and motion control accuracy. We hope our method offers novel insights into integrating 3D priors with 2D supervisions for improved 4D head avatar creation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. fMRI2Face: A Full-HD fMRI-Video Dataset and Geometry-Guided Neural Decoding Framework for Dynamic Human Face Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A 62,856-sample fMRI dataset of full-HD digital-human face videos plus a geometry-guided video-diffusion decoder that reconstructs facial identity and motion from brain signals.

  2. Total-Editing: Head Avatar with Editable Appearance, Motion, and Lighting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Total-Editing is a unified 3D head avatar framework that separately controls appearance, motion, and lighting through an intrinsically decomposed neural radiance field, and reports stronger identity, expression, pose,...

  3. FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A one-shot model for animatable 3D/4D Gaussian head reconstruction that adds attention regularization, decoupled reconstruction-animation training, and autoregressive visibility-gated fusion, reporting consistent metr...

Pith tools