REVIEW 3 cited by
Copy Motion From One to Another: Fake Motion Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
One compelling application of artificial intelligence is to generate a video of a target person performing arbitrary desired motion (from a source person). While the state-of-the-art methods are able to synthesize a video demonstrating similar broad stroke motion details, they are generally lacking in texture details. A pertinent manifestation appears as distorted face, feet, and hands, and such flaws are very sensitively perceived by human observers. Furthermore, current methods typically employ GANs with a L2 loss to assess the authenticity of the generated videos, inherently requiring a large amount of training samples to learn the texture details for adequate video generation. In this work, we tackle these challenges from three aspects: 1) We disentangle each video frame into foreground (the person) and background, focusing on generating the foreground to reduce the underlying dimension of the network output. 2) We propose a theoretically motivated Gromov-Wasserstein loss that facilitates learning the mapping from a pose to a foreground image. 3) To enhance texture details, we encode facial features with geometric guidance and employ local GANs to refine the face, feet, and hands. Extensive experiments show that our method is able to generate realistic target person videos, faithfully copying complex motions from a source person.
Forward citations
Cited by 3 Pith papers
-
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
STDPose combines a Dynamic-Aware Mask, temporal feature and heatmap fusion, and a mutual information loss to improve video pose propagation and pose estimation on sparsely-labeled videos, achieving new state-of-the-ar...
-
Optimizing Human Pose Estimation Through Focused Human and Joint Regions
VREMD combines human and keypoint masks with bidirectional deformable cross-attention to reach state-of-the-art mAP on three PoseTrack benchmarks.
-
GC-ConsFlow: Leveraging Optical Flow Residuals and Global Context for Robust Deepfake Detection
A dual-stream deepfake detector using global context attention and optical flow residuals reports high accuracy on FF++ and Celeb-DF, but with no error bars or code and only marginal gains over prior work.
Discussion (0). Continue with ORCID to comment.