REVIEW 2 cited by
Layered Neural Rendering for Retiming People in Video
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a method for retiming people in an ordinary, natural video -- manipulating and editing the time in which different motions of individuals in the video occur. We can temporally align different motions, change the speed of certain actions (speeding up/slowing down, or entirely "freezing" people), or "erase" selected people from the video altogether. We achieve these effects computationally via a dedicated learning-based layered video representation, where each frame in the video is decomposed into separate RGBA layers, representing the appearance of different people in the video. A key property of our model is that it not only disentangles the direct motions of each person in the input video, but also correlates each person automatically with the scene changes they generate -- e.g., shadows, reflections, and motion of loose clothing. The layers can be individually retimed and recombined into a new video, allowing us to achieve realistic, high-quality renderings of retiming effects for real-world videos depicting complex actions and involving multiple individuals, including dancing, trampoline jumping, or group running.
Forward citations
Cited by 2 Pith papers
-
Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation
A neural surface representation with foreground and background modules performs self-supervised object segmentation from multi-view images, producing finer masks than NeRF-based counterparts.
-
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
VDP decomposes a single test video into layers and opacity maps via two U-Nets, achieving strong unsupervised video object segmentation, dehazing, and relighting, though the relighting model reduces to gamma correction.
Discussion (0). Continue with ORCID to comment.