Pith. sign in

REVIEW 1 cited by

Video Generation from Single Semantic Label Map

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.04480 v1 pith:OIXVVVKL submitted 2019-03-11 cs.CV

classification cs.CV
keywords generationvideosemanticsinglelabelconditionedcontentflow
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes the novel task of video generation conditioned on a SINGLE semantic label map, which provides a good balance between flexibility and quality in the generation process. Different from typical end-to-end approaches, which model both scene content and dynamics in a single step, we propose to decompose this difficult task into two sub-problems. As current image generation methods do better than video generation in terms of detail, we synthesize high quality content by only generating the first frame. Then we animate the scene based on its semantic meaning to obtain the temporally coherent video, giving us excellent results overall. We employ a cVAE for predicting optical flow as a beneficial intermediate step to generate a video sequence conditioned on the initial single frame. A semantic label map is integrated into the flow prediction module to achieve major improvements in the image-to-video generation process. Extensive experiments on the Cityscapes dataset show that our method outperforms all competing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optical-Flow Guided Prompt Optimization for Coherent Video Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MotionPrompt improves temporal consistency in text-to-video diffusion models by optimizing learnable prompt tokens during sampling, guided by an optical-flow discriminator.

Pith tools