Pith. sign in

REVIEW 3 cited by

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07954 v2 pith:S4AOVWTM submitted 2023-06-13 cs.CV

classification cs.CV
keywords frameworkframestranslationvideodiffusionmodelsconsistencyexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large text-to-image diffusion models have exhibited impressive proficiency in generating high-quality images. However, when applying these models to video domain, ensuring temporal consistency across video frames remains a formidable challenge. This paper proposes a novel zero-shot text-guided video-to-video translation framework to adapt image models to videos. The framework includes two parts: key frame translation and full video translation. The first part uses an adapted diffusion model to generate key frames, with hierarchical cross-frame constraints applied to enforce coherence in shapes, textures and colors. The second part propagates the key frames to other frames with temporal-aware patch matching and frame blending. Our framework achieves global style and local texture temporal consistency at a low cost (without re-training or optimization). The adaptation is compatible with existing image diffusion techniques, allowing our framework to take advantage of them, such as customizing a specific subject with LoRA, and introducing extra spatial guidance with ControlNet. Extensive experimental results demonstrate the effectiveness of our proposed framework over existing methods in rendering high-quality and temporally-coherent videos.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FloAt: Flow Warping of Self-Attention for Clothing Animation Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FloAtControlNet animates clothing by flow-warping self-attention maps in a normal-map-conditioned ControlNet, improving temporal coherence and reducing background flicker.

  2. GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

    cs.CV 2026-08 conditional novelty 5.0 of 10

    A geometry-aware, training-free inference framework that refines pretrained video diffusion predictions with projected static history content and view-conditioned routing achieves fifth place on AI City Challenge Track 5.

  3. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

Pith tools