Pith. sign in

REVIEW 3 cited by

ViBiDSampler: Enhancing Video Interpolation Using Bidirectional Diffusion Sampler

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05651 v3 pith:X2AKIZAF submitted 2024-10-08 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords interpolationdiffusionframesgenerationmethodbackwardbidirectionalconditioned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent progress in large-scale text-to-video (T2V) and image-to-video (I2V) diffusion models has greatly enhanced video generation, especially in terms of keyframe interpolation. However, current image-to-video diffusion models, while powerful in generating videos from a single conditioning frame, need adaptation for two-frame (start & end) conditioned generation, which is essential for effective bounded interpolation. Unfortunately, existing approaches that fuse temporally forward and backward paths in parallel often suffer from off-manifold issues, leading to artifacts or requiring multiple iterative re-noising steps. In this work, we introduce a novel, bidirectional sampling strategy to address these off-manifold issues without requiring extensive re-noising or fine-tuning. Our method employs sequential sampling along both forward and backward paths, conditioned on the start and end frames, respectively, ensuring more coherent and on-manifold generation of intermediate frames. Additionally, we incorporate advanced guidance techniques, CFG++ and DDS, to further enhance the interpolation process. By integrating these, our method achieves state-of-the-art performance, efficiently generating high-quality, smooth videos between keyframes. On a single 3090 GPU, our method can interpolate 25 frames at 1024 x 576 resolution in just 195 seconds, establishing it as a leading solution for keyframe interpolation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    TLB-VFI reports state-of-the-art FID, LPIPS, and FloLPIPS on VFI benchmarks using a latent Brownian bridge diffusion model with temporal-aware encoding and wavelet gating.

  2. Semantic Frame Interpolation

    cs.CV 2025-07 reject novelty 5.0 of 10

    The authors define Semantic Frame Interpolation, build a 300k-clip dataset and benchmark, and propose a Mixture-of-LoRA adaptation of Wan2.1 that improves temporal smoothness but does not preserve the given start and ...

  3. DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DiffuseSlide boosts the frame rate of latent diffusion videos via latent interpolation, noise re-injection, and sliding-window denoising, reporting better FVD, PSNR, and SSIM than several baselines on WebVid-10M.

Pith tools