REVIEW 7 cited by
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present a method for generating video sequences with coherent motion between a pair of input key frames. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving forward in time from a single input image) for key frame interpolation, i.e., to produce a video in between two input frames. We accomplish this adaptation through a lightweight fine-tuning technique that produces a version of the model that instead predicts videos moving backwards in time from a single input image. This model (along with the original forward-moving model) is subsequently used in a dual-directional diffusion sampling process that combines the overlapping model estimates starting from each of the two keyframes. Our experiments show that our method outperforms both existing diffusion-based methods and traditional frame interpolation techniques.
Forward citations
Cited by 7 Pith papers
-
Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping
For video face swapping, adaptively adding swapped anchor frames at the moments of worst identity drift should make synthetic training pairs more faithful than the current first-and-last-frame-only scheme.
-
CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics
CHIMERA performs zero-shot diffusion image morphing with adaptive multi-scale feature-cache injection (ACI) and VLM-generated semantic anchor prompting (SAP), and proposes a new morphing-quality metric, GLCS.
-
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
A $500 LoRA fine-tune of Wan2.1-T2V with per-frame random timesteps matches Wan-I2V's benchmark quality and adds zero-shot start-end and video-extension capabilities.
-
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.
-
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
Adding bidirectional cycle-consistent training with learnable direction tokens improves long-video interpolation quality without added inference cost.
-
Semantic Frame Interpolation
The authors define Semantic Frame Interpolation, build a 300k-clip dataset and benchmark, and propose a Mixture-of-LoRA adaptation of Wan2.1 that improves temporal smoothness but does not preserve the given start and ...
-
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
DiffuseSlide boosts the frame rate of latent diffusion videos via latent interpolation, noise re-injection, and sliding-window denoising, reporting better FVD, PSNR, and SSIM than several baselines on WebVid-10M.
Discussion (0). Continue with ORCID to comment.