Pith. sign in

REVIEW 7 cited by

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15239 v2 pith:6LFKFBJU submitted 2024-08-27 cs.CV

classification cs.CV
keywords modelinputinterpolationdiffusionframeframesimageimage-to-video
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a method for generating video sequences with coherent motion between a pair of input key frames. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving forward in time from a single input image) for key frame interpolation, i.e., to produce a video in between two input frames. We accomplish this adaptation through a lightweight fine-tuning technique that produces a version of the model that instead predicts videos moving backwards in time from a single input image. This model (along with the original forward-moving model) is subsequently used in a dual-directional diffusion sampling process that combines the overlapping model estimates starting from each of the two keyframes. Our experiments show that our method outperforms both existing diffusion-based methods and traditional frame interpolation techniques.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping

    cs.CV 2026-07 conditional novelty 6.0 of 10

    For video face swapping, adaptively adding swapped anchor frames at the moments of worst identity drift should make synthetic training pairs more faithful than the current first-and-last-frame-only scheme.

  2. CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics

    cs.CV 2025-12 conditional novelty 6.0 of 10

    CHIMERA performs zero-shot diffusion image morphing with adaptive multi-scale feature-cache injection (ACI) and VLM-generated semantic anchor prompting (SAP), and proposes a new morphing-quality metric, GLCS.

  3. Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A $500 LoRA fine-tune of Wan2.1-T2V with per-frame random timesteps matches Wan-I2V's benchmark quality and adds zero-shot start-end and video-extension capabilities.

  4. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MultiTalk is the first framework to generate multi-person conversational videos from multi-stream audio, using Label Rotary Position Embedding to bind each voice to the correct person.

  5. Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation

    cs.CV 2026-04 conditional novelty 5.0 of 10

    Adding bidirectional cycle-consistent training with learnable direction tokens improves long-video interpolation quality without added inference cost.

  6. Semantic Frame Interpolation

    cs.CV 2025-07 reject novelty 5.0 of 10

    The authors define Semantic Frame Interpolation, build a 300k-clip dataset and benchmark, and propose a Mixture-of-LoRA adaptation of Wan2.1 that improves temporal smoothness but does not preserve the given start and ...

  7. DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DiffuseSlide boosts the frame rate of latent diffusion videos via latent interpolation, noise re-injection, and sliding-window denoising, reporting better FVD, PSNR, and SSIM than several baselines on WebVid-10M.

Pith tools