Pith. sign in

REVIEW 3 cited by

VFIMamba: Video Frame Interpolation with State Space Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02315 v2 pith:FKYOGWWP submitted 2024-07-02 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelingframesacrossframeinter-frameinterpolationmodelsvfimamba
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Inter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based models, which often either lack sufficient receptive fields or entail significant computational overheads. Recently, Selective State Space Models (S6) have emerged, tailored specifically for long sequence modeling, offering both linear complexity and data-dependent modeling capabilities. In this paper, we propose VFIMamba, a novel frame interpolation method for efficient and dynamic inter-frame modeling by harnessing the S6 model. Our approach introduces the Mixed-SSM Block (MSB), which initially rearranges tokens from adjacent frames in an interleaved fashion and subsequently applies multi-directional S6 modeling. This design facilitates the efficient transmission of information across frames while upholding linear complexity. Furthermore, we introduce a novel curriculum learning strategy that progressively cultivates proficiency in modeling inter-frame dynamics across varying motion magnitudes, fully unleashing the potential of the S6 model. Experimental findings showcase that our method attains state-of-the-art performance across diverse benchmarks, particularly excelling in high-resolution scenarios. In particular, on the X-TEST dataset, VFIMamba demonstrates a noteworthy improvement of 0.80 dB for 4K frames and 0.96 dB for 2K frames.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    TLB-VFI reports state-of-the-art FID, LPIPS, and FloLPIPS on VFI benchmarks using a latent Brownian bridge diffusion model with temporal-aware encoding and wavelet gating.

  2. LumosFlow: Motion-Guided Long Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    LumosFlow generates long videos by combining large-motion key frame generation, latent optical flow diffusion, and a ControlNet-style refinement module.

  3. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

Pith tools