Pith. sign in

REVIEW 10 cited by

CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18613 v2 pith:WMZR3532 submitted 2024-11-27 cs.CV

classification cs.CV
keywords videocat4dmulti-viewnoveldiffusiondynamicmodelmonocular
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera poses and timestamps. Combined with a novel sampling approach, this model can transform a single monocular video into a multi-view video, enabling robust 4D reconstruction via optimization of a deformable 3D Gaussian representation. We demonstrate competitive performance on novel view synthesis and dynamic scene reconstruction benchmarks, and highlight the creative capabilities for 4D scene generation from real or generated videos. See our project page for results and interactive demos: https://cat-4d.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoMat: Extracting PBR Materials from Video Diffusion Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    VideoMat uses a finetuned video diffusion model, intrinsic decomposition, and differentiable path tracing to extract PBR material maps for known 3D geometry from text or image prompts.

  2. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.

  3. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  4. AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Texture-aware SuperCluster pruning plus an adaptive Gaussian head lets feed-forward 3DGS models hit a user budget β while outperforming post-hoc pruners on RE10K, ACID, DL3DV and DTU.

  5. InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A training-free, near-zero-overhead inverse solver for novel-view video generation and inpainting that projects masks into continuous multi-channel latent masks and applies DDS with conjugate gradient in latent space.

  6. Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...

  7. Voyaging into Perpetual Dynamic Scenes from a Single View

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-view dynamic scene can be extended into an unbounded fly-through video by iteratively outpainting partial views of a learned 4D point cloud with ray distance guidance.

  8. ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular Inputs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion-enhanced 4D reconstruction pipeline for monocular video that achieves state-of-the-art results on DyCheck by supervising Gaussian splatting with personalized diffusion-generated pseudo-views.

  9. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  10. Restereo: Diffusion stereo video generation and restoration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion model fine-tuned on synthetically degraded stereo videos simultaneously generates a consistent stereo pair and restores low-resolution or compressed input, outperforming prior stereo generators on low-qual...

Pith tools