Pith. sign in

REVIEW 13 cited by

MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17421 v2 pith:XWQSHMNN submitted 2024-05-27 cs.CV cs.GR

classification cs.CVcs.GR
keywords moscadynamicmotionvideosgaussiannovelscaffoldsadditionally
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundational vision models and lift the video data to a novel Motion Scaffold (MoSca) representation, which compactly and smoothly encodes the underlying motions/deformations. The scene geometry and appearance are then disentangled from the deformation field and are encoded by globally fusing the Gaussians anchored onto the MoSca and optimized via Gaussian Splatting. Additionally, camera focal length and poses can be solved using bundle adjustment without the need of any other pose estimation tools. Experiments demonstrate state-of-the-art performance on dynamic rendering benchmarks and its effectiveness on real videos.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Variational Gaussian Processes with spatio-temporal kernels supply probabilistic deformation priors to 4DGS, improving sparse-view reconstruction while yielding calibrated motion uncertainty and temporal extrapolation.

  2. UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.

  3. 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.

  4. DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A training-free pipeline that feeds MLLMs reprojected motion arrows plus a structured trace graph lifts 4D spatio-temporal QA accuracy on three benchmarks.

  5. STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A decoder-only Transformer with causal attention and cached past-frame features performs incremental 3D reconstruction from streaming images, beating the RNN-based CUT3R on several benchmark metrics.

  6. TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view Videos

    cs.CV 2025-08 conditional novelty 6.0 of 10

    TRACE predicts future frames of dynamic 3D scenes by learning a per-particle translation-rotation dynamics system inside 3D Gaussian Splatting, without labels.

  7. Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Video-rewinding joint training preserves geometry while re-animating a single-video scene with novel motion from a text prompt and an image-to-video model.

  8. SpatialTrackerV2: 3D Point Tracking Made Easy

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.

  9. 4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    4D-Animal fits SMAL animal models to video using silhouette, part, pixel, and tracking losses from off-the-shelf 2D models, removing the need for sparse keypoint annotations.

  10. HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A deformable Gaussian splatting framework with hierarchical rigid, skeleton-driven, and flow-based warping reconstructs dynamic scenes from long video captures with fast training and rendering.

  11. Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.

  12. Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.

  13. SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Grouping

    cs.GR 2025-06 conditional novelty 5.0 of 10

    Temporal sensitivity pruning plus grouped SE(3) motion distillation speeds up DeformableGS rendering by 6.78x to 13.71x and training by about 2.5x across 50 dynamic scenes in MonoDyGauBench.

Pith tools