REVIEW 13 cited by
MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundational vision models and lift the video data to a novel Motion Scaffold (MoSca) representation, which compactly and smoothly encodes the underlying motions/deformations. The scene geometry and appearance are then disentangled from the deformation field and are encoded by globally fusing the Gaussians anchored onto the MoSca and optimized via Gaussian Splatting. Additionally, camera focal length and poses can be solved using bundle adjustment without the need of any other pose estimation tools. Experiments demonstrate state-of-the-art performance on dynamic rendering benchmarks and its effectiveness on real videos.
Forward citations
Cited by 13 Pith papers
-
GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
Variational Gaussian Processes with spatio-temporal kernels supply probabilistic deformation priors to 4DGS, improving sparse-view reconstruction while yielding calibrated motion uncertainty and temporal extrapolation.
-
UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
UniWorld-View couples an occlusion-aware point cloud renderer with a dual-stream video diffusion model to synthesize large-baseline novel views from monocular video.
-
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.
-
DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs
A training-free pipeline that feeds MLLMs reprojected motion arrows plus a structured trace graph lifts 4D spatio-temporal QA accuracy on three benchmarks.
-
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
A decoder-only Transformer with causal attention and cached past-frame features performs incremental 3D reconstruction from streaming images, beating the RNN-based CUT3R on several benchmark metrics.
-
TRACE: Learning 3D Gaussian Physical Dynamics from Multi-view Videos
TRACE predicts future frames of dynamic 3D scenes by learning a per-particle translation-rotation dynamics system inside 3D Gaussian Splatting, without labels.
-
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
Video-rewinding joint training preserves geometry while re-animating a single-video scene with novel motion from a text prompt and an image-to-video model.
-
SpatialTrackerV2: 3D Point Tracking Made Easy
A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.
-
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
4D-Animal fits SMAL animal models to video using silhouette, part, pixel, and tracking losses from off-the-shelf 2D models, removing the need for sparse keypoint annotations.
-
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
A deformable Gaussian splatting framework with hierarchical rigid, skeleton-driven, and flow-based warping reconstructs dynamic scenes from long video captures with fast training and rendering.
-
Leveraging 2D Priors and SDF Guidance for Dynamic Urban Scene Rendering
UGSDF achieves state-of-the-art novel-view rendering of dynamic urban objects without LiDAR or 3D motion annotations by jointly optimizing SDFs and 3D Gaussians under 2D depth and point-tracking priors.
-
Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors
A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.
-
SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Grouping
Temporal sensitivity pruning plus grouped SE(3) motion distillation speeds up DeformableGS rendering by 6.78x to 13.71x and training by about 2.5x across 50 dynamic scenes in MonoDyGauBench.
Discussion (0). Sign in to comment.