REVIEW 22 cited by
Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Reconstructing dynamic 3D scenes from 2D images and generating diverse views over time is challenging due to scene complexity and temporal dynamics. Despite advancements in neural implicit models, limitations persist: (i) Inadequate Scene Structure: Existing methods struggle to reveal the spatial and temporal structure of dynamic scenes from directly learning the complex 6D plenoptic function. (ii) Scaling Deformation Modeling: Explicitly modeling scene element deformation becomes impractical for complex dynamics. To address these issues, we consider the spacetime as an entirety and propose to approximate the underlying spatio-temporal 4D volume of a dynamic scene by optimizing a collection of 4D primitives, with explicit geometry and appearance modeling. Learning to optimize the 4D primitives enables us to synthesize novel views at any desired time with our tailored rendering routine. Our model is conceptually simple, consisting of a 4D Gaussian parameterized by anisotropic ellipses that can rotate arbitrarily in space and time, as well as view-dependent and time-evolved appearance represented by the coefficient of 4D spherindrical harmonics. This approach offers simplicity, flexibility for variable-length video and end-to-end training, and efficient real-time rendering, making it suitable for capturing complex dynamic scene motions. Experiments across various benchmarks, including monocular and multi-view scenarios, demonstrate our 4DGS model's superior visual quality and efficiency.
Forward citations
Cited by 22 Pith papers
-
ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment
ASTRA jointly estimates camera time offsets and dynamic Gaussian geometry by aligning projected 3D motion with observed 2D trajectory tracks, improving robustness to large asynchrony.
-
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.
-
DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization
DeGS restructures 3DGS rendering into span parsing, task reorganization, and dense blending stages, achieving 1.8x-7.2x speedup and >80% scaling utilization over prior 3DGS accelerators.
-
4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.
-
ECoNGS: Efficient Compressive Neural Gaussian Splats for Volume Visualization
ECoNGS compresses volume-visualization scenes into entropy-coded neural Gaussian splats that are up to 6x smaller, train up to 6x faster, and render more accurately than the prior iVR-GS method.
-
ChronoGS: Disentangling Invariants and Changes in Multi-Period Scenes
A single shared Gaussian scaffold with per-period features and opacity gating reconstructs multi-period scenes better than static and dynamic baselines on a new 12-scene benchmark.
-
Style4D-Bench: A Benchmark Suite for 4D Stylization
Style4D-Bench introduces a 12-metric evaluation protocol and a 4DGS-based baseline, Style4D, claimed to achieve state-of-the-art 4D stylization.
-
Laplacian Analysis Meets Dynamics Modelling: Gaussian Splatting for 4D Reconstruction
A Laplacian-enhanced hybrid encoding method for 4D Gaussian Splatting that claims better reconstruction fidelity for dynamic scenes.
-
High-fidelity 3D Gaussian Inpainting: preserving multi-view consistency and photorealistic details
A 3D Gaussian inpainting framework with automatic mask refinement and depth-initialized uncertainty weighting balances multi-view consistency and visual detail, reporting the best LPIPS on the SPIn-NeRF dataset.
-
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.
-
Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation
Vid2Sim recovers 3D geometry, appearance, and elastic material parameters from multi-view videos using a feed-forward network plus a fast refinement, enabling mesh-free reduced-order simulation.
-
FreeTimeGS: Free Gaussian Primitives at Anytime and Anywhere for Dynamic Scene Reconstruction
A dynamic-scene representation where Gaussian primitives live freely in 4D space-time with linear motion and Gaussian time windows achieves state-of-the-art novel-view quality on complex-motion benchmarks.
-
Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features
A video-to-4D generation method that decouples dynamic and static features in DINOv2 space and fuses similar dynamic information across views reports state-of-the-art scores on Consistent4D and Objaverse.
-
3D Gaussian Representations with Motion Trajectory Field for Dynamic Scene Reconstruction
A 3D Gaussian Splatting model whose Gaussian centers are represented as a learned combination of shared global motion bases recovers dynamic scenes and motion trajectories from monocular video.
-
SD-GS: Structured Deformable 3D Gaussians for Efficient Dynamic Scene Reconstruction
SD-GS combines anchor-based 3D Gaussians with a deformation field and a deformation-aware densification strategy to reconstruct dynamic scenes more compactly and faster than prior 4D Gaussian methods.
-
LocalDyGS: Multi-view Global Dynamic Scene Modeling via Adaptive Local Implicit Feature Decoupling
LocalDyGS reconstructs dynamic scenes by decomposing space into seed-based local regions and generating time-varying Temporal Gaussians, though its claim of being first for large-scale scenes omits the existing Swift4...
-
RoboPearls: Editable Video Simulation for Robot Manipulation
RoboPearls is a 3D Gaussian Splatting based framework that edits demonstration videos into varied photorealistic simulations, and training on them improves robot manipulation success rates on RLBench and COLOSSEUM.
-
SkinningGS: Editable Dynamic Human Scene Reconstruction Using Gaussian Splatting Based on a Skinning Model
A UV-texture-driven Gaussian splatting avatar method claims faster, leaner, and better human-scene reconstruction than HUGS, but its tables contain internal inconsistencies.
-
UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery using Gaussian Splatting
UAV4D reconstructs 4D scenes from monocular drone video by fitting a single global scale to align human meshes with the background mesh, then renders with separate Gaussian splats.
-
SuperGS: Consistent and Detailed 3D Super-Resolution Scene Reconstruction via Gaussian Splatting
SuperGS outperforms prior Gaussian-splatting methods on high-resolution novel view synthesis by combining a latent feature field, multi-view voting densification, and variational uncertainty weighting.
-
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
A Gaussian-splatting method densifies sparse points and combines object and camera motion models to produce sharp novel views from blurry monocular video.
-
Advances and Trends in the 3D Reconstruction of the Shape and Motion of Animals
A structured review of 3D animal reconstruction covering explicit, parametric, implicit, and Gaussian splatting representations, with a comparison of six methods and a dataset overview.
Discussion (0). Continue with ORCID to comment.