Pith. sign in

REVIEW 25 cited by

DreamGaussian4D: Generative 4D Gaussian Splatting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.17142 v3 pith:PHMKCN6D submitted 2023-12-28 cs.CV cs.GR

classification cs.CVcs.GR
keywords generationgaussiandg4ddreamgaussian4defficientframeworkgeneratedmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

4D content generation has achieved remarkable progress recently. However, existing methods suffer from long optimization times, a lack of motion controllability, and a low quality of details. In this paper, we introduce DreamGaussian4D (DG4D), an efficient 4D generation framework that builds on Gaussian Splatting (GS). Our key insight is that combining explicit modeling of spatial transformations with static GS makes an efficient and powerful representation for 4D generation. Moreover, video generation methods have the potential to offer valuable spatial-temporal priors, enhancing the high-quality 4D generation. Specifically, we propose an integral framework with two major modules: 1) Image-to-4D GS - we initially generate static GS with DreamGaussianHD, followed by HexPlane-based dynamic generation with Gaussian deformation; and 2) Video-to-Video Texture Refinement - we refine the generated UV-space texture maps and meanwhile enhance their temporal consistency by utilizing a pre-trained image-to-video diffusion model. Notably, DG4D reduces the optimization time from several hours to just a few minutes, allows the generated 3D motion to be visually controlled, and produces animated meshes that can be realistically rendered in 3D engines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BANG: Dividing 3D Assets via Generative Exploded Dynamics

    cs.GR 2025-07 conditional novelty 7.0 of 10

    A diffusion-based method that generates smooth exploded-view sequences of 3D objects, enabling part-level decomposition, control, and reassembly.

  2. AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A feed-forward VAE plus rectified-flow model animates arbitrary static meshes from text prompts in seconds, with a new 4M-sequence training dataset.

  3. 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on 60,000 fitted 4D Gaussian Splatting human clips generates text-prompted, view-consistent dynamic humans directly in 4D, over 10x faster than video-first pipelines.

  4. AniGS: Bridging Rendering and Diffusion Prior for 3D Scene Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AniGS animates a static 3D Gaussian Splatting scene by iteratively distilling video-diffusion motion into a time-conditioned deformation field while keeping static regions fixed.

  5. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  6. SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Weakly supervised fitting recovers consistent pseudo-skeletons from mesh sequences, then a text-conditioned transformer with Motion-GRPO generates editable skeleton-driven 4D mesh animations.

  7. MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A geometry-consistent multi-view RGBD 4D world model for robot manipulation, whose actions are recovered by test-time optimization of a learned trajectory latent, outperforming single- and dual-view world-model baseli...

  8. SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation

    cs.RO 2026-02 conditional novelty 6.0 of 10

    SoMA couples robot joint actions, environmental forces, and learned Gaussian-splat dynamics into a single neural simulator, improving resimulation and generalization on real robot soft-body manipulation by about 20% o...

  9. Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion

    cs.GR 2025-09 conditional novelty 6.0 of 10

    A diffusion model over per-person UNet weights generates new dynamic human avatars that render pose-dependent 3D Gaussians in real time.

  10. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  11. 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.

  12. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  13. Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...

  14. Voyaging into Perpetual Dynamic Scenes from a Single View

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single-view dynamic scene can be extended into an unbounded fly-through video by iteratively outpainting partial views of a learned 4D point cloud with ray distance guidance.

  15. Efficient multi-view training for 3D Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Training 3D Gaussian Splatting with multiple images per iteration, using partial rendering and a 3D-aware SSIM loss, improves novel-view synthesis quality over single-view training.

  16. RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation

    cs.RO 2025-10 unverdicted novelty 5.0 of 10

    Abstract describes RoDyn but full text describes iMoWM; the record is internally inconsistent and the headline claims are absent from the body.

  17. Align 3D Representation and Text Embedding for 3D Content Personalization

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Invert3D projects 3D content into a CLIP-style text-aligned embedding, allowing text-prompt personalization without per-scene retraining.

  18. DIP-GS: Deep Image Prior For Gaussian Splatting Sparse View Recovery

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DIP-GS applies a deep image prior in a coarse-to-fine manner to enable 3D Gaussian Splatting for sparse-view reconstruction.

  19. TextMesh4D: Zero-shot Text-to-4D Mesh Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.

  20. Generative 4D Scene Gaussian Splatting with Object View-Synthesis Priors

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A test-time optimization method that jointly fits deformable per-object 3D Gaussians with object-centric diffusion priors to generate 4D scenes and point tracks from monocular multi-object videos.

  21. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A video-conditioned latent diffusion model generates mesh vertex trajectories that deform an input 3D asset into render-ready 4D animations.

  22. CTRL-GS: Cascaded Temporal Residue Learning for 4D Gaussian Splatting

    cs.CV 2025-05 conditional novelty 5.0 of 10

    CTRL-GS represents dynamic Gaussian scenes as cascaded video-segment-frame residuals, improving reconstruction quality over 4D-GS on several dynamic-view benchmarks.

  23. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

  24. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

  25. DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes

    cs.CV 2025-08 conditional novelty 4.0 of 10

    DrivingGaussian++ reconstructs dynamic surround-view driving scenes and performs training-free multi-task editing (weather, texture, object manipulation) using Gaussians, diffusion models, and LLM-generated trajectories.

Pith tools