REVIEW 13 cited by
LDM3D: Latent Diffusion Model for 3D
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This research paper proposes a Latent Diffusion Model for 3D (LDM3D) that generates both image and depth map data from a given text prompt, allowing users to generate RGBD images from text prompts. The LDM3D model is fine-tuned on a dataset of tuples containing an RGB image, depth map and caption, and validated through extensive experiments. We also develop an application called DepthFusion, which uses the generated RGB images and depth maps to create immersive and interactive 360-degree-view experiences using TouchDesigner. This technology has the potential to transform a wide range of industries, from entertainment and gaming to architecture and design. Overall, this paper presents a significant contribution to the field of generative AI and computer vision, and showcases the potential of LDM3D and DepthFusion to revolutionize content creation and digital experiences. A short video summarizing the approach can be found at https://t.ly/tdi2.
Forward citations
Cited by 13 Pith papers
-
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
DiffSplat repurposes image diffusion models to generate multi-view Gaussian splat grids, using a rendering loss for 3D consistency and achieving state-of-the-art text- and image-conditioned 3D generation.
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
360Anything: Geometry-Free Lifting of Images and Videos to 360{\deg}
360Anything lifts perspective images and videos to 360° panoramas with a diffusion transformer and sequence concatenation, requiring no camera metadata at test time.
-
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.
-
DreamCube: 3D Panorama Generation via Multi-plane Synchronization
A synchronized multi-plane adaptation of 2D diffusion operators enables seam-consistent cubemap generation, and DreamCube extends this to joint RGB-D panorama generation and 3D scene lifting.
-
JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers
JointDiT is a Flux-based diffusion transformer that jointly generates RGB images and depth maps, and also handles depth estimation and depth-to-image generation by controlling the noise level of each branch.
-
Latent Radiance Fields with 3D-aware 2D Representations
A three-stage pipeline makes VAE latent codes 3D-consistent and builds a latent radiance field, improving photorealistic novel-view synthesis in latent space.
-
PhysMotion: Physics-Grounded Dynamics From a Single Image
PhysMotion generates physically plausible videos from a single image by simulating 3D object motion with a material point method, then enhancing the rendering with a diffusion model.
-
Direct and Explicit 3D Generation from a Single Image
A modified Stable Diffusion model generates six views of depth, color, and 3D Gaussian features from one image, then lifts them into a textured mesh or splatted scene in 15 to 25 seconds.
-
InsTex: Indoor Scenes Stylized Texture Synthesis
InsTex generates style-consistent textures for indoor 3D scenes using a coarse-to-fine diffusion pipeline with global image guidance, reporting faster and higher-scoring results than four baselines.
-
Joint Learning of Depth and Appearance for Portrait Image Animation
A single diffusion model jointly generates portrait RGB images and aligned depth maps, and its fine-tuned variants can estimate depth, edit from depth, relight, and produce audio-driven talking heads with depth.
-
Nested Annealed Training Scheme for Generative Adversarial Networks
NATS, a nested annealed training scheme for GANs, is claimed to improve FID/IS on CIFAR10, LSUN, CelebA, and ImageNet64, but its key theoretical justification is not provided in the preprint.
-
3D Scene Generation: A Survey
The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.
Discussion (0). Continue with ORCID to comment.