REVIEW 7 cited by
Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A diffusion model learns to predict a vector field of gradients. We propose to apply chain rule on the learned gradients, and back-propagate the score of a diffusion model through the Jacobian of a differentiable renderer, which we instantiate to be a voxel radiance field. This setup aggregates 2D scores at multiple camera viewpoints into a 3D score, and repurposes a pretrained 2D model for 3D data generation. We identify a technical challenge of distribution mismatch that arises in this application, and propose a novel estimation mechanism to resolve it. We run our algorithm on several off-the-shelf diffusion image generative models, including the recently released Stable Diffusion trained on the large-scale LAION dataset.
Forward citations
Cited by 7 Pith papers
-
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.
-
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.
-
Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation
Sharp-It fine-tunes a multi-view diffusion model to enhance low-quality Shap-E renderings into high-quality multi-view sets that can be reconstructed into detailed 3D assets.
-
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.
-
SHaDe: Compact and Consistent Dynamic 3D Reconstruction via Tri-Plane Deformation and Latent Diffusion
SHaDe combines explicit tri-plane deformation, SH attention rendering, and latent diffusion refinement to reconstruct dynamic 3D scenes from sparse multi-view images and reports state-of-the-art results on D-NeRF.
-
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.
-
CRAFT: Designing Creative and Functional 3D Objects
A mesh deformation system that jointly optimizes semantic alignment with text or image prompts and body fit, producing body-aware 3D objects.
Discussion (0). Continue with ORCID to comment.