REVIEW 3 cited by
SceneWiz3D: Towards Text-guided 3D Scene Composition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We are witnessing significant breakthroughs in the technology for generating 3D objects from text. Existing approaches either leverage large text-to-image models to optimize a 3D representation or train 3D generators on object-centric datasets. Generating entire scenes, however, remains very challenging as a scene contains multiple 3D objects, diverse and scattered. In this work, we introduce SceneWiz3D, a novel approach to synthesize high-fidelity 3D scenes from text. We marry the locality of objects with globality of scenes by introducing a hybrid 3D representation: explicit for objects and implicit for scenes. Remarkably, an object, being represented explicitly, can be either generated from text using conventional text-to-3D approaches, or provided by users. To configure the layout of the scene and automatically place objects, we apply the Particle Swarm Optimization technique during the optimization process. Furthermore, it is difficult for certain parts of the scene (e.g., corners, occlusion) to receive multi-view supervision, leading to inferior geometry. We incorporate an RGBD panorama diffusion model to mitigate it, resulting in high-quality geometry. Extensive evaluation supports that our approach achieves superior quality over previous approaches, enabling the generation of detailed and view-consistent 3D scenes.
Forward citations
Cited by 3 Pith papers
-
3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
A self-improving vision-language-model policy iteratively crafts 3D environments from text, and renderings of those environments serve as effective synthetic pretraining data for vision models.
-
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
A training-free pipeline that generates editable 3D scenes from text by using a generated 2D image as an intermediary to extract object shapes, appearances, positions, and poses.
-
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.
Discussion (0). Continue with ORCID to comment.