Pith. sign in

REVIEW 7 cited by

CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13843 v5 pith:4DBBQNEQ submitted 2023-03-24 cs.CV

classification cs.CV
keywords scenemulti-objectlayoutcomponerfconsistencydiffusioneditableguidance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-3D form plays a crucial role in creating editable 3D scenes for AR/VR. Recent advances have shown promise in merging neural radiance fields (NeRFs) with pre-trained diffusion models for text-to-3D object generation. However, one enduring challenge is their inadequate capability to accurately parse and regenerate consistent multi-object environments. Specifically, these models encounter difficulties in accurately representing quantity and style prompted by multi-object texts, often resulting in a collapse of the rendering fidelity that fails to match the semantic intricacies. Moreover, amalgamating these elements into a coherent 3D scene is a substantial challenge, stemming from generic distribution inherent in diffusion models. To tackle the issue of 'guidance collapse' and further enhance scene consistency, we propose a novel framework, dubbed CompoNeRF, by integrating an editable 3D scene layout with object-specific and scene-wide guidance mechanisms. It initiates by interpreting a complex text into the layout populated with multiple NeRFs, each paired with a corresponding subtext prompt for precise object depiction. Next, a tailored composition module seamlessly blends these NeRFs, promoting consistency, while the dual-level text guidance reduces ambiguity and boosts accuracy. Noticeably, our composition design permits decomposition. This enables flexible scene editing and recomposition into new scenes based on the edited layout or text prompts. Utilizing the open-source Stable Diffusion model, CompoNeRF generates multi-object scenes with high fidelity. Remarkably, our framework achieves up to a \textbf{54\%} improvement by the multi-view CLIP score metric. Our user study indicates that our method has significantly improved semantic accuracy, multi-view consistency, and individual recognizability for multi-object scene generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Sat2City generates explicit 3D city geometry and appearance from a height-map condition using cascaded latent diffusion on sparse voxel grids, beating prior methods on a new synthetic city dataset.

  2. PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PartCrafter generates several separable 3D part meshes at once from a single image by fine-tuning a pretrained 3D diffusion transformer with part identity tokens and local-global attention.

  3. ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-free pipeline that generates editable 3D scenes from text by using a generated 2D image as an intermediary to extract object shapes, appearances, positions, and poses.

  4. UrbanCraft: Urban View Extrapolation via Hierarchical Sem-Geometric Priors

    cs.CV 2025-05 reject novelty 6.0 of 10

    UrbanCraft uses hierarchical semantic-geometric priors to condition diffusion-based score distillation, enabling extrapolated view synthesis for urban 3D Gaussian Splatting scenes.

  5. NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

    cs.CV 2025-07 conditional novelty 4.0 of 10

    NeuroVoxel-LM combines dynamic multi-resolution voxelization with attention-based pooling of NeRF weights, reporting faster 3D feature extraction and modestly better NeRF captioning than fixed-resolution and max-pooli...

  6. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

  7. From 2D to 3D Cognition: A Brief Survey of General World Models

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.

Pith tools