Pith. sign in

REVIEW 8 cited by

3D-GPT: Procedural 3D Modeling with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12945 v2 pith:CM7TXTY3 submitted 2023-10-19 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords d-gptmodelingagentproceduralgenerationintegratesllmscreation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the pursuit of efficient automated content creation, procedural generation, leveraging modifiable parameters and rule-based systems, emerges as a promising approach. Nonetheless, it could be a demanding endeavor, given its intricate nature necessitating a deep understanding of rules, algorithms, and parameters. To reduce workload, we introduce 3D-GPT, a framework utilizing large language models~(LLMs) for instruction-driven 3D modeling. 3D-GPT positions LLMs as proficient problem solvers, dissecting the procedural 3D modeling tasks into accessible segments and appointing the apt agent for each task. 3D-GPT integrates three core agents: the task dispatch agent, the conceptualization agent, and the modeling agent. They collaboratively achieve two objectives. First, it enhances concise initial scene descriptions, evolving them into detailed forms while dynamically adapting the text based on subsequent instructions. Second, it integrates procedural generation, extracting parameter values from enriched text to effortlessly interface with 3D software for asset creation. Our empirical investigations confirm that 3D-GPT not only interprets and executes instructions, delivering reliable results but also collaborates effectively with human designers. Furthermore, it seamlessly integrates with Blender, unlocking expanded manipulation possibilities. Our work highlights the potential of LLMs in 3D modeling, offering a basic framework for future advancements in scene generation and animation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ShapeLib: Designing a library of programmatic 3D shape abstractions with Large Language Models

    cs.CV 2025-02 conditional novelty 7.0 of 10

    ShapeLib guides LLMs, validated with geometric checks against a small seed set, to author reusable programmatic shape abstraction libraries that generalize to new 3D shapes.

  2. Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A clarification-first 3D agent, trained by simulated multi-turn dialogue, reaches 60.4% and 43.3% success on single- and multi-step 3D tool tasks, more than doubling prior baselines.

  3. LL3M: Large Language 3D Modelers

    cs.GR 2025-08 conditional novelty 6.0 of 10

    A multi-agent LLM system generates editable 3D assets as Blender Python code, using documentation retrieval and visual self-critique to refine results.

  4. IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.

  5. Conversational Interfaces for Parametric Conceptual Architectural Design: Integrating Mixed Reality with LLM-driven Interaction

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A conversational MR interface with three LLM agents converts speech and gestures into parametric modeling code; a 27-person user study reports lower barriers and successful compilation, though quantitative evidence is...

  6. LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation

    cs.AI 2025-09 reject novelty 5.0 of 10

    A multimodal LLM framework generates interactive Unreal-based 3D environments from text and height maps, claiming superior layout accuracy and over 90x faster production than manual methods.

  7. TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs

    cs.CV 2025-05 conditional novelty 5.0 of 10

    TACO couples thinking with final answers, filters unstable training samples, reweights easy or hard samples, and adds multi-scale test inference, improving LVLM visual reasoning accuracy over VLM-R1.

  8. RoomCraft: Controllable and Complete 3D Indoor Scene Generation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    RoomCraft generates 3D indoor scenes from text, sketches, or images by extracting structured furniture relations with a VLM and resolving placement conflicts with a weighted positioning heuristic.

Pith tools