Pith. sign in

REVIEW 3 cited by

Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21333 v1 pith:OZHBEUGU submitted 2024-07-31 cs.CV

classification cs.CV
keywords layoutfurnituremllmsinteractivevisualagentgenerationprompting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic furniture layout is long desired for convenient interior design. Leveraging the remarkable visual reasoning capabilities of multimodal large language models (MLLMs), recent methods address layout generation in a static manner, lacking the feedback-driven refinement essential for interactive user engagement. We introduce Chat2Layout, a novel interactive furniture layout generation system that extends the functionality of MLLMs into the realm of interactive layout design. To achieve this, we establish a unified vision-question paradigm for in-context learning, enabling seamless communication with MLLMs to steer their behavior without altering model weights. Within this framework, we present a novel training-free visual prompting mechanism. This involves a visual-text prompting technique that assist MLLMs in reasoning about plausible layout plans, followed by an Offline-to-Online search (O2O-Search) method, which automatically identifies the minimal set of informative references to provide exemplars for visual-text prompting. By employing an agent system with MLLMs as the core controller, we enable bidirectional interaction. The agent not only comprehends the 3D environment and user requirements through linguistic and visual perception but also plans tasks and reasons about actions to generate and arrange furniture within the virtual space. Furthermore, the agent iteratively updates based on visual feedback from execution results. Experimental results demonstrate that our approach facilitates language-interactive generation and arrangement for diverse and complex 3D furniture.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KinemaFX: A Kinematic-Driven Interactive System for Particle Effect Exploration and Customization

    cs.HC 2025-07 conditional novelty 6.0 of 10

    A kinematic-aware search and exploration system helps non-experts create customized particle effect artworks through text, simple shapes, motion paths, and implicit preference-guided iteration.

  2. Handle-based Mesh Deformation Guided By Vision Language Model

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A VLM selects deformation handles and target positions, and multi-view voting produces a text-guided mesh deformation with low distortion.

  3. WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A text-to-image pipeline that lets users restyle individual characters or regions of artistic typography and iteratively refine them with region-specific prompts.

Pith tools