Pith. sign in

REVIEW 8 cited by

Generic 3D Diffusion Adapter Using Controlled Multi-View Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12032 v2 pith:NL7QRLHV submitted 2024-03-18 cs.CV cs.GR

classification cs.CVcs.GR
keywords diffusionmulti-viewmveditqualitysynthesisviewsachievesadapter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-domain 3D object synthesis has been lagging behind image synthesis due to limited data and higher computational complexity. To bridge this gap, recent works have investigated multi-view diffusion but often fall short in either 3D consistency, visual quality, or efficiency. This paper proposes MVEdit, which functions as a 3D counterpart of SDEdit, employing ancestral sampling to jointly denoise multi-view images and output high-quality textured meshes. Built on off-the-shelf 2D diffusion models, MVEdit achieves 3D consistency through a training-free 3D Adapter, which lifts the 2D views of the last timestep into a coherent 3D representation, then conditions the 2D views of the next timestep using rendered views, without uncompromising visual quality. With an inference time of only 2-5 minutes, this framework achieves better trade-off between quality and speed than score distillation. MVEdit is highly versatile and extendable, with a wide range of applications including text/image-to-3D generation, 3D-to-3D editing, and high-quality texture synthesis. In particular, evaluations demonstrate state-of-the-art performance in both image-to-3D and text-guided texture generation tasks. Additionally, we introduce a method for fine-tuning 2D latent diffusion models on small 3D datasets with limited resources, enabling fast low-resolution text-to-3D initialization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

    cs.GR 2026-07 conditional novelty 6.5 of 10

    Automatic view scheduling via a directed generation graph plus object-level identity and adherence conditioning enables high-quality outdoor 3DGS scenes from arbitrary input geometry without user camera paths.

  2. Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A unified 3D multimodal model combines understanding, text-to-3D generation, instruction-guided editing, and part generation in one architecture, trained on an 87M-sample corpus, with claimed state-of-the-art results.

  3. TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Per-token tangent-space steering, with strength set by velocity-direction mismatch, improves localized training-free 3D editing over global-scaling baselines.

  4. EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.

  5. Articulate3D: Zero-Shot Text-Driven 3D Object Posing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A zero-shot pipeline that reposes 3D meshes by generating text-conditioned target images with rewired multi-view attention and aligning mesh keypoints to them.

  6. Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage diffusion model generates realistic depth noise on synthetic CAD data, and pretraining 3D networks on the resulting data improves few-shot real-world 3D tasks.

  7. Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing

    cs.GR 2025-05 conditional novelty 6.0 of 10

    Pro3D-Editor chooses the most editing-salient view, propagates the edit to other key views with per-view LoRA experts, and refines the 3D scene, improving multi-view consistency.

  8. Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    C33D blends a 3D model with an object category by generating a fused front view, then using texture and shape multi-view diffusion plus adaptive inversion to reconstruct a novel, consistent 3D model.

Pith tools