Pith. sign in

REVIEW 11 cited by

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15008 v3 pith:GYCG4PRJ submitted 2023-10-23 cs.CV

classification cs.CV
keywords cross-domaindiffusionmulti-viewconsistencyefficiencygeometryhigh-qualityimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images.Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works directly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of image-to-3D tasks, we propose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure consistency, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a geometry-aware normal fusion algorithm that extracts high-quality surfaces from the multi-view 2D representations. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and reasonably good efficiency compared to prior works.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  2. SILICA: Repurposing Diffusion Priors for Joint Glass Segmentation and Depth Estimation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CLIP-conditioned single-step diffusion regression jointly does glass segmentation and affine-invariant depth, then aligns metric depth by masking bad sensor returns, beating prior glass methods on Mirage 18k and public sets.

  3. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  4. SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SOV-CAD recovers CAD modeling sequences from orthographic images via Decision-Transformer offline RL conditioned on stepwise three-views plus sketch canvas and IoU-based rewards, beating holistic baselines with better...

  5. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.

  6. CLONE: Continuous Latent Optimization for Normal Estimation via 3D Gaussian Splatting

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    CLONE estimates surface normals from images via a differentiable 3D Gaussian splatting loop, using photometric loss instead of normal labels.

  7. Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PointSD uses a frozen Stable Diffusion model, conditioned on point clouds through rendered images, to generate training targets for point cloud self-supervised learning.

  8. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...

  9. NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video diffusion model fine-tuned to output both color and normal maps, aligned by a geometry-temporal attention block, reconstructs textured 3D meshes from a single image.

  10. Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.

  11. Advances in 4D Representation: Geometry, Motion, and Interaction

    cs.CV 2025-10 conditional novelty 4.0 of 10

    A representation-centric survey of 4D generation and reconstruction, organized by geometry, motion, and interaction, with qualitative trade-off comparisons across seven representation families.

Pith tools