REVIEW 24 cited by
ImageDream: Image-Prompt Multi-view Diffusion for 3D Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce "ImageDream," an innovative image-prompt, multi-view diffusion model for 3D object generation. ImageDream stands out for its ability to produce 3D models of higher quality compared to existing state-of-the-art, image-conditioned methods. Our approach utilizes a canonical camera coordination for the objects in images, improving visual geometry accuracy. The model is designed with various levels of control at each block inside the diffusion model based on the input image, where global control shapes the overall object layout and local control fine-tunes the image details. The effectiveness of ImageDream is demonstrated through extensive evaluations using a standard prompt list. For more information, visit our project page at https://Image-Dream.github.io.
Forward citations
Cited by 24 Pith papers
-
AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
A feed-forward VAE plus rectified-flow model animates arbitrary static meshes from text prompts in seconds, with a new 4M-sequence training dataset.
-
MVGBench: Comprehensive Benchmark for Multi-view Generation Models
MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.
-
Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
A diffusion image-editing model conditioned on Plücker ray-map tokens and text-defined NOCS fronts generates high-fidelity novel views with absolute global pose control from unposed inputs.
-
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow
LATO.2 factorizes mesh generation into a vertex-generation flow and a vertex-conditioned connectivity flow, beating joint-latent and autoregressive baselines on geometric fidelity and connectivity quality.
-
MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image Structure
MarkSplatter embeds arbitrary messages into 3D Gaussian Splatting models in one forward pass via Splatter Image conversion, reaching about 94% bit-level accuracy at about 2.5 seconds per model.
-
CharacterShot: Controllable and Consistent 4D Character Animation
A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.
-
DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...
-
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.
-
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.
-
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
EarthCrafter generates 600-meter-scale 3D Earth scenes using separate latent diffusion models for structure and texture, conditioned on semantics, images, or nothing.
-
Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
A sliding iterative denoising scheme that alternates spatial and temporal passes, combined with skeleton conditioning, lets a diffusion model create spatio-temporally consistent multi-view human videos from sparse inp...
-
Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
A single-image 3D keypoint estimator trained without 3D labels, using multi-view diffusion-generated views and features as self-supervision.
-
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.
-
FlexPainter: Flexible and Multi-View Consistent Texture Generation
FlexPainter combines multi-view grid generation, UV-space view synchronization with a learned weighting network, and multi-modal embedding control to generate consistent, high-resolution textures from text and image prompts.
-
Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing
Pro3D-Editor chooses the most editing-salient view, propagates the edit to other key views with per-view LoRA experts, and refines the 3D scene, improving multi-view consistency.
-
Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
Direct3D-S2 uses a new Spatial Sparse Attention mechanism to train a sparse-volume diffusion transformer at 1024^3 resolution on 8 GPUs.
-
S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
A three-stage pipeline generates animatable 3D Gaussian head avatars from one image by diffusion-based splat synthesis, FLAME fitting, and inverse-distance binding with scale adaptation.
-
ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency
ObjFiller3D jointly optimizes a dense 360-degree ring of views to inpaint 3D objects with cross-view-consistent textures, reporting higher PSNR and LPIPS than per-view baselines at much lower runtime.
-
Collaborative Multi-Modal Coding for High-Quality 3D Generation
TriMM fuses RGB, RGB-D, and point-cloud encoding into a shared triplane latent space and generates 3D assets from a single image with a latent diffusion model.
-
EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction
A masked fine-tuned diffusion model simultaneously completes occluded views and synthesizes novel viewpoints, enabling fast feed-forward 3D reconstruction.
-
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.
-
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.
Discussion (0). Sign in to comment.