REVIEW 6 cited by
Text-Guided Texturing by Synchronized Multi-View Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces a novel approach to synthesize texture to dress up a given 3D object, given a text prompt. Based on the pretrained text-to-image (T2I) diffusion model, existing methods usually employ a project-and-inpaint approach, in which a view of the given object is first generated and warped to another view for inpainting. But it tends to generate inconsistent texture due to the asynchronous diffusion of multiple views. We believe such asynchronous diffusion and insufficient information sharing among views are the root causes of the inconsistent artifact. In this paper, we propose a synchronized multi-view diffusion approach that allows the diffusion processes from different views to reach a consensus of the generated content early in the process, and hence ensures the texture consistency. To synchronize the diffusion, we share the denoised content among different views in each denoising step, specifically blending the latent content in the texture domain from views with overlap. Our method demonstrates superior performance in generating consistent, seamless, highly detailed textures, comparing to state-of-the-art methods.
Forward citations
Cited by 6 Pith papers
-
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
EmbodiedGen is a modular pipeline that generates physically annotated 3D assets and scenes in URDF format from images or text, for direct use in robot simulators.
-
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces
StochSync generates images on arbitrary surfaces such as spheres and meshes by alternating non-overlapping denoised views, maximum stochasticity, and multi-step clean-image prediction from a pretrained diffusion model.
-
ArtNVG: Content-Style Separated Artistic Neighboring-View Gaussian Stylization
ArtNVG combines CSGO-style content/style separation with neighboring-view attention sharing to produce locally consistent stylized 3D Gaussian Splatting scenes from a single style reference image.
-
MCMat: Multiview-Consistent and Physically Accurate PBR Material Generation
A two-stage diffusion transformer pipeline generates multi-view-consistent, relightable PBR material maps for 3D meshes, reporting state-of-the-art FID/KID scores on 70 Objaverse models and higher user-study ratings t...
-
RoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing
A two-stage, zero-shot diffusion pipeline that textures room-scale meshes with global style consistency and per-instance repainting.
-
Make-A-Texture: Fast Shape-Aware Texture Generation in 3 Seconds
A texture-generation pipeline that produces 1024x1024 textures from text in 3.07 seconds on an H100, with quality comparable to SyncMVD and other prior methods.
Discussion (0). Continue with ORCID to comment.