REVIEW 3 cited by
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views, especially when there is a significant camera pose difference, leading to poor-quality 3D geometries and textures. We attribute this issue to their treatment of all target views with equal priority according to our empirical observation that the target views closer to the input views exhibit higher fidelity. With this inspiration, we propose AR-1-to-3, a novel next-view prediction paradigm based on diffusion models that first generates views close to the input views, which are then utilized as contextual information to progressively synthesize farther views. To encode the generated view subsequences as local and global conditions for the next-view prediction, we accordingly develop a stacked local feature encoding strategy (Stacked-LE) and an LSTM-based global feature encoding strategy (LSTM-GE). Extensive experiments demonstrate that our method significantly improves the consistency between the generated views and the input views, producing high-fidelity 3D assets.
Forward citations
Cited by 3 Pith papers
-
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
A masked discrete-diffusion transformer generates multiple consistent object views from a single image or text, reporting the best average PSNR/SSIM/LPIPS on GSO and 3D-FUTURE.
-
SDMatte: Grafting Diffusion Models for Interactive Matting
SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.
-
DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
DIPO generates articulated 3D objects from a closed and an open image, and the new PM-X dataset improves generalization to complex objects.
Discussion (0). Continue with ORCID to comment.