REVIEW 20 cited by
Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we introduce Era3D, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resulting in poor-quality multiview images. Specifically, these methods assume that the input images should comply with a predefined camera type, e.g. a perspective camera with a fixed focal length, leading to distorted shapes when the assumption fails. Moreover, the full-image or dense multiview attention they employ leads to an exponential explosion of computational complexity as image resolution increases, resulting in prohibitively expensive training costs. To bridge the gap between assumption and reality, Era3D first proposes a diffusion-based camera prediction module to estimate the focal length and elevation of the input image, which allows our method to generate images without shape distortions. Furthermore, a simple but efficient attention layer, named row-wise attention, is used to enforce epipolar priors in the multiview diffusion, facilitating efficient cross-view information fusion. Consequently, compared with state-of-the-art methods, Era3D generates high-quality multiview images with up to a 512*512 resolution while reducing computation complexity by 12x times. Comprehensive experiments demonstrate that Era3D can reconstruct high-quality and detailed 3D meshes from diverse single-view input images, significantly outperforming baseline multiview diffusion methods. Project page: https://penghtyx.github.io/Era3D/.
Forward citations
Cited by 20 Pith papers
-
MV-RAG: Retrieval Augmented Multiview Diffusion
A retrieval-augmented multiview diffusion model conditions on web images to generate 3D-consistent views of rare concepts, trained with a hybrid 3D/2D objective and evaluated on a new OOD benchmark.
-
MVGBench: Comprehensive Benchmark for Multi-view Generation Models
MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.
-
Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
A tuning-free dual-pipeline that injects original normal latents into an edited multi-view diffusion stream, preserving geometry during 2D-to-3D appearance editing.
-
DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...
-
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.
-
PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...
-
Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...
-
A 3D Facial Reconstruction Evaluation Methodology: Comparing Smartphone Scans with Deep Learning Based Methods Using Geometry and Morphometry Criteria
A new benchmarking framework using geometric morphometrics shows smartphone 3D facial scans preserve shape better than deep learning reconstructions from 2D images.
-
Pippo: High-Resolution Multi-View Humans from a Single Image
A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.
-
Consistent Flow Distillation for Text-to-3D Generation
Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.
-
Grid: Omni Visual Generation
GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.
-
MVReward: Better Aligning and Evaluating Multi-View Diffusion Models with Human Preferences
MVReward, trained on 16,000 human comparisons, matches human ranking of seven multi-view generation methods perfectly (Spearman 1.00) and MVP fine-tuning improves human preference for Wonder3D and Era3D.
-
Trajectory Attention for Fine-grained Video Motion Control
An auxiliary trajectory attention branch, added to temporal attention in video diffusion models, improves camera motion control precision while preserving generation quality.
-
PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
A pipeline that uses multi-view diffusion-generated depth images as shape priors, fused with the partial point cloud via attention and confidence filtering, achieves state-of-the-art completion on custom single-view b...
-
DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters
A diffusion-based pipeline that rigs 3D Gaussian characters, including hair and clothing, using a newly curated dataset of 9,420 anime meshes.
-
Direct and Explicit 3D Generation from a Single Image
A modified Stable Diffusion model generates six views of depth, color, and 3D Gaussian features from one image, then lifts them into a textured mesh or splatted scene in 15 to 25 seconds.
-
MV-Adapter: Multi-view Consistent Image Generation Made Easy
An adapter bolts multi-view generation onto frozen text-to-image diffusion models, producing consistent views at up to 768 resolution on SDXL.
-
MVBoost: Boost 3D Reconstruction with Multi-View Refinement
MVBoost generates pseudo-ground-truth multi-view images by diffusing renderings of a base 3D model, trains a boosted reconstruction model on them, and reports SOTA on GSO.
-
Boosting 3D Object Generation through PBR Materials
A plug-and-play pipeline that upgrades single-image 3D generators with PBR materials and refined normals, tested on CRM, Wonder3D, TripoSR and InstantMesh.
-
RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-aware Learning
RIGI improves image-to-3D generation by estimating pixel-wise uncertainty from the difference between two 3D Gaussian models and using it to reweight the reconstruction loss, reducing artifacts from inconsistent multi...
Discussion (0). Continue with ORCID to comment.