Pith. sign in

REVIEW 20 cited by

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11616 v3 pith:EYEH5G3F submitted 2024-05-19 cs.CV

classification cs.CV
keywords multiviewera3dimagesattentioncameradiffusionmethodsefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce Era3D, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resulting in poor-quality multiview images. Specifically, these methods assume that the input images should comply with a predefined camera type, e.g. a perspective camera with a fixed focal length, leading to distorted shapes when the assumption fails. Moreover, the full-image or dense multiview attention they employ leads to an exponential explosion of computational complexity as image resolution increases, resulting in prohibitively expensive training costs. To bridge the gap between assumption and reality, Era3D first proposes a diffusion-based camera prediction module to estimate the focal length and elevation of the input image, which allows our method to generate images without shape distortions. Furthermore, a simple but efficient attention layer, named row-wise attention, is used to enforce epipolar priors in the multiview diffusion, facilitating efficient cross-view information fusion. Consequently, compared with state-of-the-art methods, Era3D generates high-quality multiview images with up to a 512*512 resolution while reducing computation complexity by 12x times. Comprehensive experiments demonstrate that Era3D can reconstruct high-quality and detailed 3D meshes from diverse single-view input images, significantly outperforming baseline multiview diffusion methods. Project page: https://penghtyx.github.io/Era3D/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MV-RAG: Retrieval Augmented Multiview Diffusion

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A retrieval-augmented multiview diffusion model conditions on web images to generate 3D-consistent views of rare concepts, trained with a hybrid 3D/2D objective and evaluated on a new OOD benchmark.

  2. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  3. Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A tuning-free dual-pipeline that injects original normal latents into an edited multi-view diffusion stream, preserving geometry during 2D-to-3D appearance editing.

  4. DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...

  5. 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage cascaded video diffusion model generates 16-view consistent videos from a monocular video, enabling higher-quality 4D content reconstruction.

  6. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...

  7. Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Zero-P-to-3 fuses multi-view diffusion, a restoration prior, and a coarse 3D Gaussian rendering in DDIM sampling, then refines with rotated views, and reports improved invisible-region reconstruction from partial-view...

  8. A 3D Facial Reconstruction Evaluation Methodology: Comparing Smartphone Scans with Deep Learning Based Methods Using Geometry and Morphometry Criteria

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A new benchmarking framework using geometric morphometrics shows smartphone 3D facial scans preserve shape better than deep learning reconstructions from 2D images.

  9. Pippo: High-Resolution Multi-View Humans from a Single Image

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.

  10. Consistent Flow Distillation for Text-to-3D Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.

  11. Grid: Omni Visual Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GRID shows that fine-tuning an image diffusion model on videos arranged as grid images can generate coherent video and multi-view sequences with far less data and compute than specialized video models.

  12. MVReward: Better Aligning and Evaluating Multi-View Diffusion Models with Human Preferences

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MVReward, trained on 16,000 human comparisons, matches human ranking of seven multi-view generation methods perfectly (Spearman 1.00) and MVP fine-tuning improves human preference for Wonder3D and Era3D.

  13. Trajectory Attention for Fine-grained Video Motion Control

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An auxiliary trajectory attention branch, added to temporal attention in video diffusion models, improves camera motion control precision while preserving generation quality.

  14. PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A pipeline that uses multi-view diffusion-generated depth images as shape priors, fused with the partial point cloud via attention and confidence filtering, achieves state-of-the-art completion on custom single-view b...

  15. DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion-based pipeline that rigs 3D Gaussian characters, including hair and clothing, using a newly curated dataset of 9,420 anime meshes.

  16. Direct and Explicit 3D Generation from a Single Image

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A modified Stable Diffusion model generates six views of depth, color, and 3D Gaussian features from one image, then lifts them into a textured mesh or splatted scene in 15 to 25 seconds.

  17. MV-Adapter: Multi-view Consistent Image Generation Made Easy

    cs.CV 2024-12 conditional novelty 5.0 of 10

    An adapter bolts multi-view generation onto frozen text-to-image diffusion models, producing consistent views at up to 768 resolution on SDXL.

  18. MVBoost: Boost 3D Reconstruction with Multi-View Refinement

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MVBoost generates pseudo-ground-truth multi-view images by diffusing renderings of a base 3D model, trains a boosted reconstruction model on them, and reports SOTA on GSO.

  19. Boosting 3D Object Generation through PBR Materials

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A plug-and-play pipeline that upgrades single-image 3D generators with PBR materials and refined normals, tested on CRM, Wonder3D, TripoSR and InstantMesh.

  20. RIGI: Rectifying Image-to-3D Generation Inconsistency via Uncertainty-aware Learning

    cs.CV 2024-11 conditional novelty 4.0 of 10

    RIGI improves image-to-3D generation by estimating pixel-wise uncertainty from the difference between two 3D Gaussian models and using it to reweight the reconstruction loss, reducing artifacts from inconsistent multi...

Pith tools