Pith. sign in

REVIEW 10 cited by

Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.13928 v2 pith:UFH2BRMX submitted 2025-01-23 cs.CV cs.AIcs.GRcs.RO

classification cs.CVcs.AIcs.GRcs.RO
keywords reconstructionfast3rimagesmulti-viewalignmentapplicationsdust3rforward
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R's Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  2. LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An incremental 3D Gaussian Splatting pipeline that jointly optimizes camera poses and scene geometry using MASt3R priors and density-adaptive octree anchors achieves state-of-the-art novel view synthesis on casual lon...

  3. Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Puzzles synthesizes posed video-depth clips from single images and keyframes, letting 3D reconstruction models match full-data accuracy using only 10% of the data.

  4. Test3R: Learning to Reconstruct 3D at Test Time

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Test3R improves 3D reconstruction by optimizing visual prompts at test time so that pointmaps from different image pairs are geometrically consistent.

  5. PoseIDON: 6DoF Pose Estimation with Foundation Model Features for Marine Sediment Burial Mapping

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PoseIDON estimates 6DoF object pose and seafloor plane from ROV video using DINOv2/FoundPose and photogrammetry, achieving about 10 cm mean burial-depth error on 54 buried objects.

  6. X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography

    eess.IV 2025-05 conditional novelty 6.0 of 10

    A large transformer with fixed-voxel Gaussian splatting reconstructs CT volumes from 6-10 X-ray projections in under a second, substantially beating prior sparse-view methods in simulation.

  7. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  8. DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A pose-free deblurring 3D Gaussian Splatting pipeline using DUSt3R point clouds, confidence-balanced sampling, and event-decoded latent image supervision.

  9. SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SceneCompleter jointly denoises RGB and depth latents, conditioned on projected depth and global scene features, yielding higher quality and more pose-consistent novel views than 2D-inpainting baselines.

  10. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

Pith tools