Pith. sign in

REVIEW 6 cited by

MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11673 v1 pith:KN3Y5YGR submitted 2024-01-22 cs.CV

classification cs.CV
keywords attentionmvsformerdetailsdifferentfeaturemechanismsmethodmodules
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches have not thoroughly investigated the profound influence of transformers on different MVS modules, resulting in limited depth estimation capabilities. In this paper, we introduce MVSFormer++, a method that prudently maximizes the inherent characteristics of attention to enhance various components of the MVS pipeline. Formally, our approach involves infusing cross-view information into the pre-trained DINOv2 model to facilitate MVS learning. Furthermore, we employ different attention mechanisms for the feature encoder and cost volume regularization, focusing on feature and spatial aggregations respectively. Additionally, we uncover that some design details would substantially impact the performance of transformer modules in MVS, including normalized 3D positional encoding, adaptive attention scaling, and the position of layer normalization. Comprehensive experiments on DTU, Tanks-and-Temples, BlendedMVS, and ETH3D validate the effectiveness of the proposed method. Notably, MVSFormer++ achieves state-of-the-art performance on the challenging DTU and Tanks-and-Temples benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  2. MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network

    cs.CV 2025-07 reject novelty 6.0 of 10

    Injecting monocular depth and feature priors from a pretrained single-image model into a multi-view stereo network yields state-of-the-art point cloud accuracy on DTU and Tanks-and-Temples.

  3. High-Fidelity and Generalizable Neural Surface Reconstruction with Sparse Feature Volumes

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SVRecon predicts coarse voxel occupancies and builds sparse high-resolution feature volumes only near surfaces, enabling 512^3 generalizable reconstruction with over 50x less storage than dense-volume methods.

  4. A View-consistent Sampling Method for Regularized Training of Neural Radiance Fields

    cs.CV 2025-07 conditional novelty 6.0 of 10

    View-consistent sampling driven by color and distilled DINOv2 features, plus a depth-pushing loss, improves NeRF novel-view synthesis over depth-based regularizers.

  5. 2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

    cs.CV 2026-03 conditional novelty 5.5 of 10

    Entropy-guided sparse refinement upgrades frozen low-resolution geometric foundation models to accurate 2K depth and pointmap outputs at a fraction of full-resolution cost.

  6. LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    LiteMVS improves efficient multi-view stereo depth estimation by injecting semantic descriptors, MoE cost aggregation, and pseudo-labels from monocular foundation models.

Pith tools