REVIEW 6 cited by
MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches have not thoroughly investigated the profound influence of transformers on different MVS modules, resulting in limited depth estimation capabilities. In this paper, we introduce MVSFormer++, a method that prudently maximizes the inherent characteristics of attention to enhance various components of the MVS pipeline. Formally, our approach involves infusing cross-view information into the pre-trained DINOv2 model to facilitate MVS learning. Furthermore, we employ different attention mechanisms for the feature encoder and cost volume regularization, focusing on feature and spatial aggregations respectively. Additionally, we uncover that some design details would substantially impact the performance of transformer modules in MVS, including normalized 3D positional encoding, adaptive attention scaling, and the position of layer normalization. Comprehensive experiments on DTU, Tanks-and-Temples, BlendedMVS, and ETH3D validate the effectiveness of the proposed method. Notably, MVSFormer++ achieves state-of-the-art performance on the challenging DTU and Tanks-and-Temples benchmarks.
Forward citations
Cited by 6 Pith papers
-
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.
-
MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network
Injecting monocular depth and feature priors from a pretrained single-image model into a multi-view stereo network yields state-of-the-art point cloud accuracy on DTU and Tanks-and-Temples.
-
High-Fidelity and Generalizable Neural Surface Reconstruction with Sparse Feature Volumes
SVRecon predicts coarse voxel occupancies and builds sparse high-resolution feature volumes only near surfaces, enabling 512^3 generalizable reconstruction with over 50x less storage than dense-volume methods.
-
A View-consistent Sampling Method for Regularized Training of Neural Radiance Fields
View-consistent sampling driven by color and distilled DINOv2 features, plus a depth-pushing loss, improves NeRF novel-view synthesis over depth-based regularizers.
-
2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction
Entropy-guided sparse refinement upgrades frozen low-resolution geometric foundation models to accurate 2K depth and pointmap outputs at a fraction of full-resolution cost.
-
LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation
LiteMVS improves efficient multi-view stereo depth estimation by injecting semantic descriptors, MoE cost aggregation, and pseudo-labels from monocular foundation models.
Discussion (0). Continue with ORCID to comment.