REVIEW 9 cited by
Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce Metric3D v2, a geometric foundation model for zero-shot metric depth and surface normal estimation from a single image, which is crucial for metric 3D recovery. While depth and normal are geometrically related and highly complimentary, they present distinct challenges. SoTA monocular depth methods achieve zero-shot generalization by learning affine-invariant depths, which cannot recover real-world metrics. Meanwhile, SoTA normal estimation methods have limited zero-shot performance due to the lack of large-scale labeled data. To tackle these issues, we propose solutions for both metric depth estimation and surface normal estimation. For metric depth estimation, we show that the key to a zero-shot single-view model lies in resolving the metric ambiguity from various camera models and large-scale data training. We propose a canonical camera space transformation module, which explicitly addresses the ambiguity problem and can be effortlessly plugged into existing monocular models. For surface normal estimation, we propose a joint depth-normal optimization module to distill diverse data knowledge from metric depth, enabling normal estimators to learn beyond normal labels. Equipped with these modules, our depth-normal models can be stably trained with over 16 million of images from thousands of camera models with different-type annotations, resulting in zero-shot generalization to in-the-wild images with unseen camera settings. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. Our project page is at https://JUGGHM.github.io/Metric3Dv2.
Forward citations
Cited by 9 Pith papers
-
DVPSFormer: Efficient Online Depth-aware Video Panoptic Segmentation for Autonomous Driving
DVPSFormer runs depth-aware video panoptic segmentation online, using segmentation queries as a scene-discretization prior for a single-pass metric depth head, and beats prior DVPQ scores on Cityscapes-DVPS and SemKITTI-DVPS.
-
SeeSE3: Emergence of 3D Space in Vision Features
Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.
-
Towards Consistent Video Geometry Estimation
ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.
-
UAVScenes: A Multi-Modal Dataset for UAVs
UAVScenes adds frame-wise image and LiDAR semantic labels, reconstructed 6-DoF poses, and 3D maps to 120k frames of the MARS-LVIG dataset, with six benchmark tasks.
-
DCHM: Depth-Consistent Human Modeling for Multiview Detection
DCHM uses superpixel-based Gaussian Splatting to make monocular depth estimates multiview-consistent, producing point clouds that yield state-of-the-art label-free pedestrian detection on Wildtrack, Terrace, and MultiviewX.
-
ODG: Occupancy Prediction Using Dual Gaussians
ODG uses separate static and dynamic Gaussian query sets, refined coarse-to-fine, plus rendering supervision, and reports state-of-the-art occupancy prediction on Occ3D-nuScenes and Occ3D-Waymo.
-
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
NDF treats a fixed-image depth estimator as an implicit field and optimizes it on observed depth at test time, improving inpainting accuracy and cross-view consistency.
-
LTM: Large-scale Terrain Model for Wildfire-prone Landscapes
A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.
Discussion (0). Sign in to comment.