Pith. sign in

REVIEW 14 cited by

MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.14460 v1 pith:OVJQSIK2 submitted 2023-07-26 cs.CV

classification cs.CV
keywords midasvisiondepthestimationmodelstransformersqualityrelease
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We release MiDaS v3.1 for monocular depth estimation, offering a variety of new models based on different encoder backbones. This release is motivated by the success of transformers in computer vision, with a large variety of pretrained vision transformers now available. We explore how using the most promising vision transformers as image encoders impacts depth estimation quality and runtime of the MiDaS architecture. Our investigation also includes recent convolutional approaches that achieve comparable quality to vision transformers in image classification tasks. While the previous release MiDaS v3.0 solely leverages the vanilla vision transformer ViT, MiDaS v3.1 offers additional models based on BEiT, Swin, SwinV2, Next-ViT and LeViT. These models offer different performance-runtime tradeoffs. The best model improves the depth estimation quality by 28% while efficient models enable downstream tasks requiring high frame rates. We also describe the general process for integrating new backbones. A video summarizing the work can be found at https://youtu.be/UjaeNNFf9sE and the code is available at https://github.com/isl-org/MiDaS.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth

    cs.CV 2026-07 conditional novelty 7.0 of 10

    EpiDistill uses depth-guided epipolar attention and learnable rectified stereo tokens to distill multi-view scale knowledge into single-view monocular depth models.

  2. Evaluating Robustness of Monocular Depth Estimation with Procedural Scene Perturbations

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A procedural-generation benchmark (PDE) shows that depth models are surprisingly vulnerable to camera changes and occlusion, while resisting lighting changes.

  3. Breaking the Horizontal Prior: From Long-Tailed Orientation Bias to Roll-Robust Monocular Depth Estimation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Auxiliary rotation-stable depth cues, added during fine-tuning, reduce monocular depth-error degradation under camera roll across five benchmarks.

  4. DGSfM: Depth-Guided Scale-Aware Global Structure-from-Motion

    cs.CV 2026-07 accept novelty 6.0 of 10

    Monocular depth priors turn scale-ambiguous global SfM into a scale-aware pipeline that measurably improves camera pose accuracy on ETH3D and IMC2021.

  5. AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    AerialMetric is a new benchmark dataset and evaluation suite for adapting monocular metric depth estimation models to real-world UAV aerial views.

  6. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.

  7. Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Across 69 monocular depth estimators, human-likeness of error patterns peaks near human-level accuracy and declines for the most accurate models: accuracy does not guarantee human-like depth perception.

  8. SpatialTrackerV2: 3D Point Tracking Made Easy

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.

  9. RaCalNet: Radar Calibration Network for Sparse-Supervised Metric Depth Estimation

    cs.CV 2025-06 reject novelty 6.0 of 10

    A radar-camera depth estimation framework that recalibrates sparse radar points and aligns a frozen monocular depth model using sparse LiDAR labels, claiming state-of-the-art accuracy with roughly 1% supervision density.

  10. LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    LiteMVS improves efficient multi-view stereo depth estimation by injecting semantic descriptors, MoE cost aggregation, and pseudo-labels from monocular foundation models.

  11. ER-LoRA: Effective-Rank Guided Adaptation for Weather-Generalized Depth Estimation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Tuning only 8.7M parameters of a frozen DINOv2 on daytime data is reported to beat prior PEFT, full fine-tuning, synthetic-data depth methods, and Depth Anything V2 on zero-shot adverse-weather benchmarks.

  12. Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Concatenating monocular depth maps as an extra input channel improves video instance segmentation and reaches 56.2 AP, a new state of the art on OVIS.

  13. Depth Anything at Any Condition

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.

  14. URS-Stereo: Uncertainty-Guided Residual Search for Real-Time Stereo Matching

    cs.CV 2026-07 conditional novelty 4.5 of 10

    Uncertainty-modulated residual offsets relocate local cost-volume centers in coarse-to-fine stereo matching, improving zero-shot disparity accuracy while keeping real-time speed.

Pith tools