Pith. sign in

REVIEW 4 major objections 4 minor 140 references

The Fourth Monocular Depth Estimation Challenge

T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The fourth Monocular Depth Estimation Challenge reports a new best zero-shot result on SYNS-Patches: a 23.05% reconstruction F-Score, up from 22.58% in the previous edition, under a revised least-squares alignment protocol that puts…

desk verdict A competent challenge report that is honest about marginal gains; the headline 0.47-point F-score improvement is real under the stated protocol but fragile because the protocol changed and no uncertainty is reported. read the letter →

arxiv 2504.17787 v1 pith:NNACREFK submitted 2025-04-24 cs.CV

classification cs.CV
keywords monoculardepthestimationzero-shotgeneralizationSYNS-Patchesbenchmarkleast-squaresalignmentaffine-invariantreconstructionF-Scorefoundationmodelssaturation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is the report of the fourth Monocular Depth Estimation Challenge, a zero-shot generalization benchmark on the SYNS-Patches dataset. It claims that after switching the evaluation to least-squares scale-and-shift alignment, the winning method improved the reconstruction F-Score from 22.58% to 23.05%, the current best zero-shot result on this benchmark. It also documents that most competitive submissions rely on affine-invariant predictions built on top of public depth foundation models, and that the benchmark is close to saturation.

What carries the argument

The central object is the evaluation protocol: each prediction is bilinearly upsampled, disparity maps are inverted, and predictions are aligned to ground truth by a least-squares fit of an affine scale and shift before computing the pointcloud-based reconstruction F-Score used for ranking. This two-degree-of-freedom alignment is what makes it possible to compare metric, disparity, and affine-invariant predictions on equal footing on the SYNS-Patches benchmark.

What would settle it

Recompute the Table 1 ranking after fixing each prediction to a single global scale learned from held-out images rather than per-image least-squares scale and shift; if the winning method changes or its F-Score drops below the previous edition's result, the reported improvement is an artifact of the alignment protocol.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes a new best zero-shot monocular depth result on SYNS-Patches: the top submission reaches a reconstruction F-Score of 23.05%, ahead of the previous edition's winner re-evaluated at 22.58% under the same least-squares alignment. The paper also shows that the leading approaches almost all predict affine-invariant depth and are fine-tuned from public foundation models, and that only one reporting team outperformed the previous champion. It concludes that despite sharper qualitative predictions, the benchmark is nearing saturation, with web-scale foundation-model priors now dominating.

Load-bearing premise

The entire ranking assumes that least-squares scale-and-shift alignment is a fair common evaluation for metric, disparity, and affine-invariant depth predictions; if that alignment hides true metric errors, the F-Score comparison is an artifact of the protocol.

Editorial extensions

If this is right

  • Under this protocol, affine-invariant predictions are not a handicap: most top-ranked methods chose them, and the winner is one of them.
  • Zero-shot performance on SYNS-Patches appears close to saturation: the F-Score gain over the previous edition is 0.47 points, while qualitative sharpness improved more.
  • Web-scale foundation models such as Depth Anything v2 and Marigold are now the effective starting point for competitive depth estimation; custom architectures that avoid them only beat baselines on a single metric.
  • Further progress will likely require harder settings such as non-Lambertian surfaces, metric depth recovery, and alternative 3D representations rather than more data or larger foundation models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the organizers leave implicit: because the two-degree-of-freedom alignment absorbs a global scale and shift per image, a model that predicts only relative order can score as well as a metric model on the F-Score, so a metric-only track would likely change which method wins.
  • The gap between first and second place is 0.10 F-Score points, smaller than the cross-edition gain of 0.47, so a testable extension would be to report confidence intervals or per-scene variance before declaring a new state of the art.
  • The near-duplicate rows among anonymous entries suggest participants can re-submit minor variants; a future edition could detect and merge such duplicates before ranking.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper reports the results of the fourth Monocular Depth Estimation Challenge (MDEC), held with CVPR 2025, on the SYNS-Patches benchmark. Relative to previous editions, the evaluation protocol was changed from median scaling to a two-degree-of-freedom least-squares (LSE) alignment in order to admit metric, disparity, scale-invariant, and affine-invariant predictions, and the baselines were updated to Depth Anything v2 and Marigold. Twenty-four submissions are reported as outperforming the baselines on at least one metric, and ten teams additionally provided written method reports. The headline result is that team HRI achieves a reconstruction F-score of 23.05%, which the abstract presents as an improvement over the previous edition's best result of 22.58% (PICO-MR, re-evaluated under the new protocol). The paper contains a full results table, qualitative comparisons, and short technical descriptions of each reported method.

Significance. If the results hold, the paper provides the community with (i) a current snapshot of zero-shot monocular depth estimation on SYNS-Patches, (ii) an evaluation protocol aligned with current affine-invariant evaluation practice, and (iii) evidence that fine-tuned foundation models with affine-invariant output dominate the leaderboard. The organizational effort is substantial and worth acknowledging: 24 submissions, 10 method reports, updated and publicly described baselines, a public starter-pack repository, and a re-evaluation of the previous winner under the new protocol. The leaderboard is an external empirical measurement rather than a derived claim, and the flagged duplicate rows make the table more honest than many challenge reports. The main weakness is that the headline improvement is small (0.47 F-score points) and is demonstrated only under a single alignment protocol with no uncertainty estimates, so the paper's central numerical claim is less robust than its framing suggests.

major comments (4)
  1. [§5 vs. §3 (Evaluation)] There is a direct inconsistency between the two descriptions of the evaluation protocol. Section 3 states that all submissions were aligned using least squares with two degrees of freedom (scale and shift), whereas Section 5 states that predictions were aligned 'according to median depth scaling or least-squares alignment, as requested by the participants.' If participants could select their alignment, Table 1 is not computed under a single protocol and the headline improvement from 22.58% to 23.05% is not well-defined; if the Section 5 sentence is carried over from the previous editions' protocol, it must be removed or corrected.
  2. [Abstract and §5.1, Table 1] The claimed improvement from 22.58% to 23.05% requires the two numbers to be evaluated under the same protocol, but the manuscript is ambiguous about which protocol produced the 22.58% figure: the abstract calls it 'the previous edition's best result' (third edition, median scaling), while Table 1 reports PICO-MR at 22.58 under the new LSE protocol. The authors should state explicitly whether these two 22.58% values coincide exactly, by rounding, or only coincidentally, and should report PICO-MR's score under both the old and the new protocols so that the reader can see how much the re-evaluation changed its result.
  3. [§3 Evaluation; Table 1] The rank ordering between HRI (affine-invariant, marked 'A' in Table 1) and PICO-MR (not marked affine-invariant) depends on the choice of a two-degree-of-freedom LSE alignment, because affine-invariant methods receive an oracle per-image scale-and-shift correction while metric or disparity predictions are thereby also allowed a free shift that they should not need. Since the margin is only 0.47 F-score points, the authors should either provide a sensitivity analysis of the top entries under scale-only alignment and under the previous median-scaling protocol, or argue explicitly why the 2-DOF protocol is the only principled way to compare all three prediction types; as it stands, the paper does not rule out a protocol artifact.
  4. [§5.1, Table 1] The headline gap of 0.47 F-score points is presented without any uncertainty estimate, and Table 1 itself contains tied or near-identical rows (ranks 11–12, 13, and 15–16), indicating that per-image variability is not negligible. The authors should add bootstrapped confidence intervals or per-image standard deviations for at least the top three entries, and should temper the precision of the 'raising it from 22.58% to 23.05%' claim accordingly.
minor comments (4)
  1. [§3, Evaluation] The description of the new alignment procedure is too terse to reproduce; please specify whether the LSE fit is computed per image on depth or inverse-depth values, whether it is weighted, and whether upsampling precedes or follows the fit, and point to the exact evaluation script in the starter pack.
  2. [§3 and §5] The paper reports only submissions that beat the baselines on at least one metric; for transparency it should state the total number of test-phase submissions received and how many were excluded, so that the reported '24 submissions' is interpretable.
  3. [Table 1] The flagged near-duplicate rows (ranks 11–13 and 15–16, the latter identical to Marigold) are retained in the count; the text says '24 teams,' but the number of distinct methods is smaller and should be stated explicitly.
  4. [Throughout] Minor editorial issues include 'alongside to' (Section 1), inconsistent spelling of 'DepthAnything' versus 'Depth Anything', broken superscripts in the delta-metric column headers of Table 1, and the space in the URL 'toshas/mdec benchmark' (footnote 2).

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical challenge-leaderboard report, and the headline F-Score comparison is an external measurement rather than a derived quantity.

full rationale

The paper does not contain a derivation chain in the sense that would admit circular reduction. Its central claim is that the HRI submission achieved a 3D F-Score of 23.05% and thereby improved on the previous edition's best result of 22.58% (abstract, Table 1). Both numbers are empirical measurements of submitted models on the SYNS-Patches test split. The 22.58% figure for PICO-MR is a re-evaluation of a previous submission under the same least-squares alignment protocol, as stated in Section 5.1: 'we highlighted ... the results achieved by the winning team of the third edition according to the new aligning protocol.' No fitted parameter is relabeled as a prediction; the per-image least-squares alignment used in Section 3 is an evaluation protocol common for comparing disparity, affine-invariant, and metric outputs, not a parameter fitted by the paper and then reported as a derived result. The self-citations to the authors' previous challenge report [99] are load-bearing only as a source of an empirical baseline score, and that score is re-computed under the current evaluation protocol rather than imported as an unverified theorem. Concerns about whether two-degree-of-freedom alignment is the fairest protocol, or whether the 0.47-point margin is within sampling noise, are validity and robustness questions, not circularity: they do not show that the reported F-Score was constructed from its own inputs. The leaderboard is therefore self-contained as an external empirical benchmark, and no circular step can be exhibited from the paper's text.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. The load-bearing premises are benchmark quality, alignment validity, and submission integrity.

assumptions (3)
  • domain assumption SYNS-Patches ground-truth LiDAR depths are accurate and dense enough to support the reported metrics.
    The evaluation and all conclusions rely on this ground truth being a reliable reference; the paper states average coverage of 78.20% and manual artifact removal but does not independently verify the test set.
  • domain assumption Least-squares alignment with two degrees of freedom is a valid way to compare metric, disparity, and affine-invariant predictions.
    The ranking is computed after this alignment; if it unfairly benefits affine-invariant predictions or suppresses metric errors, the headline F-score comparison is not meaningful. Introduced in Section 3 Evaluation.
  • domain assumption Submitted predictions were produced honestly from single images without test-set leakage or geometric priors.
    The challenge rule states only one input image may be used; the organizers cannot verify every submission's inference process.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Fourth Monocular Depth Estimation Challenge." pith.science (2026). https://pith.science/paper/NNACREFK

@misc{pith2026250417787,
  author       = {Pith},
  title        = {Pith review of: The Fourth Monocular Depth Estimation Challenge},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NNACREFK}},
  note         = {Machine review of arXiv:2504.17787}
}
read the original abstract

This paper presents the results of the fourth edition of the Monocular Depth Estimation Challenge (MDEC), which focuses on zero-shot generalization to the SYNS-Patches benchmark, a dataset featuring challenging environments in both natural and indoor settings. In this edition, we revised the evaluation protocol to use least-squares alignment with two degrees of freedom to support disparity and affine-invariant predictions. We also revised the baselines and included popular off-the-shelf methods: Depth Anything v2 and Marigold. The challenge received a total of 24 submissions that outperformed the baselines on the test set; 10 of these included a report describing their approach, with most leading methods relying on affine-invariant predictions. The challenge winners improved the 3D F-Score over the previous edition's best result, raising it from 22.58% to 23.05%.

Figures

Figures reproduced from arXiv: 2504.17787 by the authors.

Figure 1
Figure 1. SYNS-Patches Dataset. The dataset contains samples from diverse scenes, including complex urban, natural, and indoor spaces. High-quality ground-truth depth measurements cover ∼78.20% of pixels; edges and object boundary annotations are also available. urban driving scenarios [32], resulting in highly-specialized solutions that failed to generalize beyond their specific do￾main of expertise. To address this limitati… view at source ↗
Figure 2
Figure 2. Qualitatives on SYNS-Patches. Best viewed in color and zoomed in. Methods are ranked based on their F-Score in [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

140 extracted references · 59 canonical work pages

  1. [1]

    The Southampton-York Natural Scenes (SYNS) dataset: Statis- tics of surface attitude

    Wendy J Adams, James H Elder, Erich W Graf, Julian Leyland, Arthur J Lugtigheid, and Alexander Muryy. The Southampton-York Natural Scenes (SYNS) dataset: Statis- tics of surface attitude. Scientific Reports, 6(1):35805, 2016

  2. [2]

    Attention attention everywhere: Monocular depth prediction with skip atten- tion

    Ashutosh Agarwal and Chetan Arora. Attention attention everywhere: Monocular depth prediction with skip atten- tion. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5861–5870, 2023

  3. [3]

    Generative adversarial networks for unsupervised monocular depth prediction

    Filippo Aleotti, Fabio Tosi, Matteo Poggi, and Stefano Mat- toccia. Generative adversarial networks for unsupervised monocular depth prediction. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018

  4. [4]

    Real-time single image depth perception in the wild with handheld devices

    Filippo Aleotti, Giulio Zaccaroni, Luca Bartolomei, Matteo Poggi, Fabio Tosi, and Stefano Mattoccia. Real-time single image depth perception in the wild with handheld devices. Sensors, 21(1):15, 2020

  5. [5]

    Enhancing self-supervised monocular depth estimation with traditional visual odometry

    Lorenzo Andraghetti, Panteleimon Myriokefalitakis, Pier Luigi Dovesi, Belen Luque, Matteo Poggi, Alessandro Pieropan, and Stefano Mattoccia. Enhancing self-supervised monocular depth estimation with traditional visual odometry. In 2019 International Conference on 3D Vision (3DV) , pages 424–433. IEEE, 2019

  6. [6]

    Adabins: Depth estimation using adaptive bins

    Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Adabins: Depth estimation using adaptive bins. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4009–4018, 2021

  7. [7]

    Localbins: Improving depth estimation by learning local distributions

    Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Localbins: Improving depth estimation by learning local distributions. In Computer Vision–ECCV 2022: 17th Eu- ropean Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part I, pages 480–496. Springer, 2022

  8. [8]

    Zoedepth: Zero-shot trans- fer by combining relative and metric depth

    Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias M ¨uller. Zoedepth: Zero-shot trans- fer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023

Show all 140 references
  1. [9]

    Unsu- pervised Scale-consistent Depth and Ego-motion Learning from Monocular Video

    Jiawang Bian, Zhichao Li, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, and Ian Reid. Unsu- pervised Scale-consistent Depth and Ego-motion Learning from Monocular Video. In Advances in Neural Information Processing Systems, volume 32, 2019

  2. [10]

    Depth pro: Sharp monocular metric depth in less than a second

    Aleksei Bochkovskii, Ama ˜AG ¸l Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073 , 2024

  3. [11]

    Virtual kitti 2

    Yohann Cabon, Naila Murray, and Martin Humenberger. Virtual kitti 2. arXiv preprint arXiv:2001.10773, 2020

  4. [12]

    Depth prediction without the sensors: Lever- aging structure for unsupervised learning from monocular videos

    Vincent Casser, Soeren Pirk, Reza Mahjourian, and Anelia Angelova. Depth prediction without the sensors: Lever- aging structure for unsupervised learning from monocular videos. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8001–8008, 2019

  5. [13]

    Video depth anything: Consistent depth estimation for super-long videos

    Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zi- long Huang, Jiashi Feng, and Bingyi Kang. Video depth anything: Consistent depth estimation for super-long videos. arXiv preprint arXiv:2501.12375, 2025

  6. [14]

    Single- image depth perception in the wild

    Weifeng Chen, Zhao Fu, Dawei Yang, and Jia Deng. Single- image depth perception in the wild. Advances in neural information processing systems, 29, 2016

  7. [15]

    OASIS: A large-scale dataset for single image 3d in the wild

    Weifeng Chen, Shengyi Qian, David Fan, Noriyuki Kojima, Max Hamilton, and Jia Deng. OASIS: A large-scale dataset for single image 3d in the wild. In CVPR, 2020

  8. [16]

    Swin-depth: Us- ing transformers and multi-scale fusion for monocular-based depth estimation

    Zeyu Cheng, Yi Zhang, and Chengkai Tang. Swin-depth: Us- ing transformers and multi-scale fusion for monocular-based depth estimation. IEEE Sensors Journal , 21(23):26912– 26920, 2021

  9. [17]

    Diml/cvl rgb-d dataset: 2m rgb-d im- ages of natural indoor and outdoor scenes

    Jaehoon Cho, Dongbo Min, Youngjung Kim, and Kwanghoon Sohn. Diml/cvl rgb-d dataset: 2m rgb-d im- ages of natural indoor and outdoor scenes. arXiv preprint arXiv:2110.11590, 2021

  10. [18]

    Adaptive confidence thresholding for monocular depth es- timation

    Hyesong Choi, Hunsang Lee, Sunkyung Kim, Sunok Kim, Seungryong Kim, Kwanghoon Sohn, and Dongbo Min. Adaptive confidence thresholding for monocular depth es- timation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12808–12818, 2021

  11. [19]

    Energy-quality scalable monocular depth estimation on low- power cpus

    Antonio Cipolletta, Valentino Peluso, Andrea Calimera, Mat- teo Poggi, Fabio Tosi, Filippo Aleotti, and Stefano Mattoccia. Energy-quality scalable monocular depth estimation on low- power cpus. IEEE Internet of Things Journal, 9(1):25–36, 2021

  12. [20]

    The cityscapes dataset for semantic urban scene understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognitio...

  13. [21]

    Learn- ing depth estimation for transparent and mirror surfaces

    Alex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, and Luigi Di Stefano. Learn- ing depth estimation for transparent and mirror surfaces. In The IEEE International Conference on Computer Vision ,

  14. [22]

    Self-supervised object motion and depth estimation from video

    Qi Dai, Vaishakh Patil, Simon Hecker, Dengxin Dai, Luc Van Gool, and Konrad Schindler. Self-supervised object motion and depth estimation from video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2020

  15. [23]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009

  16. [24]

    DiffusionDepth: Diffusion denoising approach for monocular depth estima- tion

    Yiqun Duan, Xianda Guo, and Zheng Zhu. DiffusionDepth: Diffusion denoising approach for monocular depth estima- tion. arXiv preprint arXiv:2303.05021, 2023

  17. [25]

    Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans

    Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In ICCV, pages 10786–10796, 2021

  18. [26]

    Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-scale Convolutional Architecture

    David Eigen and Rob Fergus. Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-scale Convolutional Architecture. In International Conference on Computer Vision, pages 2650–2658, 2015

  19. [27]

    Deep ordinal regression network for monocular depth estimation

    Huan Fu, Mingming Gong, Chaohui Wang, Kayhan Bat- manghelich, and Dacheng Tao. Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE conference on computer vision and pattern recog- 9 nition, pages 2002–2011, 2018

  20. [28]

    Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image

    Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image. In European Conference on Computer Vision, pages 241–258. Springer, 2024

  21. [29]

    Unsupervised CNN for Single View Depth Estimation: Ge- ometry to the Rescue

    Ravi Garg, Vijay Kumar, Gustavo Carneiro, and Ian Reid. Unsupervised CNN for Single View Depth Estimation: Ge- ometry to the Rescue. In European Conference on Computer Vision, pages 740–756, 2016

  22. [30]

    R4dyn: Exploring radar for self-supervised monocular depth estima- tion of dynamic scenes

    Stefano Gasperini, Patrick Koch, Vinzenz Dallabetta, Nassir Navab, Benjamin Busam, and Federico Tombari. R4dyn: Exploring radar for self-supervised monocular depth estima- tion of dynamic scenes. In 2021 International Conference on 3D Vision (3DV), pages 751–760. IEEE, 2021

  23. [31]

    Robust monocular depth estimation under challenging conditions

    Stefano Gasperini, Nils Morbitzer, HyunJun Jung, Nassir Navab, and Federico Tombari. Robust monocular depth estimation under challenging conditions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023

  24. [32]

    Vision meets robotics: The KITTI dataset

    A Geiger, P Lenz, C Stiller, and R Urtasun. Vision meets robotics: The KITTI dataset. International Journal of Robotics Research, 32(11):1231–1237, 2013

  25. [33]

    Clement Godard, Oisin Mac Aodha, and Gabriel J. Brostow. Unsupervised Monocular Depth Estimation with Left-Right Consistency. Conference on Computer Vision and Pattern Recognition, pages 6602–6611, 2017

  26. [34]

    Digging Into Self-Supervised Monocular Depth Estimation

    Clement Godard, Oisin Mac Aodha, Michael Firman, and Gabriel Brostow. Digging Into Self-Supervised Monocular Depth Estimation. International Conference on Computer Vision, 2019-Octob:3827–3837, 2019

  27. [35]

    PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single- View Depth Estimation with Neural Positional Encoding and Distilled Matting Loss

    Juan Luis Gonzalez Bello and Munchurl Kim. PLADE-Net: Towards Pixel-Level Accuracy for Self-Supervised Single- View Depth Estimation with Neural Positional Encoding and Distilled Matting Loss. In Conference on Computer Vision and Pattern Recognition, pages 6847–6856, 2021

  28. [36]

    Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras

    Ariel Gordon, Hanhan Li, Rico Jonschkowski, and Anelia Angelova. Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras. In Pro- ceedings of the IEEE/CVF International Conference on Com- puter Vision, pages 8977–8986, 2019

  29. [37]

    3D packing for self- supervised monocular depth estimation

    Vitor Guizilini, Ambrus Ambrus, Sudeep Pillai, Allan Raventos, and Adrien Gaidon. 3D packing for self- supervised monocular depth estimation. Conference on Computer Vision and Pattern Recognition, pages 2482–2491, 2020

  30. [38]

    Towards zero-shot scale-aware monoc- ular depth estimation

    Vitor Guizilini, Igor Vasiljevic, Dian Chen, Rares, Ambrus,, and Adrien Gaidon. Towards zero-shot scale-aware monoc- ular depth estimation. In ICCV, 2023

  31. [39]

    Lotus: Diffusion-based visual foundation model for high-quality dense prediction

    Jing He, Haodong Li, Wei Yin, Yixun Liang, Leheng Li, Kaiqiang Zhou, Hongbo Liu, Bingbing Liu, and Ying- Cong Chen. Lotus: Diffusion-based visual foundation model for high-quality dense prediction. arXiv preprint arXiv:2409.18124, 2024

  32. [40]

    Denoising diffu- sion probabilistic models, 2020

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffu- sion probabilistic models, 2020

  33. [41]

    Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation

    Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE Transactions on Pattern Analysis...

  34. [42]

    Depthcrafter: Generating consistent long depth sequences for open-world videos

    Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xiaodong Cun, Yong Zhang, Long Quan, and Ying Shan. Depthcrafter: Generating consistent long depth sequences for open-world videos. arXiv preprint arXiv:2409.02095, 2024

  35. [43]

    Densely connected convolutional net- works

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. Densely connected convolutional net- works. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017

  36. [44]

    The apolloscape open dataset for autonomous driving and its application

    Xinyu Huang, Peng Wang, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, and Ruigang Yang. The apolloscape open dataset for autonomous driving and its application. IEEE transactions on pattern analysis and machine intelligence, 42(10):2702–2719, 2019

  37. [45]

    Lightweight monocular depth with a novel neural architecture search method

    Lam Huynh, Phong Nguyen, Ji ˇr´ı Matas, Esa Rahtu, and Janne Heikkil¨a. Lightweight monocular depth with a novel neural architecture search method. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 3643–3653, 2022

  38. [46]

    DDP: Diffusion model for dense visual prediction

    Yuanfeng Ji, Zhe Chen, Enze Xie, Lanqing Hong, Xihui Liu, Zhaoqiang Liu, Tong Lu, Zhenguo Li, and Ping Luo. DDP: Diffusion model for dense visual prediction. In ICCV, 2023

  39. [47]

    Self-Supervised Monocular Trained Depth Estimation Using Self-Attention and Discrete Disparity V olume

    Adrian Johnston and Gustavo Carneiro. Self-Supervised Monocular Trained Depth Estimation Using Self-Attention and Discrete Disparity V olume. InConference on Computer Vision and Pattern Recognition, pages 4755–4764, 2020

  40. [48]

    Video depth without video models

    Bingxin Ke, Dominik Narnhofer, Shengyu Huang, Lei Ke, Torben Peters, Katerina Fragkiadaki, Anton Obukhov, and Konrad Schindler. Video depth without video models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025

  41. [49]

    Repurpos- ing diffusion-based image generators for monocular depth estimation

    Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Met- zger, Rodrigo Caye Daudt, and Konrad Schindler. Repurpos- ing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9492–...

  42. [50]

    Supervising the New with the Old: Learning SFM from SFM

    Maria Klodt and Andrea Vedaldi. Supervising the New with the Old: Learning SFM from SFM. In European Conference on Computer Vision, pages 713–728, 2018

  43. [51]

    Evaluation of CNN-Based Single-Image Depth Estimation Methods

    Tobias Koch, Lukas Liebel, Friedrich Fraundorfer, and Marco K ¨orner. Evaluation of CNN-Based Single-Image Depth Estimation Methods. In European Conference on Computer Vision Workshops, pages 331–348, 2018

  44. [52]

    Deeper depth prediction with fully convolutional residual networks

    Iro Laina, Christian Rupprecht, Vasileios Belagiannis, Fed- erico Tombari, and Nassir Navab. Deeper depth prediction with fully convolutional residual networks. International Conference on 3D Vision, pages 239–248, 2016

  45. [53]

    Deep attention-based classification network for robust depth prediction

    Ruibo Li, Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, and Lingxiao Hang. Deep attention-based classification network for robust depth prediction. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, Revised Selected Paper...

  46. [54]

    Mannequin- challenge: Learning the depths of moving people by watch- 10 ing frozen people

    Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker, Noah Snavely, Ce Liu, and William T Freeman. Mannequin- challenge: Learning the depths of moving people by watch- 10 ing frozen people. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(12):4229–4241, 2020

  47. [55]

    Megadepth: Learning single- view depth prediction from internet photos

    Zhengqi Li and Noah Snavely. Megadepth: Learning single- view depth prediction from internet photos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2041–2050, 2018

  48. [56]

    Deep con- volutional neural fields for depth estimation from a single image

    Fayao Liu, Chunhua Shen, and Guosheng Lin. Deep con- volutional neural fields for depth estimation from a single image. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5162–5170, 2015

  49. [57]

    Sm4depth: Seamless monoc- ular metric depth estimation across multiple cameras and scenes by one model

    Yihao Liu, Feng Xue, Anlong Ming, Mingshuai Zhao, Huadong Ma, and Nicu Sebe. Sm4depth: Seamless monoc- ular metric depth estimation across multiple cameras and scenes by one model. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 3469–3478, 2024

  50. [58]

    Swiftdepth: An efficient hybrid cnn-transformer model for self-supervised monocular depth estimation on mobile devices

    Albert Luginov and Ilya Makarov. Swiftdepth: An efficient hybrid cnn-transformer model for self-supervised monocular depth estimation on mobile devices. In 2023 IEEE Interna- tional Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct), pages 642–647. IEEE, 2023

  51. [59]

    Nimbled: Enhanc- ing self-supervised monocular depth estimation with pseudo- labels and large-scale video pre-training

    Albert Luginov and Muhammad Shahzad. Nimbled: Enhanc- ing self-supervised monocular depth estimation with pseudo- labels and large-scale video pre-training. arXiv preprint arXiv:2408.14177, 2024

  52. [60]

    Every pixel counts ++: Joint learning of geometry and motion with 3d holistic understanding

    Chenxu Luo, Zhenheng Yang, Peng Wang, Yang Wang, Wei Xu, Ram Nevatia, and Alan Yuille. Every pixel counts ++: Joint learning of geometry and motion with 3d holistic understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10):2624–2641, 2020

  53. [61]

    HR-Depth: High Resolution Self-Supervised Monocular Depth Estima- tion

    Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen, and Yi Yuan. HR-Depth: High Resolution Self-Supervised Monocular Depth Estima- tion. AAAI Conference on Artificial Intelligence, 35(3):2294– 2301, 2021

  54. [62]

    Un- supervised Learning of Depth and Ego-Motion from Monoc- ular Video Using 3D Geometric Constraints

    Reza Mahjourian, Martin Wicke, and Anelia Angelova. Un- supervised Learning of Depth and Ego-Motion from Monoc- ular Video Using 3D Geometric Constraints. Conference on Computer Vision and Pattern Recognition, pages 5667–5675, 2018

  55. [63]

    Fine-tuning image-conditional diffusion models is easier than you think

    Gonzalo Martin Garcia, Karim Abou Zeid, Christian Schmidt, Daan de Geus, Alexander Hermans, and Bastian Leibe. Fine-tuning image-conditional diffusion models is easier than you think. In Proceedings of the IEEE/CVF Win- ter Conference on Applications of Computer Vision (WACV), 2025

  56. [64]

    Boosting monocular depth estimation models to high-resolution via content-adaptive multi-resolution merging

    S Mahdi H Miangoleh, Sebastian Dille, Long Mai, Syl- vain Paris, and Yagiz Aksoy. Boosting monocular depth estimation models to high-resolution via content-adaptive multi-resolution merging. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...

  57. [65]

    Indoor segmentation and support inference from rgbd images

    Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012

  58. [66]

    Dinov2: Learning robust visual features without supervision

    Maxime Oquab, Timoth´ee Darcet, Th´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023

  59. [67]

    From 2D to 3D: Re-thinking Benchmarking of Monocular Depth Prediction

    Evin Pinar ¨Ornek, Shristi Mudgal, Johanna Wald, Yida Wang, Nassir Navab, and Federico Tombari. From 2D to 3D: Re-thinking Benchmarking of Monocular Depth Prediction. arXiv preprint, 2022

  60. [68]

    Codalab competitions: An open source platform to organize scientific challenges

    Adrien Pavao, Isabelle Guyon, Anne-Catherine Letour- nel, Xavier Bar ´o, Hugo Escalante, Sergio Escalera, Tyler Thomas, and Zhen Xu. Codalab competitions: An open source platform to organize scientific challenges. Technical report, 2022

  61. [69]

    Monocular depth perception on microcontrollers for edge applications

    Valentino Peluso, Antonio Cipolletta, Andrea Calimera, Mat- teo Poggi, Fabio Tosi, Filippo Aleotti, and Stefano Mattoccia. Monocular depth perception on microcontrollers for edge applications. IEEE Transactions on Circuits and Systems for Video Technology, 32(3):1524–1536, 2021

  62. [70]

    Enabling energy-efficient unsupervised monocular depth estimation on armv7-based platforms

    Valentino Peluso, Antonio Cipolletta, Andrea Calimera, Mat- teo Poggi, Fabio Tosi, and Stefano Mattoccia. Enabling energy-efficient unsupervised monocular depth estimation on armv7-based platforms. In 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE), pag...

  63. [71]

    Excavating the potential capacity of self- supervised monocular depth estimation

    Rui Peng, Ronggang Wang, Yawen Lai, Luyang Tang, and Yangang Cai. Excavating the potential capacity of self- supervised monocular depth estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15560–15569, 2021

  64. [72]

    Unidepthv2: Universal monocular metric depth estimation made simpler

    Luigi Piccinelli, Christos Sakaridis, Yung-Hsu Yang, Mat- tia Segu, Siyuan Li, Wim Abbeloos, and Luc Van Gool. Unidepthv2: Universal monocular metric depth estimation made simpler. arXiv preprint arXiv:2502.20110, 2025

  65. [73]

    Unidepth: Universal monocular metric depth estimation

    Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segu, Siyuan Li, Luc Van Gool, and Fisher Yu. Unidepth: Universal monocular metric depth estimation. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10106–10116, 2024

  66. [74]

    Su- perDepth: Self-Supervised, Super-Resolved Monocular Depth Estimation

    Sudeep Pillai, Rare s ¸Ambrus ¸, and Adrien Gaidon. Su- perDepth: Self-Supervised, Super-Resolved Monocular Depth Estimation. In International Conference on Robotics and Automation, pages 9250–9256, 2019

  67. [75]

    Towards real-time unsupervised monocular depth es- timation on cpu

    Matteo Poggi, Filippo Aleotti, Fabio Tosi, and Stefano Mat- toccia. Towards real-time unsupervised monocular depth es- timation on cpu. In 2018 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 5848–5854. IEEE, 2018

  68. [76]

    On the Uncertainty of Self-Supervised Monocular Depth Estimation

    Matteo Poggi, Filippo Aleotti, Fabio Tosi, and Stefano Mat- toccia. On the Uncertainty of Self-Supervised Monocular Depth Estimation. In Conference on Computer Vision and Pattern Recognition, pages 3224–3234, 2020

  69. [77]

    On the synergies between machine learning and binocular stereo for depth estimation from images: a survey

    Matteo Poggi, Fabio Tosi, Konstantinos Batsos, Philippos Mordohai, and Stefano Mattoccia. On the synergies between machine learning and binocular stereo for depth estimation from images: a survey. IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 44(9):5314–5334, 2021

  70. [78]

    Learning Monocular Depth Estimation with Unsupervised Trinocular Assumptions

    Matteo Poggi, Fabio Tosi, and Stefano Mattoccia. Learning Monocular Depth Estimation with Unsupervised Trinocular Assumptions. In International Conference on 3D Vision , pages 324–333, 2018

  71. [79]

    Booster: a benchmark for depth from images of specular and transparent surfaces

    Pierluigi Zama Ramirez, Alex Costanzino, Fabio Tosi, Mat- teo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi 11 Di Stefano. Booster: a benchmark for depth from images of specular and transparent surfaces. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  72. [80]

    Open chal- lenges in deep stereo: the booster dataset

    Pierluigi Zama Ramirez, Fabio Tosi, Matteo Poggi, Samuele Salti, Stefano Mattoccia, and Luigi Di Stefano. Open chal- lenges in deep stereo: the booster dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21168–21178, 2022

  73. [81]

    Vi- sion transformers for dense prediction

    Ren´e Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vi- sion transformers for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 12179–12188, October 2021

  74. [82]

    Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer

    Ren´e Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocu- lar depth estimation: Mixing datasets for zero-shot cross- dataset transfer. IEEE transactions on pattern analysis and machine intelligence, 2020

  75. [83]

    Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation

    Anurag Ranjan, Varun Jampani, Lukas Balles, Kihwan Kim, Deqing Sun, Jonas Wulff, and Michael J Black. Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation. In Proceed- ings of the IEEE/CVF conference on computer v...

  76. [84]

    Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding

    Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. Hypersim: A photorealistic syn- thetic dataset for holistic indoor scene understanding. In Proceedings of the IEEE/CVF international conference o...

  77. [85]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022

  78. [86]

    Monocular depth esti- mation using neural regression forest

    Anirban Roy and Sinisa Todorovic. Monocular depth esti- mation using neural regression forest. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 5506–5514, 2016

  79. [87]

    Deep Virtual Stereo Odometry: Leveraging Deep Depth Pre- diction for Monocular Direct Sparse Odometry

    Rui, St ¨uckler J¨org, Cremers Daniel Yang Nan, and Wang. Deep Virtual Stereo Odometry: Leveraging Deep Depth Pre- diction for Monocular Direct Sparse Odometry. InEuropean Conference on Computer Vision, pages 835–852, 2018

  80. [88]

    Saurabh Saxena, Charles Herrmann, Junhwa Hur, Abhishek Kar, Mohammad Norouzi, Deqing Sun, and David J. Fleet. The surprising effectiveness of diffusion models for opti- cal flow and monocular depth estimation. arXiv preprint arXiv:2306.01923, 2023

  81. [89]

    Monocular depth estimation using diffusion models

    Saurabh Saxena, Abhishek Kar, Mohammad Norouzi, and David J Fleet. Monocular depth estimation using diffusion models. arXiv preprint arXiv:2302.14816, 2023

  82. [90]

    Swiftformer: Efficient additive attention for transformer- based real-time mobile vision applications

    Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Swiftformer: Efficient additive attention for transformer- based real-time mobile vision applications. In Proceedings of the IEEE/CVF international conference on computer ...

  83. [91]

    Learning temporally consistent video depth from video diffusion priors

    Jiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang, Yujun Shen, Vitor Guizilini, Yue Wang, Matteo Poggi, and Yiyi Liao. Learning temporally consistent video depth from video diffusion priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...

  84. [92]

    Denoising diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020

  85. [93]

    DeFeat-Net: General monocular depth via simultaneous unsupervised representation learning

    Jaime Spencer, Richard Bowden, and Simon Hadfield. DeFeat-Net: General monocular depth via simultaneous unsupervised representation learning. In Conference on Computer Vision and Pattern Recognition , pages 14390– 14401, 2020

  86. [94]

    The monoc- ular depth estimation challenge

    Jaime Spencer, C Stella Qian, Chris Russell, Simon Hadfield, Erich Graf, Wendy Adams, Andrew J Schofield, James H Elder, Richard Bowden, Heng Cong, et al. The monoc- ular depth estimation challenge. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer V...

  87. [95]

    The second monocular depth estimation challenge

    Jaime Spencer, C Stella Qian, Michaela Trescakova, Chris Russell, Simon Hadfield, Erich W Graf, Wendy J Adams, Andrew J Schofield, James Elder, Richard Bowden, et al. The second monocular depth estimation challenge. In Pro- ceedings of the IEEE/CVF Conference on Computer Visio...

  88. [96]

    Deconstructing self-supervised monocular recon- struction: The design decisions that matter

    Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden. Deconstructing self-supervised monocular recon- struction: The design decisions that matter. Transactions on Machine Learning Research, 2022. Reproducibility Certifi- cation

  89. [97]

    Kick back & relax: Learning to reconstruct the world by watching slowtv

    Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden. Kick back & relax: Learning to reconstruct the world by watching slowtv. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15768–15779, October 2023

  90. [98]

    Kick back & relax++: Scaling beyond ground- truth depth with slowtv & cribstv

    Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden. Kick back & relax++: Scaling beyond ground- truth depth with slowtv & cribstv. arXiv preprint arXiv:2403.01569, 2024

  91. [99]

    Jaime Spencer, Fabio Tosi, Matteo Poggi, Ripudaman Singh Arora, Chris Russell, Simon Hadfield, Richard Bowden, Guangyuan Zhou, Zhengxin Li, Qiang Rao, Yiping Bao, Xiao Liu, Dohyeong Kim, Jinseong Kim, Myunghyun Kim, Mykola Lavreniuk, Rui Li, Qing Mao, Jiang Wu, Yu Zhu, Jinqiu ...

  92. [100]

    Sturm, N

    J. Sturm, N. Engelhard, F. Endres, W. Burgard, and D. Cre- mers. A benchmark for the evaluation of rgb-d slam systems. In Proc. of the International Conference on Intelligent Robot Systems (IROS), Oct. 2012

  93. [101]

    Mind the edge: Refining depth edges in sparsely-supervised monocular depth estimation

    Lior Talker, Aviad Cohen, Erez Yosef, Alexandra Dana, and Michael Dinerstein. Mind the edge: Refining depth edges in sparsely-supervised monocular depth estimation. In Conference on Computer Vision and Pattern Recognition, 12

  94. [102]

    Learning monocular depth estimation infusing tra- ditional stereo knowledge

    Fabio Tosi, Filippo Aleotti, Matteo Poggi, and Stefano Mat- toccia. Learning monocular depth estimation infusing tra- ditional stereo knowledge. Conference on Computer Vision and Pattern Recognition, 2019-June:9791–9801, 2019

  95. [103]

    Distilled semantics for comprehensive scene under- standing from videos

    Fabio Tosi, Filippo Aleotti, Pierluigi Zama Ramirez, Matteo Poggi, Samuele Salti, Luigi Di Stefano, and Stefano Mat- toccia. Distilled semantics for comprehensive scene under- standing from videos. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern reco...

  96. [104]

    Smd-nets: Stereo mixture density networks

    Fabio Tosi, Yiyi Liao, Carolin Schmitt, and Andreas Geiger. Smd-nets: Stereo mixture density networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8942–8952, 2021

  97. [105]

    Dif- fusion models for monocular depth estimation: Overcoming challenging conditions

    Fabio Tosi, Pierluigi Zama Ramirez, and Matteo Poggi. Dif- fusion models for monocular depth estimation: Overcoming challenging conditions. In European Conference on Com- puter Vision (ECCV), 2024

  98. [106]

    Unsupervised monocu- lar depth estimation for night-time images using adversar- ial domain feature adaptation

    Madhu Vankadari, Sourav Garg, Anima Majumder, Swa- gat Kumar, and Ardhendu Behera. Unsupervised monocu- lar depth estimation for night-time images using adversar- ial domain feature adaptation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, ...

  99. [107]

    When the sun goes down: Repairing photometric losses for all-day depth estimation

    Madhu Vankadari, Stuart Golodetz, Sourav Garg, Sangyun Shin, Andrew Markham, and Niki Trigoni. When the sun goes down: Repairing photometric losses for all-day depth estimation. In Conference on Robot Learning, pages 1992–

  100. [108]

    Dai, Andrea F

    Igor Vasiljevic, Nick Kolkin, Shanyi Zhang, Ruotian Luo, Haochen Wang, Falcon Z. Dai, Andrea F. Daniele, Moham- madreza Mostajabi, Steven Basart, Matthew R. Walter, and Gregory Shakhnarovich. DIODE: A Dense Indoor and Out- door DEpth Dataset. CoRR, abs/1908.00463, 2019

  101. [109]

    Marigold-dc: Zero-shot monocular depth completion with guided diffusion, 2024

    Massimiliano Viola, Kevin Qu, Nando Metzger, Bingxin Ke, Alexander Becker, Konrad Schindler, and Anton Obukhov. Marigold-dc: Zero-shot monocular depth completion with guided diffusion, 2024

  102. [110]

    Learning Depth from Monocular Videos Using Direct Methods

    Chaoyang Wang, Jose Miguel Buenaposada, Rui Zhu, and Simon Lucey. Learning Depth from Monocular Videos Using Direct Methods. Conference on Computer Vision and Pattern Recognition, pages 2022–2030, 2018

  103. [111]

    Web stereo video supervision for depth prediction from dynamic scenes

    Chaoyang Wang, Simon Lucey, Federico Perazzi, and Oliver Wang. Web stereo video supervision for depth prediction from dynamic scenes. In 2019 International Conference on 3D Vision (3DV), pages 348–357. IEEE, 2019

  104. [112]

    Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision

    Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, and Jiaolong Yang. Moge: Unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. arXiv preprint arXiv:2410.19115, 2024

  105. [113]

    Dust3r: Geometric 3d vision made easy

    Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, pages 20697–20709, 2024

  106. [114]

    Self-supervised monocular depth hints

    Jamie Watson, Michael Firman, Gabriel Brostow, and Dani- yar Turmukhambetov. Self-supervised monocular depth hints. International Conference on Computer Vision, 2019- Octob:2162–2171, 2019

  107. [115]

    Fastdepth: Fast monocular depth estima- tion on embedded systems

    Diana Wofk, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, and Vivienne Sze. Fastdepth: Fast monocular depth estima- tion on embedded systems. In 2019 International Confer- ence on Robotics and Automation (ICRA), pages 6101–6108. IEEE, 2019

  108. [116]

    Monocular relative depth per- ception with web stereo data supervision

    Ke Xian, Chunhua Shen, Zhiguo Cao, Hao Lu, Yang Xiao, Ruibo Li, and Zhenbo Luo. Monocular relative depth per- ception with web stereo data supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 311–320, 2018

  109. [117]

    Channel-Wise Attention-Based Network for Self-Supervised Monocular Depth Estimation

    Jiaxing Yan, Hong Zhao, Penghui Bu, and YuSheng Jin. Channel-Wise Attention-Based Network for Self-Supervised Monocular Depth Estimation. In International Conference on 3D Vision, pages 464–473, 2021

  110. [118]

    Depth anything: Unleashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In CVPR, 2024

  111. [119]

    Depth anything v2

    Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xi- aogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth anything v2. Advances in Neural Information Processing Systems, 37:21875–21911, 2024

  112. [120]

    Unimatch v2: Pushing the limit of semi-supervised semantic segmen- tation

    Lihe Yang, Zhen Zhao, and Hengshuang Zhao. Unimatch v2: Pushing the limit of semi-supervised semantic segmen- tation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  113. [121]

    D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry

    Nan Yang, Lukas von Stumberg, Rui Wang, and Daniel Cre- mers. D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry. In Conference on Com- puter Vision and Pattern Recognition , pages 1278–1289, 2020

  114. [122]

    Di- versedepth: Affine-invariant depth prediction using diverse data

    Wei Yin, Xinlong Wang, Chunhua Shen, Yifan Liu, Zhi Tian, Songcen Xu, Changming Sun, and Dou Renyin. Di- versedepth: Affine-invariant depth prediction using diverse data. arXiv preprint arXiv:2002.00569, 2020

  115. [123]

    Met- ric3D: Towards zero-shot metric 3d prediction from a single image

    Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. Met- ric3D: Towards zero-shot metric 3d prediction from a single image. In ICCV, 2023

  116. [124]

    Towards accurate reconstruction of 3d scene shape from a single monocular image

    Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Si- mon Chen, Yifan Liu, and Chunhua Shen. Towards accurate reconstruction of 3d scene shape from a single monocular image. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(5):6480–6494, 2022

  117. [125]

    Learning to recover 3d scene shape from a single image

    Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 204–213, 2021

  118. [126]

    Learning to recover 3d scene shape from a single image

    Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen. Learning to recover 3d scene shape from a single image. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 204–213, June 2021

  119. [127]

    Geonet: Unsupervised learn- ing of dense depth, optical flow and camera pose

    Zhichao Yin and Jianping Shi. Geonet: Unsupervised learn- ing of dense depth, optical flow and camera pose. In Pro- ceedings of the IEEE conference on computer vision and 13 pattern recognition, pages 1983–1992, 2018

  120. [128]

    Bdd100k: A diverse driving dataset for heterogeneous multitask learning

    Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Dar- rell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition,...

  121. [129]

    Neural window fully-connected crfs for monoc- ular depth estimation

    Weihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu, and Ping Tan. Neural window fully-connected crfs for monoc- ular depth estimation. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3906–3915, 2022

  122. [130]

    NTIRE 2023 challenge on HR depth from images of specular and transparent surfaces

    Pierluigi Zama Ramirez, Tosi Fabio, Luigi Di Stefano, Radu Timofte, Alex Costanzino, Matteo Poggi, Samuele Salti, Ste- fano Mattoccia, Jun Shi, Dafeng Zhang, Yong A, Yixiang Jin, Dingzhe Li, Chao Li, Zhiwen Liu, Qi Zhang, Yixing Wang, and Shi Yin. NTIRE 2023 challenge on HR de...

  123. [131]

    Geometry meets seman- tics for semi-supervised monocular depth estimation

    Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi, Stefano Mattoccia, and Luigi Di Stefano. Geometry meets seman- tics for semi-supervised monocular depth estimation. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, R...

  124. [132]

    NTIRE 2024 challenge on HR depth from images of specular and transpar- ent surfaces

    Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano, Radu Timofte, Alex Costanzino, Matteo Poggi, et al. NTIRE 2024 challenge on HR depth from images of specular and transpar- ent surfaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...

  125. [133]

    NTIRE 2025 challenge on HR depth from images of specular and transpar- ent surfaces

    Pierluigi Zama Ramirez, Fabio Tosi, Luigi Di Stefano, Radu Timofte, Alex Costanzino, Matteo Poggi, et al. NTIRE 2025 challenge on HR depth from images of specular and transpar- ent surfaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...

  126. [134]

    Huangying Zhan, Ravi Garg, Chamara Saroj Weerasekera, Kejie Li, Harsh Agarwal, and Ian M. Reid. Unsupervised Learning of Monocular Depth Estimation and Visual Odom- etry with Deep Feature Reconstruction. Conference on Com- puter Vision and Pattern Recognition, pages 340–349, 2018

  127. [135]

    Hierarchical normalization for robust monoc- ular depth estimation

    Chi Zhang, Wei Yin, Billzb Wang, Gang Yu, Bin Fu, and Chunhua Shen. Hierarchical normalization for robust monoc- ular depth estimation. NIPS, 35, 2022

  128. [136]

    Betterdepth: Plug-and-play diffu- sion refiner for zero-shot monocular depth estimation

    Xiang Zhang, Bingxin Ke, Hayko Riemenschneider, Nando Metzger, Anton Obukhov, Markus Gross, Konrad Schindler, and Christopher Schroers. Betterdepth: Plug-and-play diffu- sion refiner for zero-shot monocular depth estimation. arXiv preprint arXiv:2407.17952, 2024

  129. [137]

    Unsupervised monocular depth estimation in highly complex environments

    Chaoqiang Zhao, Yang Tang, and Qiyu Sun. Unsupervised monocular depth estimation in highly complex environments. IEEE Transactions on Emerging Topics in Computational Intelligence, 6(5):1237–1246, 2022

  130. [138]

    Monovit: Self-supervised monocular depth estimation with a vision transformer

    Chaoqiang Zhao, Youmin Zhang, Matteo Poggi, Fabio Tosi, Xianda Guo, Zheng Zhu, Guan Huang, Yang Tang, and Stefano Mattoccia. Monovit: Self-supervised monocular depth estimation with a vision transformer. International Conference on 3D Vision, 2022

  131. [139]

    Self- Supervised Monocular Depth Estimation with Internal Fea- ture Fusion

    Hang Zhou, David Greenwood, and Sarah Taylor. Self- Supervised Monocular Depth Estimation with Internal Fea- ture Fusion. In British Machine Vision Conference, 2021

  132. [140]

    Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. Unsupervised Learning of Depth and Ego-Motion from Video. Conference on Computer Vision and Pattern Recognition, pages 6612–6619, 2017. 14

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.