Pith. sign in

REVIEW 9 cited by

DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.05021 v4 pith:NMH5FBKY submitted 2023-03-09 cs.CV

classification cs.CV
keywords depthapproachestimationmonocularprocessdenoisingdiffusiontask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Monocular depth estimation is a challenging task that predicts the pixel-wise depth from a single 2D image. Current methods typically model this problem as a regression or classification task. We propose DiffusionDepth, a new approach that reformulates monocular depth estimation as a denoising diffusion process. It learns an iterative denoising process to `denoise' random depth distribution into a depth map with the guidance of monocular visual conditions. The process is performed in the latent space encoded by a dedicated depth encoder and decoder. Instead of diffusing ground truth (GT) depth, the model learns to reverse the process of diffusing the refined depth of itself into random depth distribution. This self-diffusion formulation overcomes the difficulty of applying generative models to sparse GT depth scenarios. The proposed approach benefits this task by refining depth estimation step by step, which is superior for generating accurate and highly detailed depth maps. Experimental results on KITTI and NYU-Depth-V2 datasets suggest that a simple yet efficient diffusion approach could reach state-of-the-art performance in both indoor and outdoor scenarios with acceptable inference time.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

    cs.CV 2024-12 conditional novelty 7.0 of 10

    CrossFlow turns text directly into images, and images into text, depth, and higher resolution, by flowing between modality latents without a noise prior or cross-attention.

  2. Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A dual-branch stereo network fusing stereo correlation volumes with monocular depth foundation model priors achieves state-of-the-art zero-shot generalization, including on mirrors and transparencies.

  3. RoughNet: Mapping Arctic Sea Ice Roughness Using Diffusion-Based Super-Resolution of Satellite Imagery

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A conditional diffusion model maps Sentinel-2 optical imagery to 1 m sea-ice roughness residuals with ~9 cm RMSE on an unseen region, though pointwise correlation is weak (ZNCC≈0.11).

  4. Distillation of Diffusion Features for Semantic Correspondence

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A DINOv2 student trained with LoRA to imitate DINOv2-plus-SDXL-Turbo similarity maps, then fine-tuned on 3D-derived correspondences, sets new state-of-the-art on three semantic correspondence benchmarks.

  5. Video Depth without Video Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.

  6. Scalable Autoregressive Monocular Depth Estimation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An autoregressive transformer that predicts depth maps in low-to-high resolution steps with recursively refined depth bins achieves state-of-the-art monocular depth accuracy on KITTI and NYU Depth v2.

  7. Scaling Properties of Diffusion Models for Perceptual Tasks

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Diffusion models for depth, optical flow, and amodal segmentation improve along power laws as training and test-time compute scale, and the fitted recipes match prior specialist models with less data.

  8. Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Sparse depth points injected as test-time guidance into a pretrained monocular depth diffusion model achieve strong zero-shot depth completion across indoor and outdoor scenes.

  9. A Critical Synthesis of Uncertainty Quantification and Foundation Models in Monocular Depth Estimation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    Fine-tuning DepthAnythingV2 with a Gaussian negative log-likelihood loss yields the most reliable pixel-wise uncertainty estimates on indoor, street, and object scenes, but it fails on aerial large-depth data.

Pith tools