REVIEW 9 cited by
DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Monocular depth estimation is a challenging task that predicts the pixel-wise depth from a single 2D image. Current methods typically model this problem as a regression or classification task. We propose DiffusionDepth, a new approach that reformulates monocular depth estimation as a denoising diffusion process. It learns an iterative denoising process to `denoise' random depth distribution into a depth map with the guidance of monocular visual conditions. The process is performed in the latent space encoded by a dedicated depth encoder and decoder. Instead of diffusing ground truth (GT) depth, the model learns to reverse the process of diffusing the refined depth of itself into random depth distribution. This self-diffusion formulation overcomes the difficulty of applying generative models to sparse GT depth scenarios. The proposed approach benefits this task by refining depth estimation step by step, which is superior for generating accurate and highly detailed depth maps. Experimental results on KITTI and NYU-Depth-V2 datasets suggest that a simple yet efficient diffusion approach could reach state-of-the-art performance in both indoor and outdoor scenarios with acceptable inference time.
Forward citations
Cited by 9 Pith papers
-
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
CrossFlow turns text directly into images, and images into text, depth, and higher resolution, by flowing between modality latents without a noise prior or cross-attention.
-
Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail
A dual-branch stereo network fusing stereo correlation volumes with monocular depth foundation model priors achieves state-of-the-art zero-shot generalization, including on mirrors and transparencies.
-
RoughNet: Mapping Arctic Sea Ice Roughness Using Diffusion-Based Super-Resolution of Satellite Imagery
A conditional diffusion model maps Sentinel-2 optical imagery to 1 m sea-ice roughness residuals with ~9 cm RMSE on an unseen region, though pointwise correlation is weak (ZNCC≈0.11).
-
Distillation of Diffusion Features for Semantic Correspondence
A DINOv2 student trained with LoRA to imitate DINOv2-plus-SDXL-Turbo similarity maps, then fine-tuned on 3D-derived correspondences, sets new state-of-the-art on three semantic correspondence benchmarks.
-
Video Depth without Video Models
A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.
-
Scalable Autoregressive Monocular Depth Estimation
An autoregressive transformer that predicts depth maps in low-to-high resolution steps with recursively refined depth bins achieves state-of-the-art monocular depth accuracy on KITTI and NYU Depth v2.
-
Scaling Properties of Diffusion Models for Perceptual Tasks
Diffusion models for depth, optical flow, and amodal segmentation improve along power laws as training and test-time compute scale, and the fitted recipes match prior specialist models with less data.
-
Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion
Sparse depth points injected as test-time guidance into a pretrained monocular depth diffusion model achieve strong zero-shot depth completion across indoor and outdoor scenes.
-
A Critical Synthesis of Uncertainty Quantification and Foundation Models in Monocular Depth Estimation
Fine-tuning DepthAnythingV2 with a Gaussian negative log-likelihood loss yields the most reliable pixel-wise uncertainty estimates on indoor, street, and object scenes, but it fails on aerial large-depth data.
Discussion (0). Continue with ORCID to comment.