Pith. sign in

REVIEW 2 cited by

The Surprising Effectiveness of Diffusion Models for Optical Flow and Monocular Depth Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01923 v2 pith:3Y66V5ZE submitted 2023-06-02 cs.CV

classification cs.CV
keywords diffusiondepthflowmodelsopticaldenoisingmodeltraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Denoising diffusion probabilistic models have transformed image generation with their impressive fidelity and diversity. We show that they also excel in estimating optical flow and monocular depth, surprisingly, without task-specific architectures and loss functions that are predominant for these tasks. Compared to the point estimates of conventional regression-based methods, diffusion models also enable Monte Carlo inference, e.g., capturing uncertainty and ambiguity in flow and depth. With self-supervised pre-training, the combined use of synthetic and real data for supervised training, and technical innovations (infilling and step-unrolled denoising diffusion training) to handle noisy-incomplete training data, and a simple form of coarse-to-fine refinement, one can train state-of-the-art diffusion models for depth and optical flow estimation. Extensive experiments focus on quantitative performance against benchmarks, ablations, and the model's ability to capture uncertainty and multimodality, and impute missing values. Our model, DDVM (Denoising Diffusion Vision Model), obtains a state-of-the-art relative depth error of 0.074 on the indoor NYU benchmark and an Fl-all outlier rate of 3.26\% on the KITTI optical flow benchmark, about 25\% better than the best published method. For an overview see https://diffusion-vision.github.io.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A dual-branch stereo network fusing stereo correlation volumes with monocular depth foundation model priors achieves state-of-the-art zero-shot generalization, including on mirrors and transparencies.

  2. Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Sparse depth points injected as test-time guidance into a pretrained monocular depth diffusion model achieve strong zero-shot depth completion across indoor and outdoor scenes.

Pith tools