Pith. sign in

REVIEW 6 cited by

DiverseDepth: Affine-invariant Depth Prediction Using Diverse Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.00569 v3 pith:OSRKC3EX submitted 2020-02-03 cs.CV

classification cs.CV
keywords depthdiversedatasetscenesgeneralizationlearningmethodscene
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a method for depth estimation with monocular images, which can predict high-quality depth on diverse scenes up to an affine transformation, thus preserving accurate shapes of a scene. Previous methods that predict metric depth often work well only for a specific scene. In contrast, learning relative depth (information of being closer or further) can enjoy better generalization, with the price of failing to recover the accurate geometric shape of the scene. In this work, we propose a dataset and methods to tackle this dilemma, aiming to predict accurate depth up to an affine transformation with good generalization to diverse scenes. First we construct a large-scale and diverse dataset, termed Diverse Scene Depth dataset (DiverseDepth), which has a broad range of scenes and foreground contents. Compared with previous learning objectives, i.e., learning metric depth or relative depth, we propose to learn the affine-invariant depth using our diverse dataset to ensure both generalization and high-quality geometric shapes of scenes. Furthermore, in order to train the model on the complex dataset effectively, we propose a multi-curriculum learning method. Experiments show that our method outperforms previous methods on 8 datasets by a large margin with the zero-shot test setting, demonstrating the excellent generalization capacity of the learned model to diverse scenes. The reconstructed point clouds with the predicted depth show that our method can recover high-quality 3D shapes. Code and dataset are available at: https://tinyurl.com/DiverseDepth

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution

    cs.CV 2024-12 conditional novelty 7.0 of 10

    CrossFlow turns text directly into images, and images into text, depth, and higher resolution, by flowing between modality latents without a noise prior or cross-attention.

  2. Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A dual-branch stereo network fusing stereo correlation volumes with monocular depth foundation model priors achieves state-of-the-art zero-shot generalization, including on mirrors and transparencies.

  3. FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FiffDepth transforms a pre-trained diffusion image generator into a feed-forward monocular depth estimator that combines generative detail with DINOv2-based robustness.

  4. Video Depth without Video Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.

  5. One Diffusion to Generate Them All

    cs.CV 2024-11 conditional novelty 6.0 of 10

    OneDiffusion shows that a single 2.8B-parameter diffusion model, trained by treating all tasks as frame sequences with varying noise scales, can handle image generation and image understanding tasks bidirectionally.

  6. Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Sparse depth points injected as test-time guidance into a pretrained monocular depth diffusion model achieve strong zero-shot depth completion across indoor and outdoor scenes.

Pith tools