REVIEW 3 cited by
Stealing Stable Diffusion Prior for Robust Monocular Depth Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Monocular depth estimation is a crucial task in computer vision. While existing methods have shown impressive results under standard conditions, they often face challenges in reliably performing in scenarios such as low-light or rainy conditions due to the absence of diverse training data. This paper introduces a novel approach named Stealing Stable Diffusion (SSD) prior for robust monocular depth estimation. The approach addresses this limitation by utilizing stable diffusion to generate synthetic images that mimic challenging conditions. Additionally, a self-training mechanism is introduced to enhance the model's depth estimation capability in such challenging environments. To enhance the utilization of the stable diffusion prior further, the DINOv2 encoder is integrated into the depth model architecture, enabling the model to leverage rich semantic priors and improve its scene understanding. Furthermore, a teacher loss is introduced to guide the student models in acquiring meaningful knowledge independently, thus reducing their dependency on the teacher models. The effectiveness of the approach is evaluated on nuScenes and Oxford RobotCar, two challenging public datasets, with the results showing the efficacy of the method. Source code and weights are available at: https://github.com/hitcslj/SSD.
Forward citations
Cited by 3 Pith papers
-
Depth Anything at Any Condition
A fine-tuned Depth Anything V2 model using perturbation consistency and spatial distance constraints improves monocular depth estimation under adverse conditions without any labeled data.
-
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
Driver gaze and accident-reason text are used to train a video diffusion model that can edit and generate egocentric crash videos with the correct causal participants, with a new large gaze dataset for accidents.
-
GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstruction
A 2D Gaussian Splatting method that uses foundation-model depth/normal priors plus deferred shading to improve reconstruction and relighting of reflective objects.
Discussion (0). Continue with ORCID to comment.