REVIEW 16 cited by
From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Estimating accurate depth from a single image is challenging because it is an ill-posed problem as infinitely many 3D scenes can be projected to the same 2D scene. However, recent works based on deep convolutional neural networks show great progress with plausible results. The convolutional neural networks are generally composed of two parts: an encoder for dense feature extraction and a decoder for predicting the desired depth. In the encoder-decoder schemes, repeated strided convolution and spatial pooling layers lower the spatial resolution of transitional outputs, and several techniques such as skip connections or multi-layer deconvolutional networks are adopted to recover the original resolution for effective dense prediction. In this paper, for more effective guidance of densely encoded features to the desired depth prediction, we propose a network architecture that utilizes novel local planar guidance layers located at multiple stages in the decoding phase. We show that the proposed method outperforms the state-of-the-art works with significant margin evaluating on challenging benchmarks. We also provide results from an ablation study to validate the effectiveness of the proposed method.
Forward citations
Cited by 16 Pith papers
-
The Multipath Blind Spot: $K$-Agnostic Robust Calibration for Sparse-Anchor Metric Depth from Frozen Foundations
MRAC gates sparse anchors via Theil–Sen + MAD consistency with a frozen foundation's relative depth, repairing multipath outliers that collapse residual-on-CFA and blind VI-Depth while winning 84% of same-backbone cells.
-
Twins: Learn to Predict Unified Representations with Focal Loss
Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.
-
Boosting Robustness for All-Weather Self-Supervised Depth Estimation in Autonomous Driving
Uncertainty-weighted multi-teacher distillation plus dense bird's-eye-view radar fusion improves self-supervised depth estimation under adverse weather, cutting night absRel by ~23% on nuScenes.
-
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras
X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.
-
G3Splat: Geometrically Consistent Generalizable Gaussian Splatting
Adding ray-alignment and local-normal orientation losses to generalizable Gaussian splatting fixes geometrically degenerate splats and improves zero-shot depth, mesh, and pose estimation.
-
Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation
Across 69 monocular depth estimators, human-likeness of error patterns peaks near human-level accuracy and declines for the most accurate models: accuracy does not guarantee human-like depth perception.
-
XD-RCDepth: Lightweight Radar-Camera Depth Estimation with Explainability-Aligned and Distribution-Aware Distillation
XD-RCDepth uses explainability-aligned and depth-distribution distillation to shrink a radar-camera depth model by 29.7% parameters while improving MAE by about 8%.
-
Compact and robust optical frequency reference module based on reproducible and redistributable optical design
A reproducible, compact optical frequency reference module is claimed to maintain frequency stability for months with 4g vibration tolerance, based on openly shared design files.
-
Perfecting Depth: Uncertainty-Aware Enhancement of Metric Depth
Perfecting Depth is a two-stage pipeline that uses diffusion-sample variance to flag unreliable depth pixels and a deterministic network to refine them, beating monocular baselines on indoor depth inpainting and noisy...
-
BadDepth: Backdoor Attacks Against Monocular Depth Estimation in the Physical World
BadDepth uses poisoned depth labels and physical-world image augmentation to make a triggered object vanish from monocular depth predictions.
-
DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV
DAPM jointly estimates depth and camera pose from single drone images by injecting an ideal-ground-plane depth prior and progressive per-pixel depth bins, trained on a new 42k simulated UAV dataset.
-
BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
A single network that iteratively aligns monocular features with stereo hypotheses reduces zero-shot stereo depth error by over 40% on Middlebury and ETH3D.
-
Region-aware Depth Scale Adaptation with Sparse Measurements
A non-learning method segments an image and gives each region its own scale and shift, fitted to a few sparse depth points, to turn relative monocular depth predictions into metric depth more accurately than a single ...
-
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
A training-free diffusion-guidance framework that couples scale alignment across windows and geometric multi-view constraints inside the denoising loop yields more scale- and geometry-consistent depth for long videos.
-
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
A LiDAR-intensity projection, matched to the camera image with an attention-based detector-free network and a repeatability score, achieves state-of-the-art point-pixel registration using only single-frame LiDAR.
-
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
The thesis demonstrates that combining implicit 3D scene representations with LLM-based reasoning, using text as an interface, yields strong performance on robotic perception and spatial language tasks.
Discussion (0). Sign in to comment.