Pith. sign in

REVIEW 3 major objections 5 minor 87 references

A unified stereo framework couples a feed-forward stereo matcher with a diffusion normal branch; the pair improves both disparity and normals, achieving top zero-shot results on major benchmarks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-07-31 23:13 UTC pith:U4C3WOOC

load-bearing objection A promising and novel coupling of stereo matching and diffusion-based normal estimation, but the main ablation evidence is in-distribution and the headline disparity gains over FoundationStereo are tiny. the 3 major comments →

arxiv 2607.24024 v1 pith:U4C3WOOC submitted 2026-07-27 cs.CV

A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal

classification cs.CV
keywords stereo matchingsurface normal estimationdiffusion priorszero-shot generalizationdisparity-to-normal initializationjoint geometry estimationcross-view conditioninglatent diffusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper tries to establish that stereo matching and surface normal estimation are better solved together than apart: a feed-forward stereo network supplies a coarse geometric starting point, and a diffusion network refines normals while feeding structural priors back into disparity. The two branches are coupled by converting the initial disparity into a provisional normal map and by warping the right image into the left-view frame, both in a differentiable way. If the approach is right, the diffusion prior rescues disparity in low-light, reflective, and transparent regions where correspondence is ambiguous, while the stereo geometry anchors normal prediction. The reported zero-shot results across indoor and outdoor benchmarks support the claim's plausibility.

Core claim

The central claim is that a unified model with two complementary paradigms beats either paradigm alone. The disparity branch's coarse geometry is turned into a normal prior via a pseudo-depth log-warp and central differences, and the right view is warped to the left view with the current disparity to give aligned cross-view guidance; the diffusion U-Net then refines the normal latent conditioned on both. Because both operations are differentiable, the normal loss propagates back into the stereo branch, so diffusion priors regularize disparity in ill-posed regions while stereo geometry keeps normals grounded. On zero-shot benchmarks the paper reports the lowest disparity errors among compared

What carries the argument

The central mechanism is the two differentiable couplings between branches. First, the disparity-to-normal initialization builds a coarse normal prior from the initial disparity: after clipping, it forms Z = -log(D/d_max), constructs a pseudo point map P(u,v) = (x(u)Z, y(v)Z, Z), and estimates normals from the cross product of central-difference tangents, normalized to a consistent orientation. Second, the warp-to-left-view condition samples the right image at u - D_init to produce a spatially aligned cross-view input. Both operations are differentiable, so the normal loss from the x0-reparameterized latent diffusion U-Net flows back into the stereo branch, letting diffusion priors regulariz

Load-bearing premise

The load-bearing premise is that the coarse normal prior computed from log-warped disparity (Section 3.3) is accurate enough that the diffusion branch refines rather than fights it; if that proxy is systematically biased on slanted or curved surfaces, the normal gradients flowing back into the stereo branch could distort disparity instead of helping.

What would settle it

On a synthetic benchmark with ground-truth normals and disparity dominated by slanted and curved surfaces, compute the D2N proxy's angular error as a function of surface slant; if the proxy error is large on such surfaces and the final normal error tracks it, or if removing D2N improves disparity EPE on those scenes during joint training, the paper's central coupling claim would be undermined.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Zero-shot disparity estimation improves to the best reported EPE and D1 among compared methods on KITTI, NYUv2, Hypersim, DREDS, and ScanNet.
  • Surface normals become more accurate on indoor benchmarks, with the best accuracy-threshold scores on iBims-1 and the best mean and median angular errors on ScanNet.
  • In low-light, reflective, and transparent regions, the diffusion prior stabilizes disparity where correspondence is ambiguous.
  • Because the coupling is differentiable, normal supervision back-propagates into the stereo branch, making disparity and normals geometrically consistent rather than independently predicted.
  • A two-stage training scheme — separate optimization followed by joint consistency training — is needed to align the two branches.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the same differentiable coupling should transfer to metric-depth-plus-normal estimation, since the log-warp and warp-to-left-view operations are not specific to disparity; a monocular or multi-view depth network could replace the stereo branch.
  • Paper-stated limitation: the conclusion notes the diffusion refinement branch adds inference cost; a testable extension is distilling the diffusion branch or applying it only where disparity confidence is low.
  • Editorial inference: the D2N proxy's log-warp breaks normal invariance under depth scaling; on curved or slanted surfaces the proxy may be biased, so ablating it against a ground-truth-derived normal prior would test how much of the gain comes from the coupling versus the diffusion prior alone.
  • Editorial inference: the method's zero-shot gains suggest the normal branch acts as an unsupervised regularizer for disparity; one could quantify this by training the stereo branch alone on the same data and comparing error distributions in ill-posed regions only.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. GeoStereo couples a FoundationStereo-initialized feed-forward stereo matching branch with a diffusion-based surface-normal branch. The initial disparity is converted into a coarse normal prior (Eqs. 9–11), used to warp the right image into the left view (Eq. 12), and both differentiable pathways back-propagate normal-supervision gradients into the stereo branch. Joint training follows Eq. 14. The paper reports competitive zero-shot disparity and normal results on multiple benchmarks and argues that the diffusion branch supplies structural priors that improve disparity in ill-posed regions. The central claim is that the coupled design, not merely the assembled system, is what delivers the gains.

Significance. If the coupling claim is correct, GeoStereo would be a useful demonstration that generative diffusion priors can regularize feed-forward stereo in ambiguous regions while stereo geometry anchors normal prediction. The paper targets a real limitation of current stereo foundation models and evaluates across several public benchmarks, which is a strength. However, the paper does not provide a controlled experiment that isolates the coupling benefit on the final FoundationStereo-based system or on a truly held-out domain, so the load-bearing causal claim remains unsupported. The geometric validity of the log-disparity normal proxy is also not established. The work is potentially significant, but the evidence currently falls short of the strength of the claims.

major comments (3)
  1. [§4.4, Table 3] The central claim that the diffusion branch improves disparity is supported only by ablations run with LightStereo as the stereo backbone (Table 3), while the final system is initialized from FoundationStereo (§4.1). The improvement from EPE 2.4440 to 1.1091 on IRS is therefore measured on a different architecture. The diffusion-context contribution to the actual reported system is not isolated. Please repeat the with/without-D2N and with/without-diffusion ablations using the FoundationStereo backbone, or provide evidence that the LightStereo result transfers.
  2. [§4.2 and §4.4, Tables 3–5] The ablation studies are described as 'evaluated in the zero-shot regime,' but IRS and 3D Ken Burns are both listed in the disparity training set (~1.5M pairs) and in the normal training set (0.6M pairs). These experiments therefore measure in-distribution fitting, not zero-shot generalization. Since the paper's headline advantage is zero-shot robustness in ill-posed regions, the coupling benefit must be demonstrated on a dataset absent from the training list, e.g., one of the Table 1 or Table 2 benchmarks. Without such a held-out ablation, the claim that diffusion priors improve disparity in unseen challenging domains is not supported.
  3. [§3.3, Eqs. (9)–(11)] The disparity-to-normal initialization defines Z = -log(D/d_max) and computes normals from P(u,v) = (x(u)Z, y(v)Z, Z). Surface normals are not invariant under this log-warp of depth: for a fronto-parallel or slanted plane, the pseudo-normal computed from log-disparity differs from the true normal by a depth-dependent amount. The paper asserts that this proxy is a 'geometry-aware prior' and that back-propagated normal gradients provide reliable guidance, but no experiment or derivation validates the proxy's unbiasedness for slanted or curved surfaces. Please provide a quantitative comparison of proxy versus true normals on held-out surfaces, or justify why the bias does not undermine the coupling mechanism.
minor comments (5)
  1. [§4.1] The 'self-collected dataset' used in disparity training is not described. Please specify its size, acquisition method, and licensing/source details, since it is part of the training mixture and could affect reproducibility.
  2. [Table 1] The headline improvements over FoundationStereo are small (e.g., KITTI EPE 0.8642 vs 0.8679; NYUv2 1.3426 vs 1.3922). No error bars, multiple seeds, or statistical significance tests are reported, so it is unclear whether these differences are meaningful.
  3. [§4.4, Table 4] The warp-to-left-view ablation evaluates only normal accuracy. Since the warp pathway also back-propagates gradients to disparity, a disparity-error column would strengthen the argument that this conditioning contributes to stereo improvement.
  4. [§4.4] The text says 'all ablation variants are trained with the same training set and the same number of training iterations,' but the training schedule, batch size, and diffusion sampling steps for the ablations are not specified. Please add these details.
  5. [§5] The conclusion acknowledges the efficiency cost of the diffusion branch but does not report runtime or parameter counts. Given the stated future-work direction, including at least inference latency and GPU memory would help position the method.

Circularity Check

0 steps flagged

No significant circularity: the method is an empirically evaluated coupled-training system; no equation or fitted parameter reduces to its own input. The main caveat is that the direct coupling ablations run on training-domain datasets, which weakens the zero-shot causal claim but is not a definitional circularity.

full rationale

GeoStereo's derivation chain is an architecture design plus empirical evaluation, not a formal derivation that collapses into its inputs. The disparity-to-normal initialization (Eqs. 9-11) is a deterministic, differentiable mapping from D_init to N_init; the normal objective (Eq. 13) back-propagates through this mapping, but this is a standard auxiliary-gradient coupling rather than a self-justifying definition. The diffusion branch is initialized from Stable Diffusion v2.1 and conditioned on image and warped-image features that are not functions solely of the fitted disparity, so the predicted normals are not equal to the disparity-derived prior by construction. The headline zero-shot disparity and normal results on KITTI, NYUv2, iBims-1, ScanNet, and other benchmarks are evaluated on external datasets not used in training, giving the central performance claims independent content. The main evidentiary weakness is that the ablations supporting the causal claim that the diffusion branch improves disparity (Tables 3-5) are evaluated on IRS and 3D Ken Burns, both listed in Sec. 4.2 as training data, while being labeled 'zero-shot.' This is an in-distribution evaluation and should not be read as evidence of cross-domain generalization; it weakens the strength of the causal story, but it is not a circularity because the reported numbers are not forced to equal the training targets by the paper's equations. The paper's many self-citations (LightStereo, StereoCarla, OpenStereo, Stereo Anything, etc.) are used as baselines, training sources, or related work and are not load-bearing assumptions of the proposed method. No uniqueness theorem, self-imported ansatz, or renaming of a known result carries the argument. The score of 2 reflects minor non-load-bearing self-citation and the evaluation caveat, not a derivation that reduces to its inputs.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

No fundamentally new physical or categorical entities are postulated. D2N and warp-to-left-view are differentiable intermediate representations within the proposed architecture, not independent entities with falsifiable handles. The main borrowed assumptions are the pretrained FoundationStereo and Stable Diffusion backbones, plus the paper-specific belief that a log-disparity normal proxy is a reliable geometric prior.

free parameters (4)
  • dmax (max disparity normalization) = not reported
    Eq 9 normalizes D_init by dmax before computing Z = -log(D_bar); the value changes the pseudo-depth scale and normal prior but is not reported or ablated.
  • lambda_init (initial-disparity loss weight) = not reported
    Eq 8 balances initial vs final disparity supervision; affects stereo branch training dynamics but is not specified.
  • lambda_norm (normal loss weight) = not reported
    Eq 14 balances disparity and normal objectives; controls how strongly diffusion priors are pushed into the stereo branch, value not given.
  • Z smoothing filter kernel size = not reported
    Sec 3.3 says Z is smoothed with a local average filter before normal computation; kernel size is unspecified and could affect the coarse normal prior.
axioms (5)
  • domain assumption Rectified stereo pair with normalized image coordinates is available; disparity maps to a consistent point map through the D2N construction.
    Eqs. 9-10 assume normalized pixel coordinates (x,y) in [-1,1] and a log-disparity proxy for depth without using camera intrinsics.
  • domain assumption Large-scale diffusion models contain geometric priors that transfer to normal estimation and disparity refinement.
    Core motivation in Sec 1 and Sec 3.1; supported only by ablations, not by an independent analysis of where the priors come from.
  • ad hoc to paper The D2N log-depth normal is a sufficiently unbiased coarse prior for the diffusion branch to correct.
    Sec 3.3, Eqs. 9-11: this proxy is introduced for this paper and its geometric validity is not justified.
  • domain assumption Two-stage training (separate then joint) yields consistent coupling without instability.
    Sec 3.4 'Consistency Training Strategy' states this heuristic; no analysis of stage-wise stability or sensitivity is provided.
  • standard math Standard DDPM/DDIM/LDM machinery is correct and applicable in latent space for normal refinement.
    Preliminaries in Sec 3.2 and the loss in Eq 13 rely on established diffusion model theory.

pith-pipeline@v1.3.0-alltime-deepseek · 16042 in / 12705 out tokens · 117128 ms · 2026-07-31T23:13:26.934339+00:00 · methodology

0 comments
read the original abstract

Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produce reliable predictions in challenging regions, mainly due to the lack of strong geometric priors. In this paper, we propose $\textbf{GeoStereo}$, a unified stereo geometry estimation framework that leverages powerful diffusion priors to jointly predict disparity and surface normals. Specifically, GeoStereo couples a feed-forward stereo matching pipeline with a diffusion-based normal estimation branch. To enable effective interaction between the two tasks, we introduce a disparity to normal initialization strategy and construct a warp to left-view condition for the diffusion process. This coupled design allows the diffusion branch to provide strong structural priors that enhance disparity estimation in ill-posed regions, while the feed-forward branch offers reliable geometric guidance for accurate normal prediction. Extensive experiments show that GeoStereo performs reliably in challenging scenarios, including low-light environments, highly reflective surfaces, and transparent objects. Under zero-shot settings, it achieves Rank-1 disparity estimation on multiple benchmarks, including KITTI and NYUv2, and delivers the best normal estimation accuracy on many real indoor benchmarks, such as iBims-1 and ScanNet. Project page: https://qz-wei.github.io/GeoStereo.github.io/

Figures

Figures reproduced from arXiv: 2607.24024 by Hao Zhao, Hong Li, Qizhe Wei, Runyi Yang, Shaocong Xu, Xianda Guo.

Figure 1
Figure 1. Figure 1: Zero-shot predictions on diverse indoor scenes, including robotics setups, office, reflective surfaces, and large open [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: GeoStereo jointly predicts disparity and surface [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the proposed GeoStereo framework. A stereo pair is first processed by a weight-shared stereo encoder, [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Left: Different architectures for integrating disparity and normal estimation. Right: Effects of right-view guidance [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison of zero-shot inference on diverse scenes. All methods are trained on a mixture of public [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

87 extracted references · 19 linked inside Pith

  1. [1]

    Gwangbin Bae, Ignas Budvytis, and Roberto Cipolla. 2021. Estimating and exploiting the aleatoric uncertainty in surface normal estimation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 13137–13146

  2. [2]

    Gwangbin Bae and Andrew J Davison. 2024. Rethinking inductive biases for surface normal estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9535–9545

  3. [3]

    Luca Bartolomei, Fabio Tosi, Matteo Poggi, and Stefano Mattoccia. 2025. Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching Even Where Either Stereo or Mono Fail. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  4. [4]

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127(2023)

  5. [5]

    D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black. 2012. A naturalistic open source movie for optical flow evaluation. InEuropean Conf. on Computer Vision (ECCV) (Part IV, LNCS 7577), A. Fitzgibbon et al. (Eds.) (Ed.). Springer-Verlag, 611–625

  6. [6]

    Yohann Cabon, Naila Murray, and Martin Humenberger. 2020. Virtual kitti 2. arXiv preprint arXiv:2001.10773(2020)

  7. [7]

    Jia-Ren Chang and Yong-Sheng Chen. 2018. Pyramid stereo matching network. InProceedings of the IEEE conference on computer vision and pattern recognition. 5410–5418

  8. [8]

    Xiaoxue Chen, Ziyi Xiong, Yuantao Chen, Gen Li, Nan Wang, Hongcheng Luo, Long Chen, Haiyang Sun, Bing Wang, Guang Chen, et al . 2025. DGGT: Feed- forward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images. arXiv preprint arXiv:2512.03004(2025)

  9. [9]

    Junda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang, Zhaoxing Zhang, Yong Deng, Jinliang Zang, Yurui Chen, Zhipeng Cai, and Xin Yang. 2025. MonSter: Marry Monodepth to Stereo Unleashes Power. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6273–6282

  10. [10]

    Junda Cheng, Wei Yin, Kaixuan Wang, Xiaozhi Chen, Shijie Wang, and Xin Yang

  11. [11]

    Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, and Zongyuan Ge. 2020. Hierarchical neural architecture search for deep stereo matching.Advances in neural information processing systems33 (2020), 22158–22169

  12. [12]

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. 2017. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE conference on computer vision and pattern recognition. 5828–5839

  13. [13]

    Qiyu Dai, Jiyao Zhang, Qiwei Li, Tianhao Wu, Hao Dong, Ziyuan Liu, Ping Tan, and He Wang. 2022. Domain Randomization-Enhanced Depth Simulation and Restoration for Perceiving and Grasping Specular and Transparent Objects. In European Conference on Computer Vision (ECCV)

  14. [14]

    Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. 2021. Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3D scans. InProceedings of the IEEE/CVF International Conference on Computer Vision

  15. [15]

    David Eigen and Rob Fergus. 2015. Predicting depth, surface normals and seman- tic labels with a common multi-scale convolutional architecture. InProceedings of the IEEE International Conference on Computer Vision

  16. [16]

    Xiao Fu, Wei Yin, Mu Hu, Kaixuan Wang, Yuexin Ma, Ping Tan, Shaojie Shen, Dahua Lin, and Xiaoxiao Long. 2024. GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image. InECCV

  17. [17]

    Zhiheng Fu, Siyu Hong, Mengyi Liu, Hamid Laga, Mohammed Bennamoun, Farid Boussaid, and Yulan Guo. 2023. Multi-stage information diffusion for joint depth and surface normal estimation.Pattern Recognition141 (2023), 109660. doi:10.1016/j.patcog.2023.109660

  18. [18]

    Yuan Gao, Chen Chen, and Jiatao Gu. 2026. One layer is enough: Adapting pretrained visual encoders for image generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4688–4697

  19. [19]

    Gonzalo Martin Garcia, Karim Abou Zeid, Christian Schmidt, Daan De Geus, Alexander Hermans, and Bastian Leibe. 2025. Fine-tuning image-conditional diffusion models is easier than you think. InProceedings of the Winter Conference on Applications of Computer Vision. 753–762

  20. [20]

    Andreas Geiger, Philip Lenz, and Raquel Urtasun. 2012. Are we ready for au- tonomous driving? The KITTI vision benchmark suite. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3354–3361

  21. [21]

    Tongfan Guan, Jiaxin Guo, Chen Wang, and Yun-Hui Liu. 2025. Bridgedepth: Bridging monocular and stereo reasoning with latent alignment. InProceedings of the IEEE/CVF International Conference on Computer Vision. 27681–27691

  22. [22]

    Tongfan Guan, Chen Wang, and Yun-Hui Liu. 2024. Neural markov random field for stereo matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5459–5469

  23. [23]

    Xianda Guo, Juntao Lu, Chenming Zhang, Yiqi Wang, Yiqun Duan, Tian Yang, Zheng Zhu, and Long Chen. 2023. Openstereo: A comprehensive benchmark for stereo matching and strong baseline.arXiv preprint arXiv:2312.00343(2023)

  24. [24]

    Xiaoyang Guo, Kai Yang, Wukui Yang, Xiaogang Wang, and Hongsheng Li. 2019. Group-wise Correlation Stereo Network. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3273–3282

  25. [25]

    Xianda Guo, Chenming Zhang, Ruilin Wang, Youmin Zhang, Wenzhao Zheng, Matteo Poggi, Hao Zhao, Qin Zou, and Long Chen. 2025. StereoCarla: A High- Fidelity Driving Dataset for Generalizable Stereo.arXiv preprint arXiv:2509.12683 (2025)

  26. [26]

    Xianda Guo, Chenming Zhang, Youmin Zhang, Ruilin Wang, Dujun Nie, Wenzhao Zheng, Matteo Poggi, Hao Zhao, Mang Ye, Qin Zou, et al. 2024. Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data.arXiv preprint arXiv:2411.14053(2024)

  27. [27]

    Xianda Guo, Chenming Zhang, Youmin Zhang, Wenzhao Zheng, Dujun Nie, Matteo Poggi, and Long Chen. 2025. Lightstereo: Channel boost is all you need for efficient 2d cost aggregation. InICRA

  28. [28]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models.Advances in neural information processing systems33 (2020), 6840–6851

  29. [29]

    Barry Shichen Hu, Siyun Liang, Johannes Paetzold, Huy H Nguyen, Isao Echizen, and Jiapeng Tang. 2024. Surface Normal Estimation with Transformers.arXiv preprint arXiv:2401.05745(2024)

  30. [30]

    Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. 2024. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation.IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)

  31. [31]

    Jingwei Huang, Yichao Zhou, Thomas Funkhouser, and Leonidas J Guibas. 2019. Framenet: Learning local canonical frames of 3d surfaces from a single rgb image. InProceedings of the IEEE/CVF International Conference on Computer Vision. 8638– 8647

  32. [32]

    Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Rui Xu, Minglang Tan, Wenjie Jiang, and Rui Huang. 2025. Defom-stereo: Depth foundation model based stereo matching. InProceedings of the Computer Vision and Pattern Recognition Conference. 21857– 21867

  33. [33]

    Bingxin Ke, Kevin Qu, Tianfu Wang, Nando Metzger, Shengyu Huang, Bo Li, Anton Obukhov, and Konrad Schindler. 2025. Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis. arXiv:2505.09358 [cs.CV]

  34. [34]

    Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta, Peter Henry, Ryan Kennedy, Abraham Bachrach, and Adam Bry. 2017. End-to-End Learning of Geometry and Context for Deep Stereo Regression. arXiv:1703.04309 [cs.CV] https://arxiv.org/ abs/1703.04309

  35. [35]

    Tobias Koch, Lukas Liebel, Friedrich Fraundorfer, and Marco Korner. 2018. Eval- uation of cnn-based single-image depth estimation methods. InProceedings of the European Conference on Computer Vision (ECCV) Workshops. 0–0

  36. [36]

    Peter Kocsis, Lukas Höllein, and Matthias Nießner. 2025. Intrinsix: High-quality pbr generation using image priors.arXiv preprint arXiv:2504.01008(2025)

  37. [37]

    Hong Li, Houyuan Chen, Chongjie Ye, Zhaoxi Chen, Bohan Li, Shaocong Xu, Xianda Guo, Xuhui Liu, Yikai Wang, Baochang Zhang, et al . 2025. Light of normals: Unified feature representation for universal photometric stereo.arXiv preprint arXiv:2506.18882(2025)

  38. [38]

    Jiankun Li, Peisen Wang, Pengfei Xiong, Tao Cai, Ziwei Yan, Lei Yang, Jiangyu Liu, Haoqiang Fan, and Shuaicheng Liu. 2022. Practical stereo matching via cascaded recurrent network with adaptive correlation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16263–16272

  39. [39]

    Yixuan Li, Lihan Jiang, Linning Xu, Yuanbo Xiangli, Zhenzhi Wang, Dahua Lin, and Bo Dai. 2023. MatrixCity: A Large-scale City Dataset for City-scale Neural Rendering and Beyond. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3205–3215

  40. [40]

    Waslander

    Ziwei Liao, Jialiang Zhu, Chunyu Wang, Han Hu, and Steven L. Waslander

  41. [41]

    Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie, and Zhizhong Su. 2025. BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  42. [42]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

    Multiple View Geometry Transformers for 3D Human Pose Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  43. [43]

    Lahav Lipson, Zachary Teed, and Jia Deng. 2021. RAFT-Stereo: Multilevel Re- current Field Transforms for Stereo Matching.arXiv preprint arXiv:2109.07547 (2021)

  44. [44]

    Xinqi Lin, Fanghua Yu, Jinfan Hu, Zhiyuan You, Wu Shi, Jimmy S Ren, Jinjin Gu, and Chao Dong. 2025. Harnessing diffusion-yielded score priors for image restoration.ACM Transactions on Graphics (TOG)44, 6 (2025), 1–21

  45. [45]

    Zihua Liu, Songyan Zhang, Zhicheng Wang, and Masatoshi Okutomi. 2022. Dig- ging Into Normal Incorporated Stereo Matching. InProceedings of the 30th ACM International Conference on Multimedia. 6050–6060

  46. [46]

    Tianqi Liu, Zhaoxi Chen, Zihao Huang, Shaocong Xu, Saining Zhang, Chongjie Ye, Bohan Li, Zhiguo Cao, Wei Li, Hao Zhao, et al. 2025. Light-X: Generative 4D Video Rendering with Camera and Illumination Control.arXiv preprint arXiv:2512.05115(2025). MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Wei et al

  47. [47]

    Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. 2016. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. InProceedings of the IEEE conference on computer vision and pattern recognition. 4040–4048

  48. [48]

    Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101(2017)

  49. [49]

    Simon Niklaus, Long Mai, Jimei Yang, and Feng Liu. 2019. 3D Ken Burns Effect from a Single Image.ACM Transactions on Graphics38, 6 (2019), 184:1–184:15

  50. [50]

    Junhong Min, Youngpil Jeon, Jimin Kim, and Minyong Choi. 2025. S2M2: Scalable Stereo Matching Model for Reliable Depth Estimation. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

  51. [51]

    Mike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar, Miguel Angel Bautista, Nathan Paczan, Russ Webb, and Joshua M Susskind. 2021. Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In Proceedings of the IEEE/CVF international conference on computer vision. 10912– 10922

  52. [52]

    René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. 2021. Vision transformers for dense prediction. InProceedings of the IEEE/CVF international conference on computer vision. 12179–12188

  53. [53]

    Shuwei Shao, Zhongcai Pei, Weihai Chen, Xingming Wu, and Zhengguo Li. 2023. NDDepth: Normal-distance assisted monocular depth estimation. InProceedings of the IEEE/CVF International Conference on Computer Vision. 7931–7940

  54. [54]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752 [cs.CV] https://arxiv.org/abs/2112.10752

  55. [55]

    Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from rgbd images. InEuropean conference on computer vision. Springer, 746–760

  56. [56]

    Zhelun Shen, Yuchao Dai, and Zhibo Rao. 2021. CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  57. [57]

    Jonathan Tremblay, Thang To, and Stan Birchfield. 2018. Falling things: A syn- thetic dataset for 3d object detection and pose estimation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2038– 2041

  58. [58]

    Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502(2020)

  59. [59]

    Jiyuan Wang, Chunyu Lin, Lei Sun, Rongying Liu, Mingxing Li, Lang Nie, Kang Liao, Xiangxiang Chu, and Yao Zhao. 2025. FE2E: From Editor to Dense Geometry Estimator.arXiv preprint arXiv:2509.04338(2025)

  60. [60]

    Dai, Andrea F

    Igor Vasiljevic, Nick Kolkin, Shanyi Zhang, Ruotian Luo, Haochen Wang, Falcon Z. Dai, Andrea F. Daniele, Mohammadreza Mostajabi, Steven Basart, Matthew R. Walter, and Gregory Shakhnarovich. 2019. DIODE: A Dense Indoor and Outdoor DEpth Dataset.CoRRabs/1908.00463 (2019). http://arxiv.org/abs/1908.00463

  61. [61]

    Ruicheng Wang, Sicheng Xu, Yue Dong, Yu Deng, Jianfeng Xiang, Zelong Lv, Guangzhong Sun, Xin Tong, and Jiaolong Yang. 2025. MoGe-2: Accurate Monoc- ular Geometry with Metric Scale and Sharp Details. arXiv:2507.02546 [cs.CV] https://arxiv.org/abs/2507.02546

  62. [62]

    Qiang Wang, Shizhen Zheng, Qingsong Yan, Fei Deng, Kaiyong Zhao, and Xiaowen Chu. 2021. IRS: A Large Naturalistic Indoor Robotics Stereo Dataset to Train Deep Models for Disparity and Surface Normal Estimation. arXiv:1912.09678 [cs.CV] https://arxiv.org/abs/1912.09678

  63. [63]

    Xianqi Wang, Hao Yang, Hangtian Wang, Junda Cheng, Gangwei Xu, Min Lin, and Xin Yang. 2026. PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts.arXiv preprint arXiv:2603.01650(2026)

  64. [64]

    Xianqi Wang, Gangwei Xu, Hao Jia, and Xin Yang. 2024. Selective-stereo: Adap- tive frequency information selection for stereo matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19701–19710

  65. [65]

    Yun Wang, Longguang Wang, Chenghao Zhang, Yongjian Zhang, Zhanjie Zhang, Ao Ma, Chenyou Fan, Tin Lun Lam, and Junjie Hu. 2025. Learning robust stereo matching in the wild with selective mixture-of-experts. (2025), 21276–21287

  66. [66]

    Xianqi Wang, Hao Yang, Gangwei Xu, Junda Cheng, Min Lin, Yong Deng, Jinliang Zang, Yurui Chen, and Xin Yang. 2025. ZeroStereo: Zero-shot Stereo Matching from Single Images. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

  67. [67]

    Bowen Wen, Matthew Trepte, Joseph Aribido, Jan Kautz, Orazio Gallo, and Stan Birchfield. 2025. Foundationstereo: Zero-shot stereo matching. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5249–5260

  68. [68]

    Qizhe Wei, Yingping Liang, Shaodi You, and Ying Fu. 2026. AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation.arXiv preprint arXiv:2607.04303(2026)

  69. [69]

    Guangkai Xu, Yixun Ge, Mingyu Liu, Chengxiang Fan, Kangkang Xie, Ziyue Zhao, Haoran Chen, and Chunhua Shen. 2025. What Matters When Repurpos- ing Diffusion Models for General Dense Perception Tasks?. InThe Thirteenth International Conference on Learning Representations

  70. [70]

    Gangwei Xu, Junda Cheng, Peng Guo, and Xin Yang. 2022. Attention Concate- nation Volume for Accurate and Efficient Stereo Matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  71. [71]

    Gangwei Xu, Xianqi Wang, Zhaoxing Zhang, Junda Cheng, Chunyuan Liao, and Xin Yang. 2024. IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo Matching.arXiv preprint arXiv:2409.00638(2024)

  72. [72]

    Gangwei Xu, Xianqi Wang, Xiaohuan Ding, and Xin Yang. 2023. Iterative Ge- ometry Encoding Volume for Stereo Matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  73. [73]

    Shaocong Xu, Songlin Wei, Qizhe Wei, Zheng Geng, Hong Li, Licheng Shen, Qianpu Sun, Shu Han, Bin Ma, Bohan Li, et al . 2025. Diffusion Knows Trans- parency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation.arXiv preprint arXiv:2512.23705(2025)

  74. [74]

    Gangwei Xu, Yun Wang, Junda Cheng, Jinhui Tang, and Xin Yang. 2023. Ac- curate and efficient stereo matching via attention concatenation volume.IEEE Transactions on Pattern Analysis and Machine Intelligence46, 4 (2023), 2461–2474

  75. [75]

    Chongjie Ye, Lingteng Qiu, Xiaodong Gu, Qi Zuo, Yushuang Wu, Zilong Dong, Liefeng Bo, Yuliang Xiu, and Xiaoguang Han. 2024. StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal.ACM Transactions on Graphics (TOG)(2024)

  76. [76]

    Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024. Depth Anything V2.arXiv:2406.09414(2024)

  77. [77]

    Minghao Yin, Shangzhe Wu, and Kai Han. 2024. IBD-SLAM: Learning Image- Based Depth Fusion for Generalizable SLAM. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  78. [78]

    Chongjie Ye, Yushuang Wu, Ziteng Lu, Jiahao Chang, Xiaoyang Guo, Jiaqing Zhou, Hao Zhao, and Xiaoguang Han. 2025. Hi3dgen: High-fidelity 3d geometry generation from images via normal bridging. InProceedings of the IEEE/CVF International Conference on Computer Vision. 25050–25061

  79. [79]

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. arXiv:2302.05543 [cs.CV] https: //arxiv.org/abs/2302.05543

  80. [80]

    Feihu Zhang, Victor Prisacariu, Ruigang Yang, and Philip H. S. Torr

Showing first 80 references.