REVIEW 3 major objections 5 minor 3 cited by
MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching
T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read MatchAttention makes the query–key match offset a learnable attention center, yielding linear-complexity cross-view matching that leads Middlebury.
desk verdict A genuinely new linear-complexity attention primitive for stereo/flow with strong empirical results; the coverage concern is real but softened by direct residual supervision and the ablation's LRI fix. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
MatchAttention is a sliding-window attention whose sampling center is p_query + R_pos, with R_pos a per-token continuous 2D offset learned end-to-end. BilinearSoftmax distributes each query's attention over four integer sub-windows with bilinear weights, keeping the whole sampling path differentiable; negative L1-norm plus softmax acts as a normalized Laplace kernel suited to one-to-one matching. R_pos is updated by residual connections and fed as extra feature channels, so the network iteratively refines disparity or flow while aggregating features.
What would settle it
Construct or collect stereo/flow pairs with large, abrupt disparities and corrupt the coarse initialization—for instance by downsampling beyond 1/32 or adding repetitive texture—then check whether MatchAttention can still converge to the true match. A specific test: take a trained model, shift the initial R_pos by ±(window_radius+1) pixels at the coarsest scale, and measure whether fine-scale refinements recover; if errors jump to chance level, the coarse-to-fine coverage assumption is the binding constraint.
Extended reading notes
Core claim
The central claim is that the relative position between a query and its matched key can be treated as a learnable component of attention sampling, not as a positional embedding. MatchAttention computes, for each query, a contiguous window centered at the query position plus a learned offset; BilinearSoftmax makes sampling differentiable and sub-pixel accurate; residual connections embed the offset in feature channels so it refines layer by layer. Because the window is small and constant, complexity is linear. The paper then instantiates this in a stereo/flow decoder with negative L1-norm similarity, gated cross-attention, and a consistency-constrained loss to handle occlusion. If the claims
Load-bearing premise
The whole chain rests on the initial correlation at 1/32 scale being accurate enough that the true matching key falls inside the small sampling window (3x3 or 5x5) at every finer level; if the coarse estimate is off by more than roughly half a window, the correct key never receives attention weight or gradient.
Editorial extensions
If this is right
- High-resolution stereo and flow inference becomes affordable: the model claims 4K UHD pairs in under 0.1 seconds and KITTI-resolution in 29 ms with a small GPU footprint.
- Long-range correspondences no longer require quadratic attention; arbitrary offsets reach beyond the local window at constant window cost.
- The predicted relative position is directly interpretable as disparity or flow, and self-attention offsets visualize where occluded regions are sampling, enabling explainable occlusion handling.
- Because window size has no learnable parameters, the same weights can train with a large window and infer with a smaller one, decoupling training cost from deployment speed.
- Zero-shot transfer from synthetic data to real benchmarks is reported as state of the art, suggesting that the explicit matching constraint improves generalization.
Reading between the lines
- If MatchAttention generalizes as the architecture suggests, it could replace global cross-attention in any cross-view matching stack—sparse feature matching, multi-view stereo, or feed-forward Gaussian splatting—where quadratic cost currently forces tiling.
- A stress test worth running: feed stereo pairs with disparities whose coarse 1/32-scale initialization is deliberately biased beyond the 3x3 or 5x5 window radius; the predicted failure mode is a sharp accuracy cliff rather than graceful degradation.
- The train-large/infer-small window property implies a compute-versus-accuracy knob for deployment that most attention designs lack; measuring that trade-off directly would be a useful extension.
- The normalized-Laplace similarity suggests a closer link to kernel-based matching than to dot-product attention; exploring whether the L1 kernel can be replaced by a learned metric may extend the method to non-rectified views.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MatchAttention, a sliding-window attention operator for cross-view matching in which the sampling center for each query is p_q + r, with r a learnable relative-position field. BilinearSoftmax makes the continuous sampling differentiable, and the relative positions are refined through residual connections in a hierarchical decoder. The authors instantiate this in MatchStereo/MatchFlow variants and report state-of-the-art Middlebury average error, strong zero-shot generalization, and high-resolution inference with linear attention complexity.
Significance. If the reported results hold, MatchAttention is a significant step toward high-resolution cross-view matching: it combines explicit correspondence prediction with attention-based aggregation while avoiding the quadratic cost of global cross-attention. The paper's strengths include a clear complexity analysis (Sec. 3.3), a broad evaluation across Middlebury, KITTI, ETH3D, Sintel and Spring, a released codebase with custom CUDA kernels, and systematic ablations (Sec. 5.3). The main open risk is the unquantified dependence of the long-range claim on the coarse-to-fine initialization, which should be addressed before the claims are fully convincing.
major comments (3)
- [Sec. 4.1, Eq. (5), Table 1] The central long-range claim relies on the initial R_pos from the 1/32 correlation (Eq. 9) being accurate enough that the true matching key lies inside the small sampling window at each finer scale. At 1/16, a 5x5 window tolerates an initial offset error of only about 32 full-resolution pixels before the true match leaves W_i; at 1/8 and 1/4 the budgets shrink further. If the true key is outside W_i, it receives zero attention weight and the attention module obtains no matching evidence; the dense L1 losses (Eqs. 13-16) provide direct training gradients, but at inference no such signal exists. The LRI ablation in Sec. 5.3 concedes that 'the initial correlation may not scale well at high resolution.' Please provide a perturbation study of the initial R_pos (or LRI vs no-LRI across resolutions) and discuss the failure mode, or qualify the 'arbitrary relative position' claim to the coverage
- [Sec. 3.2 vs Eq. (10)] The final paragraph of Sec. 3.2 claims that 'w in (5) does not introduce learnable parameters' and hence one can train with large w and infer with small w using the same weights. This is not true for the full model because Eq. (10) concatenates the flattened attention weights alpha_i to the input of the projection layer, making W'_p have c_v+(w+1)^2 input columns. Changing w changes the projection dimension. Please state that the size-independence holds only for the variant without the attention-weight injection, or modify the claim.
- [Abstract / Sec. 5.5] The arXiv metadata abstract names variants 'MatchAttentionXL' and 'MatchAttentionRT' and reports edge latencies (9.3 ms on RTX 4060 Ti, 79.1 ms on Jetson Orin NX at 1024x512) that do not appear in the main text, where the variants are MatchStereo-T/S/B. In addition, Sec. 5.5 claims state-of-the-art performance on KITTI 2015, but Table 9 shows DEFOM-Stereo (D1-all 1.41) and FoundationStereo (1.26) outperform MatchStereo-B (1.50). These claims need to be aligned with the reported numbers and with the actual model names.
minor comments (5)
- [Eqs. (6)-(7)] Eq. (6) uses exp(<q_i,k_j>) inside BilinearSoftmax, while Eq. (7) defines Softmax with the negative L1 norm exp(-gamma||q_i-k_j||_1). Please make the notation consistent by defining a single similarity function s(q,k).
- [Eq. (5)] The indexing W_i = {floor(p_i^k)+(u,v) | u,v=-w/2,...,w/2+1} is ambiguous for odd w (e.g., w=5) and mixes effective and expanded window sizes. Please define integer offsets explicitly (e.g., effective window offsets -(w-1)/2,...,(w-1)/2 and expanded window w+1).
- [Sec. 5.3] The cumulative analysis states that 'when any single component is removed from the full model, performance degradation occurs,' but the subtractive row Full-N improves on Full at half resolution (8.62 vs 9.23). The text should acknowledge this exception, as the component discussion itself does.
- [Table 8] In the provided manuscript text, Table 8 appears to contain no data rows; the reader only sees the caption and numbers cited in Sec. 5.5. Please verify that the table is rendered with the actual per-method AvgErr values.
- [Sec. 5.4] The statement that fixed-resolution inference 'improves accuracy through upsampling' for lower-resolution images is not uniformly supported: on FSD-Mix training, MatchStereo-B* is worse than MatchStereo-B on Middlebury F (6.56 vs 5.67) and on ETH3D (2.68 vs 1.36). Please qualify this benefit.
Circularity Check
No circularity: relative-position attention is a supervised, externally benchmarked architecture; the acknowledged coarse-to-fine initialization risk is a robustness limitation, not a self-referential derivation.
full rationale
The paper's central mechanism — using R_pos to center a small sampling window (Eqs. 4–5) and refining R_pos through residual connections and BilinearSoftmax gradients (Eq. 8) — is an empirically trained and externally evaluated architecture, not a derivation whose output equals its input. R_pos is initialized from a 1/32-scale correlation (Eq. 9), but it is subsequently supervised against ground-truth disparity/flow in L_init, L_self, and L_cross (Eqs. 13–16), and zero-shot results are reported separately from finetuned benchmark results. The only statement in the paper that touches on a potential failure mode of the mechanism is the LRI ablation: 'The initial correlation may not scale well at high resolution and LRI can provide a more robust initial relative position.' This flags a coverage/robustness concern — if the initial R_pos is too far from the true match, the matching key lies outside the sampled window and receives no corrective gradient — but it is a limitation of the method, not circular reasoning. MatchAttention is compared against external methods on public benchmarks (Middlebury, KITTI, ETH3D, Spring), and references to the authors' own prior work (e.g., [42]) are contextual and not load-bearing. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from self-citation, and no known empirical pattern is merely relabeled. Therefore no circular step can be exhibited.
Assumptions & free parameters
free parameters (5)
- beta (R_pos input scale) =
learned; initialized (0.1, 0.1)
- A (non-occlusion mask threshold) =
1
- epsilon (consistency loss weight) =
0.01
- gamma (L1 similarity scaling) =
1/sqrt(c_k)
- sampling window sizes w =
(5,5,3,3) across decoder scales
assumptions (4)
- domain assumption Rectified epipolar geometry: stereo matching reduces to a 1D search along the x-axis
- domain assumption Coarse-to-fine initialization coverage: the 1/32-scale initial correlation and convex upsampling place the true match within the small sampling window at each finer scale
- domain assumption Bilinear interpolation and fused CUDA sampling are faithful differentiable implementations of the described BilinearSoftmax
- domain assumption Benchmark ground truth (Middlebury, KITTI, ETH3D, Sintel, Spring) is accurate and used consistently
Cite this review
Pith. "Pith review of MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching." pith.science (2026). https://pith.science/paper/L47QPQN4
@misc{pith2026251014260,
author = {Pith},
title = {Pith review of: MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching},
year = {2026},
howpublished = {\url{https://pith.science/paper/L47QPQN4}},
note = {Machine review of arXiv:2510.14260}
}
read the original abstract
Standard attention mechanisms are not well suited to stereo matching. Global attention scales quadratically and provides no explicit matching constraint, while local attention is efficient but loses long-range correspondences. We propose MatchAttention, an attention mechanism that embeds an explicit matching constraint into attention by treating the relative position between a query and its matched key as a learnable component of attention sampling. Centering a small contiguous sampling window on this learnable relative position enforces the matching constraint and supports long-range correspondence at strictly linear attention complexity. A differentiable contiguous attention sampling (CAS) operator enables sub-pixel accuracy, and cascaded MatchAttention blocks iteratively refine the relative positions through residual connections. We instantiate MatchAttention as a hierarchical coarse-to-fine stereo network with two variants. MatchAttentionXL targets accuracy and MatchAttentionRT targets real-time edge inference. MatchAttentionXL achieves state-of-the-art accuracy on Middlebury V3 and top results across KITTI 2012/2015 and ETH3D. MatchAttentionRT runs at 9.3 ms on RTX 4060 Ti and 79.1 ms on Jetson Orin NX 16 GB at 1024 x 512, making it the first stereo model to deliver real-time edge inference without sacrificing zero-shot generalization. The code is available at https://github.com/TingmanYan/MatchAttention.
Figures
Forward citations
Cited by 3 Pith papers
-
WAVE-Stereo: Warp-Aligned Volume Encoding for Stereo Matching
WAVE-Stereo unifies correlation-volume search and feature-warping residual alignment in an iterative stereo matcher, achieving real-time zero-shot generalization on five benchmarks.
-
WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching
Warping alone, with a classification head before iterative regression, matches or beats cost-volume stereo methods on ETH3D, KITTI and Middlebury at higher speed.
-
URS-Stereo: Uncertainty-Guided Residual Search for Real-Time Stereo Matching
Uncertainty-modulated residual offsets relocate local cost-volume centers in coarse-to-fine stereo matching, improving zero-shot disparity accuracy while keeping real-time speed.
Reference graph
Works this paper leans on
-
[1]
Dust3r: Geometric 3d vision made easy,
S. Wang, V . Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 20 697–20 709
2024
-
[2]
Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,
Y. Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.-J. Cham, and J. Cai, “Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,”Eur. Conf. Comput. Vis., pp. 370–386, 2024
2024
-
[3]
Splatt3r: Zero- shot gaussian splatting from uncalibrated image pairs,
B. Smart, C. Zheng, I. Laina, and V . A. Prisacariu, “Splatt3r: Zero- shot gaussian splatting from uncalibrated image pairs,”arXiv preprint arXiv:2408.13912, 2024
arXiv 2024
-
[4]
No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images,
B. Ye, S. Liu, H. Xu, L. Xueting, M. Pollefeys, M.-H. Yang, and P . Songyou, “No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images,”arXiv preprint arXiv:2410.24207, 2024
arXiv 2024
-
[5]
Unifying flow, stereo and depth estimation,
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,”IEEE T rans. Pattern Anal. Mach. Intell., vol. 45, no. 11, pp. 13 941–13 958, 2023
2023
-
[6]
Grounding image matching in 3d with mast3r,
V . Leroy, Y. Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” inEur. Conf. Comput. Vis., 2024, pp. 71–91
2024
-
[7]
LoFTR: Detector- free local feature matching with transformers,
J. Sun, Z. Shen, Y. Wang, H. Bao, and X. Zhou, “LoFTR: Detector- free local feature matching with transformers,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2021, pp. 8922–8931
2021
-
[8]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2016
2016
Show all 88 references
-
[9]
Show, attend and tell: neural image caption generation with visual attention,
K. Xu, J. L. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: neural image caption generation with visual attention,” inInt. Conf. Mach. Learn., 2015, p. 2048–2057
2015
-
[10]
Effective approaches to attention-based neural machine translation,
T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” inEmpirical Methods in Natural Language Processing, EMNLP, 2015, pp. 1412–1421
2015
-
[11]
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection.” inInt. Conf. Learn. Represent., 2021
2021
-
[12]
High-resolution stereo datasets with subpixel-accurate ground truth,
D. Scharstein, H. Hirschm ¨uller, Y. Kitajima, G. Krathwohl, N. Ne ˇsi´c, X. Wang, and P . Westling, “High-resolution stereo datasets with subpixel-accurate ground truth,” inPattern Recog. German Conf., 2014, pp. 31–42
2014
-
[13]
A multi-view stereo benchmark with high-resolution images and multi-camera videos,
T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 3260–3269
2017
-
[14]
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,
N. Mayer, E. Ilg, P . Hausser, P . Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 4040–4048
2016
-
[15]
Raft-stereo: Multilevel recurrent field transforms for stereo matching,
L. Lipson, Z. Teed, and J. Deng, “Raft-stereo: Multilevel recurrent field transforms for stereo matching,” inInt. Conf. 3D Vision, 2021, pp. 218–227
2021
-
[16]
Foundationstereo: Zero-shot stereo matching,
B. Wen, M. Trepte, J. Aribido, J. Kautz, O. Gallo, and S. Birchfield, “Foundationstereo: Zero-shot stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025, pp. 5249–5260
2025
-
[17]
Are we ready for autonomous driving? the kitti vision benchmark suite,
A. Geiger, P . Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” inIEEE Conf. Comput. Vis. Pattern Recog., 2012
2012
-
[18]
Joint 3d estimation of vehicles and scene flow,
M. Menze, C. Heipke, and A. Geiger, “Joint 3d estimation of vehicles and scene flow,”ISPRS Ann. Photogrammetry, Remote Sens. Spatial Inf. Sciences, vol. 2, p. 427, 2015
2015
-
[19]
Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo,
L. Mehl, J. Schmalfuss, A. Jahedi, Y. Nalivayko, and A. Bruhn, “Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 4981–4991
2023
-
[20]
Neural machine translation by jointly learning to align and translate,
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” inInt. Conf. Learn. Represent., 2015
2015
-
[21]
Bridging the divide: Reconsidering softmax and linear attention,
D. Han, Y. Pu, Z. Xia, Y. Han, X. Pan, X. Li, J. Lu, S. Song, and G. Huang, “Bridging the divide: Reconsidering softmax and linear attention,” inAdv. Neural Inform. Process. Syst., vol. 37, 2024, pp. 79 221–79 245
2024
-
[22]
Longformer: The long- document transformer,
I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long- document transformer,”arXiv:2004.05150, 2020
2004 arXiv
-
[23]
Big bird: Transformers for longer sequences,
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P . Pham, A. Ravula, Q. Wang, L. Yanget al., “Big bird: Transformers for longer sequences,”Adv. Neural Inform. Process. Syst., vol. 33, 2020
2020
-
[24]
Sparser is faster and less is more: Efficient sparse attention for long-range transformers,
C. Lou, Z. Jia, Z. Zheng, and K. Tu, “Sparser is faster and less is more: Efficient sparse attention for long-range transformers,” arXiv:2406.16747, 2024. 15
2024 arXiv
-
[25]
Neighborhood attention transformer,
A. Hassani, S. Walton, J. Li, S. Li, and H. Shi, “Neighborhood attention transformer,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 6185–6194
2023
-
[26]
Clear: Conv-like lineariza- tion revs pre-trained diffusion transformers up,
S. Liu, Z. Tan, and X. Wang, “Clear: Conv-like lineariza- tion revs pre-trained diffusion transformers up,”arXiv preprint arXiv:2412.16112, 2024
2024 arXiv
-
[27]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inInt. Conf. Comput. Vis., October 2021, pp. 10 012– 10 022
2021
-
[28]
Bevformer: Learning bird’s-eye-view representation from multi- camera images via spatiotemporal transformers,
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi- camera images via spatiotemporal transformers,” inEur. Conf. Comput. Vis., 2022, pp. 1–18
2022
-
[29]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” inAdv. Neural Inform. Process. Syst., vol. 30, 2017
2017
-
[30]
Swin transformer v2: Scaling up capacity and resolution,
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo, “Swin transformer v2: Scaling up capacity and resolution,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 12 009–12 019
2022
-
[31]
Transnext: Robust foveal visual perception for vision transformers,
D. Shi, “Transnext: Robust foveal visual perception for vision transformers,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 17 773–17 783
2024
-
[32]
Roformer: Enhanced transformer with rotary position embedding,
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,”Neuro- comput., vol. 568, no. C, February 2024
2024
-
[33]
The llama 3 herd of models,
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al- Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, and et al., “The llama 3 herd of models,”arXiv:2407.21783, 2024
2024 arXiv
-
[34]
Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow,
P . Weinzaepfel, T. Lucas, V . Leroy, Y. Cabon, V . Arora, R. Br´egier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud, “Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow,” inInt. Conf. Comput. Vis., 2023, pp. 17 969–17 980
2023
-
[35]
Cameras as relative positional encoding,
R. Li, B. Yi, J. Liu, H. Gao, Y. Ma, and A. Kanazawa, “Cameras as relative positional encoding,”arXiv:2507.10496, 2025
2025
-
[36]
Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,
Z. Li, X. Liu, N. Drenkow, A. Ding, F. X. Creighton, R. H. Taylor, and M. Unberath, “Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,” inInt. Conf. Comput. Vis., 2021, pp. 6197–6206
2021
-
[37]
End-to-end learning of geometry and context for deep stereo regression,
A. Kendall, H. Martirosyan, S. Dasgupta, P . Henry, R. Kennedy, A. Bachrach, and A. Bry, “End-to-end learning of geometry and context for deep stereo regression,” inInt. Conf. Comput. Vis., 2017, pp. 66–75
2017
-
[38]
Pyramid stereo matching network,
J.-R. Chang and Y.-S. Chen, “Pyramid stereo matching network,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 5410–5418
2018
-
[39]
Correlate-and-excite: Real-time stereo matching via guided cost volume excitation,
A. Bangunharcana, J. W. Cho, S. Lee, I. S. Kweon, K.-S. Kim, and S. Kim, “Correlate-and-excite: Real-time stereo matching via guided cost volume excitation,” inInt. Conf. Intell. Robots Syst., 2021
2021
-
[40]
Raft: Recurrent all-pairs field transforms for optical flow,
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” inEur. Conf. Comput. Vis.Springer, 2020, pp. 402–419
2020
-
[41]
Continuous 3D Label Stereo Matching using Local Expansion Moves,
T. Taniai, Y. Matsushita, Y. Sato, and T. Naemura, “Continuous 3D Label Stereo Matching using Local Expansion Moves,”IEEE T rans. Pattern Anal. Mach. Intell., vol. 40, no. 11, pp. 2725–2739, 2018
2018
-
[42]
Hierarchical belief propagation on image segmentation pyramid,
T. Yan, X. Yang, G. Yang, and Q. Zhao, “Hierarchical belief propagation on image segmentation pyramid,”IEEE T rans. Image Process., vol. 32, pp. 4432–4442, 2023
2023
-
[43]
Iterative geometry encod- ing volume for stereo matching,
G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encod- ing volume for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 21 919–21 928
2023
-
[44]
Selective-stereo: Adaptive frequency information selection for stereo matching,
X. Wang, G. Xu, H. Jia, and X. Yang, “Selective-stereo: Adaptive frequency information selection for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 701–19 710
2024
-
[45]
All-in-one: Transferring vision foundation models into stereo matching,
J. Zhou, H. Zhang, J. Yuan, P . Ye, T. Chen, H. Jiang, M. Chen, and Y. Zhang, “All-in-one: Transferring vision foundation models into stereo matching,” inAssoc. Advancement Artif. Intell., February 2025
2025
-
[46]
DINOv2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V . Vo, M. Szafraniec, V . Khalidov, P . Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P .-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P . Labatut...
2024
-
[47]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P . Dollar, and R. Gir- shick, “Segment anything,” inInt. Conf. Comput. Vis., October 2023, pp. 4015–4026
2023
-
[48]
Depth anything v2,
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” inAdv. Neural Inform. Process. Syst., 2024
2024
-
[49]
Vision transformers for dense prediction,
R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inInt. Conf. Comput. Vis., October 2021, pp. 12 179–12 188
2021
-
[50]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” inInt. Conf. Learn. Represent., 2020
2020
-
[51]
Trans- formers as support vector machines,
D. A. Tarzanagh, Y. Li, C. Thrampoulidis, and S. Oymak, “Trans- formers as support vector machines,” inAdv. Neural Inform. Pro- cess. Syst., 2023
2023
-
[52]
Metaformer baselines for vision,
W. Yu, C. Si, P . Zhou, M. Luo, Y. Zhou, J. Feng, S. Yan, and X. Wang, “Metaformer baselines for vision,”IEEE T rans. Pattern Anal. Mach. Intell., vol. 46, no. 2, pp. 896–912, 2024
2024
-
[53]
Gaussian error linear units (GELUs),
D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELUs),”arXiv:1606.08415, 2023
2023 arXiv
-
[54]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P . Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” 2015, pp. 234–241
2015
-
[55]
Conditional positional encodings for vision transformers,
X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, “Conditional positional encodings for vision transformers,” inInt. Conf. Learn. Represent., 2023
2023
-
[56]
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,
S. Elfwing, E. Uchibe, and K. Doya, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,”Neural Netw., vol. 107, pp. 3–11, 2018
2018
-
[57]
Adaptive multi-modal cross-entropy loss for stereo matching,
P . Xu, Z. Xiang, C. Qiao, J. Fu, and X. Zhao, “Adaptive multi-modal cross-entropy loss for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024
2024
-
[58]
Decoupled weight decay regulariza- tion,
I. Loshchilov and F. Hutter, “Decoupled weight decay regulariza- tion,” inInt. Conf. Learn. Represent., 2019
2019
-
[59]
A naturalistic open source movie for optical flow evaluation,
D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” inEur. Conf. Comput. Vis.Springer, 2012, pp. 611–625
2012
-
[60]
Practical stereo matching via cascaded recurrent network with adaptive correlation,
J. Li, P . Wang, P . Xiong, T. Cai, Z. Yan, L. Yang, J. Liu, H. Fan, and S. Liu, “Practical stereo matching via cascaded recurrent network with adaptive correlation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16 263–16 272
2022
-
[61]
Falling things: A synthetic dataset for 3d object detection and pose estimation,
J. Tremblay, T. To, and S. Birchfield, “Falling things: A synthetic dataset for 3d object detection and pose estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 2038–2041
2018
-
[62]
In- stereo2k: a large real dataset for stereo matching in indoor scenes,
W. Bao, W. Wang, Y. Xu, Y. Guo, S. Hong, and X. Zhang, “In- stereo2k: a large real dataset for stereo matching in indoor scenes,” Sci. China Inf. Sci., vol. 63, pp. 1–11, 2020
2020
-
[63]
Hierarchical deep stereo matching on high-resolution images,
G. Yang, J. Manela, M. Happold, and D. Ramanan, “Hierarchical deep stereo matching on high-resolution images,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 5515–5524
2019
-
[64]
Virtual kitti 2,
Y. Cabon, N. Murray, and M. Humenberger, “Virtual kitti 2,” arXiv:2001.10773, 2020
2001 arXiv
-
[65]
Flownet: Learning optical flow with convolutional networks,
A. Dosovitskiy, P . Fischer, E. Ilg, P . H¨ausser, C. Hazırbas ¸, V . Golkov, P . v. d. Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” inInt. Conf. Comput. Vis., 2015, pp. 2758–2766
2015
-
[66]
The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driv- ing,
D. Kondermann, R. Nair, K. Honauer, K. Krispin, J. Andrulis, A. Brock, B. G ¨ussefeld, M. Rahimimoghaddam, S. Hofmann, C. Brenner, and B. J ¨ahne, “The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driv- ing,” inIEEE Conf. Comput. Vi...
2016
-
[67]
Defom-stereo: Depth foundation model based stereo matching,
H. Jiang, Z. Lou, L. Ding, R. Xu, M. Tan, W. Jiang, and R. Huang, “Defom-stereo: Depth foundation model based stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025, pp. 21 857– 21 867
2025
-
[68]
Depth pro: Sharp monocular metric depth in less than a second,
A. Bochkovskiy, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, and V . Koltun, “Depth pro: Sharp monocular metric depth in less than a second,” inInt. Conf. Learn. Represent., April 2025
2025
-
[69]
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,
D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 8934–8943
2018
-
[70]
Learning to estimate hidden motions with global motion aggregation,
S. Jiang, D. Campbell, Y. Lu, H. Li, and R. Hartley, “Learning to estimate hidden motions with global motion aggregation,” inInt. Conf. Comput. Vis., October 2021, pp. 9772–9781. 16
2021
-
[71]
Skflow: Learning optical flow with super kernels,
S. Sun, Y. Chen, Y. Zhu, G. Guo, and G. Li, “Skflow: Learning optical flow with super kernels,”Adv. Neural Inform. Process. Syst., vol. 35, pp. 11 313–11 326, 2022
2022
-
[72]
Flowformer: A transformer architecture for optical flow,
Z. Huang, X. Shi, C. Zhang, Q. Wang, K. C. Cheung, H. Qin, J. Dai, and H. Li, “Flowformer: A transformer architecture for optical flow,” inEur. Conf. Comput. Vis.Springer, 2022, pp. 668–685
2022
-
[73]
Dip: Deep inverse patchmatch for high-resolution optical flow,
Z. Zheng, N. Nie, Z. Ling, P . Xiong, J. Liu, H. Wang, and J. Li, “Dip: Deep inverse patchmatch for high-resolution optical flow,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 8925–8934
2022
-
[74]
Explicit motion disentangling for efficient optical flow estimation,
C. Deng, A. Luo, H. Huang, S. Ma, J. Liu, and S. Liu, “Explicit motion disentangling for efficient optical flow estimation,” inInt. Conf. Comput. Vis., 2023, pp. 9521–9530
2023
-
[75]
Craft: Cross-attentional flow transformer for robust optical flow,
X. Sui, S. Li, X. Geng, Y. Wu, X. Xu, Y. Liu, R. Goh, and H. Zhu, “Craft: Cross-attentional flow transformer for robust optical flow,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 17 602–17 611
2022
-
[76]
Recurrent partial kernel network for efficient optical flow estimation,
H. Morimitsu, X. Zhu, X. Ji, and X.-C. Yin, “Recurrent partial kernel network for efficient optical flow estimation,”Assoc. Ad- vancement Artif. Intell., vol. 38, no. 5, pp. 4278–4286, March 2024
2024
-
[77]
Global matching with overlapping attention for optical flow estimation,
S. Zhao, L. Zhao, Z. Zhang, E. Zhou, and D. Metaxas, “Global matching with overlapping attention for optical flow estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 17 592–17 601
2022
-
[78]
Sea-raft: Simple, efficient, accurate raft for optical flow,
Y. Wang, L. Lipson, and J. Deng, “Sea-raft: Simple, efficient, accurate raft for optical flow,” inEur. Conf. Comput. Vis., 2024, pp. 36–54
2024
-
[79]
Unifying flow, stereo and depth estimation,
H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,”IEEE T rans. Pattern Anal. Mach. Intell., 2023
2023
-
[80]
Memflow: Optical flow estimation and prediction with memory,
Q. Dong and Y. Fu, “Memflow: Optical flow estimation and prediction with memory,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 19 068–19 078
2024
-
[81]
Domain-invariant stereo matching networks,
F. Zhang, X. Qi, R. Yang, V . Prisacariu, B. Wah, and P . Torr, “Domain-invariant stereo matching networks,” inEur. Conf. Com- put. Vis.Springer, 2020, pp. 420–439
2020
-
[82]
Cfnet: Cascade and fused cost volume for robust stereo matching,
Z. Shen, Y. Dai, and Z. Rao, “Cfnet: Cascade and fused cost volume for robust stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 13 906–13 915
2021
-
[83]
Graftnet: Towards domain general- ized stereo matching with a broad-spectrum and task-oriented feature,
B. Liu, H. Yu, and G. Qi, “Graftnet: Towards domain general- ized stereo matching with a broad-spectrum and task-oriented feature,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 13 012–13 021
2022
-
[84]
Itsa: An information-theoretic approach to automatic shortcut avoidance and domain generalization in stereo matching networks,
W. Chuah, R. Tennakoon, R. Hoseinnezhad, A. Bab-Hadiashar, and D. Suter, “Itsa: An information-theoretic approach to automatic shortcut avoidance and domain generalization in stereo matching networks,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 13 022–13 032
2022
-
[85]
Domain generalized stereo matching via hierarchical visual transformation,
T. Chang, X. Yang, T. Zhang, and M. Wang, “Domain generalized stereo matching via hierarchical visual transformation,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 9559–9568
2023
-
[86]
Neural markov random field for stereo matching,
T. Guan, C. Wang, and Y.-H. Liu, “Neural markov random field for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 5459–5469
2024
-
[87]
Mocha-stereo: Motif channel attention network for stereo matching,
Z. Chen, W. Long, H. Yao, Y. Zhang, B. Wang, Y. Qin, and J. Wu, “Mocha-stereo: Motif channel attention network for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024
2024
-
[88]
Stereo any- where: Robust zero-shot deep stereo matching even where either stereo or mono fail,
L. Bartolomei, F. Tosi, M. Poggi, and S. Mattoccia, “Stereo any- where: Robust zero-shot deep stereo matching even where either stereo or mono fail,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025, pp. 1013–1027. Tingman Yanreceived the B.S. degree from Beihang University...
2025
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.