Pith. sign in

REVIEW 3 major objections 5 minor 3 cited by

MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching

T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read MatchAttention makes the query–key match offset a learnable attention center, yielding linear-complexity cross-view matching that leads Middlebury.

desk verdict A genuinely new linear-complexity attention primitive for stereo/flow with strong empirical results; the coverage concern is real but softened by direct residual supervision and the ablation's LRI fix. read the letter →

arxiv 2510.14260 v3 pith:L47QPQN4 submitted 2025-10-16 cs.CV

classification cs.CV
keywords MatchAttentionstereomatchingmechanismrelativepositionlinearcomplexityopticalflowocclusionhandlingcross-view
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes an attention operator, MatchAttention, for stereo and optical flow. Instead of attending to all target tokens or to a fixed local window, each query attends to a small contiguous window whose center is an explicitly predicted relative position—the disparity or flow itself. Because the center is arbitrary, long-range correspondences are reachable at linear cost in the number of tokens. The authors build a hierarchical coarse-to-fine decoder around it and report first-place average error on Middlebury, strong zero-shot generalization from synthetic training, and 4K UHD processing in 0.1 seconds. A sympathetic reader would take away that explicit matching constraints and efficient attention are compatible, not opposed.

What carries the argument

MatchAttention is a sliding-window attention whose sampling center is p_query + R_pos, with R_pos a per-token continuous 2D offset learned end-to-end. BilinearSoftmax distributes each query's attention over four integer sub-windows with bilinear weights, keeping the whole sampling path differentiable; negative L1-norm plus softmax acts as a normalized Laplace kernel suited to one-to-one matching. R_pos is updated by residual connections and fed as extra feature channels, so the network iteratively refines disparity or flow while aggregating features.

What would settle it

Construct or collect stereo/flow pairs with large, abrupt disparities and corrupt the coarse initialization—for instance by downsampling beyond 1/32 or adding repetitive texture—then check whether MatchAttention can still converge to the true match. A specific test: take a trained model, shift the initial R_pos by ±(window_radius+1) pixels at the coarsest scale, and measure whether fine-scale refinements recover; if errors jump to chance level, the coarse-to-fine coverage assumption is the binding constraint.

Watch

Extended reading notes

Core claim

The central claim is that the relative position between a query and its matched key can be treated as a learnable component of attention sampling, not as a positional embedding. MatchAttention computes, for each query, a contiguous window centered at the query position plus a learned offset; BilinearSoftmax makes sampling differentiable and sub-pixel accurate; residual connections embed the offset in feature channels so it refines layer by layer. Because the window is small and constant, complexity is linear. The paper then instantiates this in a stereo/flow decoder with negative L1-norm similarity, gated cross-attention, and a consistency-constrained loss to handle occlusion. If the claims

Load-bearing premise

The whole chain rests on the initial correlation at 1/32 scale being accurate enough that the true matching key falls inside the small sampling window (3x3 or 5x5) at every finer level; if the coarse estimate is off by more than roughly half a window, the correct key never receives attention weight or gradient.

Editorial extensions

If this is right

  • High-resolution stereo and flow inference becomes affordable: the model claims 4K UHD pairs in under 0.1 seconds and KITTI-resolution in 29 ms with a small GPU footprint.
  • Long-range correspondences no longer require quadratic attention; arbitrary offsets reach beyond the local window at constant window cost.
  • The predicted relative position is directly interpretable as disparity or flow, and self-attention offsets visualize where occluded regions are sampling, enabling explainable occlusion handling.
  • Because window size has no learnable parameters, the same weights can train with a large window and infer with a smaller one, decoupling training cost from deployment speed.
  • Zero-shot transfer from synthetic data to real benchmarks is reported as state of the art, suggesting that the explicit matching constraint improves generalization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If MatchAttention generalizes as the architecture suggests, it could replace global cross-attention in any cross-view matching stack—sparse feature matching, multi-view stereo, or feed-forward Gaussian splatting—where quadratic cost currently forces tiling.
  • A stress test worth running: feed stereo pairs with disparities whose coarse 1/32-scale initialization is deliberately biased beyond the 3x3 or 5x5 window radius; the predicted failure mode is a sharp accuracy cliff rather than graceful degradation.
  • The train-large/infer-small window property implies a compute-versus-accuracy knob for deployment that most attention designs lack; measuring that trade-off directly would be a useful extension.
  • The normalized-Laplace similarity suggests a closer link to kernel-based matching than to dot-product attention; exploring whether the L1 kernel can be replaced by a learned metric may extend the method to non-rectified views.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces MatchAttention, a sliding-window attention operator for cross-view matching in which the sampling center for each query is p_q + r, with r a learnable relative-position field. BilinearSoftmax makes the continuous sampling differentiable, and the relative positions are refined through residual connections in a hierarchical decoder. The authors instantiate this in MatchStereo/MatchFlow variants and report state-of-the-art Middlebury average error, strong zero-shot generalization, and high-resolution inference with linear attention complexity.

Significance. If the reported results hold, MatchAttention is a significant step toward high-resolution cross-view matching: it combines explicit correspondence prediction with attention-based aggregation while avoiding the quadratic cost of global cross-attention. The paper's strengths include a clear complexity analysis (Sec. 3.3), a broad evaluation across Middlebury, KITTI, ETH3D, Sintel and Spring, a released codebase with custom CUDA kernels, and systematic ablations (Sec. 5.3). The main open risk is the unquantified dependence of the long-range claim on the coarse-to-fine initialization, which should be addressed before the claims are fully convincing.

major comments (3)
  1. [Sec. 4.1, Eq. (5), Table 1] The central long-range claim relies on the initial R_pos from the 1/32 correlation (Eq. 9) being accurate enough that the true matching key lies inside the small sampling window at each finer scale. At 1/16, a 5x5 window tolerates an initial offset error of only about 32 full-resolution pixels before the true match leaves W_i; at 1/8 and 1/4 the budgets shrink further. If the true key is outside W_i, it receives zero attention weight and the attention module obtains no matching evidence; the dense L1 losses (Eqs. 13-16) provide direct training gradients, but at inference no such signal exists. The LRI ablation in Sec. 5.3 concedes that 'the initial correlation may not scale well at high resolution.' Please provide a perturbation study of the initial R_pos (or LRI vs no-LRI across resolutions) and discuss the failure mode, or qualify the 'arbitrary relative position' claim to the coverage
  2. [Sec. 3.2 vs Eq. (10)] The final paragraph of Sec. 3.2 claims that 'w in (5) does not introduce learnable parameters' and hence one can train with large w and infer with small w using the same weights. This is not true for the full model because Eq. (10) concatenates the flattened attention weights alpha_i to the input of the projection layer, making W'_p have c_v+(w+1)^2 input columns. Changing w changes the projection dimension. Please state that the size-independence holds only for the variant without the attention-weight injection, or modify the claim.
  3. [Abstract / Sec. 5.5] The arXiv metadata abstract names variants 'MatchAttentionXL' and 'MatchAttentionRT' and reports edge latencies (9.3 ms on RTX 4060 Ti, 79.1 ms on Jetson Orin NX at 1024x512) that do not appear in the main text, where the variants are MatchStereo-T/S/B. In addition, Sec. 5.5 claims state-of-the-art performance on KITTI 2015, but Table 9 shows DEFOM-Stereo (D1-all 1.41) and FoundationStereo (1.26) outperform MatchStereo-B (1.50). These claims need to be aligned with the reported numbers and with the actual model names.
minor comments (5)
  1. [Eqs. (6)-(7)] Eq. (6) uses exp(<q_i,k_j>) inside BilinearSoftmax, while Eq. (7) defines Softmax with the negative L1 norm exp(-gamma||q_i-k_j||_1). Please make the notation consistent by defining a single similarity function s(q,k).
  2. [Eq. (5)] The indexing W_i = {floor(p_i^k)+(u,v) | u,v=-w/2,...,w/2+1} is ambiguous for odd w (e.g., w=5) and mixes effective and expanded window sizes. Please define integer offsets explicitly (e.g., effective window offsets -(w-1)/2,...,(w-1)/2 and expanded window w+1).
  3. [Sec. 5.3] The cumulative analysis states that 'when any single component is removed from the full model, performance degradation occurs,' but the subtractive row Full-N improves on Full at half resolution (8.62 vs 9.23). The text should acknowledge this exception, as the component discussion itself does.
  4. [Table 8] In the provided manuscript text, Table 8 appears to contain no data rows; the reader only sees the caption and numbers cited in Sec. 5.5. Please verify that the table is rendered with the actual per-method AvgErr values.
  5. [Sec. 5.4] The statement that fixed-resolution inference 'improves accuracy through upsampling' for lower-resolution images is not uniformly supported: on FSD-Mix training, MatchStereo-B* is worse than MatchStereo-B on Middlebury F (6.56 vs 5.67) and on ETH3D (2.68 vs 1.36). Please qualify this benefit.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: relative-position attention is a supervised, externally benchmarked architecture; the acknowledged coarse-to-fine initialization risk is a robustness limitation, not a self-referential derivation.

full rationale

The paper's central mechanism — using R_pos to center a small sampling window (Eqs. 4–5) and refining R_pos through residual connections and BilinearSoftmax gradients (Eq. 8) — is an empirically trained and externally evaluated architecture, not a derivation whose output equals its input. R_pos is initialized from a 1/32-scale correlation (Eq. 9), but it is subsequently supervised against ground-truth disparity/flow in L_init, L_self, and L_cross (Eqs. 13–16), and zero-shot results are reported separately from finetuned benchmark results. The only statement in the paper that touches on a potential failure mode of the mechanism is the LRI ablation: 'The initial correlation may not scale well at high resolution and LRI can provide a more robust initial relative position.' This flags a coverage/robustness concern — if the initial R_pos is too far from the true match, the matching key lies outside the sampled window and receives no corrective gradient — but it is a limitation of the method, not circular reasoning. MatchAttention is compared against external methods on public benchmarks (Middlebury, KITTI, ETH3D, Spring), and references to the authors' own prior work (e.g., [42]) are contextual and not load-bearing. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from self-citation, and no known empirical pattern is merely relabeled. Therefore no circular step can be exhibited.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No new physical or ontological entities are introduced. R_pos is a learned tensor field that parameterizes the matching offset, not a postulated external quantity. The central empirical claims rest on network hyperparameters, the coarse-to-fine initialization assumption, and public benchmark ground truth.

free parameters (5)
  • beta (R_pos input scale) = learned; initialized (0.1, 0.1)
    Multiplicative factor for relative-position channels before concatenation with features (Sec 5.1.1); stabilizes training and affects matching dynamics.
  • A (non-occlusion mask threshold) = 1
    Threshold in Eq. 12 for the left-right consistency check; hand-set, controls which pixels are treated as occluded.
  • epsilon (consistency loss weight) = 0.01
    Weight of the backward consistency term in Eq. 16; hand-set.
  • gamma (L1 similarity scaling) = 1/sqrt(c_k)
    Inverse temperature in Eq. 7 for the negative-L1 Laplace kernel; chosen by analogy to dot-product attention scaling, not learned.
  • sampling window sizes w = (5,5,3,3) across decoder scales
    Table 1; window extent is the main complexity/accuracy knob and the central coverage assumption depends on it.
assumptions (4)
  • domain assumption Rectified epipolar geometry: stereo matching reduces to a 1D search along the x-axis
    Initial correlation for stereo is computed along the epipolar line and R_pos.x is treated as disparity (Sec 4.1, Eq. 9). On non-rectified or general cross-view pairs the method needs the epipolar constraint or a 2D search.
  • domain assumption Coarse-to-fine initialization coverage: the 1/32-scale initial correlation and convex upsampling place the true match within the small sampling window at each finer scale
    Eqs. 4-5 and Table 1; if R_pos is off by more than the window half-size, the true key is outside the sampled window and cannot contribute to attention or receive a corrective gradient.
  • domain assumption Bilinear interpolation and fused CUDA sampling are faithful differentiable implementations of the described BilinearSoftmax
    The correctness of sub-pixel relative-position gradients relies on this implementation; the paper states custom CUDA kernels but does not ship or verify them in the manuscript.
  • domain assumption Benchmark ground truth (Middlebury, KITTI, ETH3D, Sintel, Spring) is accurate and used consistently
    All reported error metrics depend on public benchmark labels and masks; the paper uses standard metrics but these cannot be verified from the text alone.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching." pith.science (2026). https://pith.science/paper/L47QPQN4

@misc{pith2026251014260,
  author       = {Pith},
  title        = {Pith review of: MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L47QPQN4}},
  note         = {Machine review of arXiv:2510.14260}
}
read the original abstract

Standard attention mechanisms are not well suited to stereo matching. Global attention scales quadratically and provides no explicit matching constraint, while local attention is efficient but loses long-range correspondences. We propose MatchAttention, an attention mechanism that embeds an explicit matching constraint into attention by treating the relative position between a query and its matched key as a learnable component of attention sampling. Centering a small contiguous sampling window on this learnable relative position enforces the matching constraint and supports long-range correspondence at strictly linear attention complexity. A differentiable contiguous attention sampling (CAS) operator enables sub-pixel accuracy, and cascaded MatchAttention blocks iteratively refine the relative positions through residual connections. We instantiate MatchAttention as a hierarchical coarse-to-fine stereo network with two variants. MatchAttentionXL targets accuracy and MatchAttentionRT targets real-time edge inference. MatchAttentionXL achieves state-of-the-art accuracy on Middlebury V3 and top results across KITTI 2012/2015 and ETH3D. MatchAttentionRT runs at 9.3 ms on RTX 4060 Ti and 79.1 ms on Jetson Orin NX 16 GB at 1024 x 512, making it the first stereo model to deliver real-time edge inference without sacrificing zero-shot generalization. The code is available at https://github.com/TingmanYan/MatchAttention.

Figures

Figures reproduced from arXiv: 2510.14260 by the authors.

Figure 1
Figure 1. MatchAttention block. Left: At the l-th layer of the transformer, the input feature tokens F l and the relative position Rl pos are concatenated in channel dimension and fed as input of MatchAttention after a LayerNorm. F l and Rl pos are updated by the residual connection which outputs F l+1 and R l+1 pos , where the token features are further processed by a feed-forward network. Center and right: Between layers l … view at source ↗
Figure 2
Figure 2. MatchAttention mechanism. The input is defined as [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 4
Figure 4. Cross-view matching and feature aggregation. For a given query [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: MatchDecoder architecture. Given cross-view inputs [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Visualization of self relative positions. Top: Color reference [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WAVE-Stereo: Warp-Aligned Volume Encoding for Stereo Matching

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WAVE-Stereo unifies correlation-volume search and feature-warping residual alignment in an iterative stereo matcher, achieving real-time zero-shot generalization on five benchmarks.

  2. WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching

    cs.CV 2026-03 accept novelty 6.0 of 10

    Warping alone, with a classification head before iterative regression, matches or beats cost-volume stereo methods on ETH3D, KITTI and Middlebury at higher speed.

  3. URS-Stereo: Uncertainty-Guided Residual Search for Real-Time Stereo Matching

    cs.CV 2026-07 conditional novelty 4.5 of 10

    Uncertainty-modulated residual offsets relocate local cost-volume centers in coarse-to-fine stereo matching, improving zero-shot disparity accuracy while keeping real-time speed.

Reference graph

Works this paper leans on

88 extracted references · 8 linked inside Pith · cited by 3 Pith papers

  1. [1]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y. Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 20 697–20 709

  2. [2]

    Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,

    Y. Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.-J. Cham, and J. Cai, “Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,”Eur. Conf. Comput. Vis., pp. 370–386, 2024

  3. [3]

    Splatt3r: Zero- shot gaussian splatting from uncalibrated image pairs,

    B. Smart, C. Zheng, I. Laina, and V . A. Prisacariu, “Splatt3r: Zero- shot gaussian splatting from uncalibrated image pairs,”arXiv preprint arXiv:2408.13912, 2024

  4. [4]

    No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images,

    B. Ye, S. Liu, H. Xu, L. Xueting, M. Pollefeys, M.-H. Yang, and P . Songyou, “No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images,”arXiv preprint arXiv:2410.24207, 2024

  5. [5]

    Unifying flow, stereo and depth estimation,

    H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,”IEEE T rans. Pattern Anal. Mach. Intell., vol. 45, no. 11, pp. 13 941–13 958, 2023

  6. [6]

    Grounding image matching in 3d with mast3r,

    V . Leroy, Y. Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” inEur. Conf. Comput. Vis., 2024, pp. 71–91

  7. [7]

    LoFTR: Detector- free local feature matching with transformers,

    J. Sun, Z. Shen, Y. Wang, H. Bao, and X. Zhou, “LoFTR: Detector- free local feature matching with transformers,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2021, pp. 8922–8931

  8. [8]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2016

Show all 88 references
  1. [9]

    Show, attend and tell: neural image caption generation with visual attention,

    K. Xu, J. L. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: neural image caption generation with visual attention,” inInt. Conf. Mach. Learn., 2015, p. 2048–2057

  2. [10]

    Effective approaches to attention-based neural machine translation,

    T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” inEmpirical Methods in Natural Language Processing, EMNLP, 2015, pp. 1412–1421

  3. [11]

    Deformable detr: Deformable transformers for end-to-end object detection

    X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection.” inInt. Conf. Learn. Represent., 2021

  4. [12]

    High-resolution stereo datasets with subpixel-accurate ground truth,

    D. Scharstein, H. Hirschm ¨uller, Y. Kitajima, G. Krathwohl, N. Ne ˇsi´c, X. Wang, and P . Westling, “High-resolution stereo datasets with subpixel-accurate ground truth,” inPattern Recog. German Conf., 2014, pp. 31–42

  5. [13]

    A multi-view stereo benchmark with high-resolution images and multi-camera videos,

    T. Schops, J. L. Schonberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high-resolution images and multi-camera videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 3260–3269

  6. [14]

    A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,

    N. Mayer, E. Ilg, P . Hausser, P . Fischer, D. Cremers, A. Dosovitskiy, and T. Brox, “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 4040–4048

  7. [15]

    Raft-stereo: Multilevel recurrent field transforms for stereo matching,

    L. Lipson, Z. Teed, and J. Deng, “Raft-stereo: Multilevel recurrent field transforms for stereo matching,” inInt. Conf. 3D Vision, 2021, pp. 218–227

  8. [16]

    Foundationstereo: Zero-shot stereo matching,

    B. Wen, M. Trepte, J. Aribido, J. Kautz, O. Gallo, and S. Birchfield, “Foundationstereo: Zero-shot stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025, pp. 5249–5260

  9. [17]

    Are we ready for autonomous driving? the kitti vision benchmark suite,

    A. Geiger, P . Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” inIEEE Conf. Comput. Vis. Pattern Recog., 2012

  10. [18]

    Joint 3d estimation of vehicles and scene flow,

    M. Menze, C. Heipke, and A. Geiger, “Joint 3d estimation of vehicles and scene flow,”ISPRS Ann. Photogrammetry, Remote Sens. Spatial Inf. Sciences, vol. 2, p. 427, 2015

  11. [19]

    Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo,

    L. Mehl, J. Schmalfuss, A. Jahedi, Y. Nalivayko, and A. Bruhn, “Spring: A high-resolution high-detail dataset and benchmark for scene flow, optical flow and stereo,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 4981–4991

  12. [20]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” inInt. Conf. Learn. Represent., 2015

  13. [21]

    Bridging the divide: Reconsidering softmax and linear attention,

    D. Han, Y. Pu, Z. Xia, Y. Han, X. Pan, X. Li, J. Lu, S. Song, and G. Huang, “Bridging the divide: Reconsidering softmax and linear attention,” inAdv. Neural Inform. Process. Syst., vol. 37, 2024, pp. 79 221–79 245

  14. [22]

    Longformer: The long- document transformer,

    I. Beltagy, M. E. Peters, and A. Cohan, “Longformer: The long- document transformer,”arXiv:2004.05150, 2020

  15. [23]

    Big bird: Transformers for longer sequences,

    M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P . Pham, A. Ravula, Q. Wang, L. Yanget al., “Big bird: Transformers for longer sequences,”Adv. Neural Inform. Process. Syst., vol. 33, 2020

  16. [24]

    Sparser is faster and less is more: Efficient sparse attention for long-range transformers,

    C. Lou, Z. Jia, Z. Zheng, and K. Tu, “Sparser is faster and less is more: Efficient sparse attention for long-range transformers,” arXiv:2406.16747, 2024. 15

  17. [25]

    Neighborhood attention transformer,

    A. Hassani, S. Walton, J. Li, S. Li, and H. Shi, “Neighborhood attention transformer,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 6185–6194

  18. [26]

    Clear: Conv-like lineariza- tion revs pre-trained diffusion transformers up,

    S. Liu, Z. Tan, and X. Wang, “Clear: Conv-like lineariza- tion revs pre-trained diffusion transformers up,”arXiv preprint arXiv:2412.16112, 2024

  19. [27]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inInt. Conf. Comput. Vis., October 2021, pp. 10 012– 10 022

  20. [28]

    Bevformer: Learning bird’s-eye-view representation from multi- camera images via spatiotemporal transformers,

    Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai, “Bevformer: Learning bird’s-eye-view representation from multi- camera images via spatiotemporal transformers,” inEur. Conf. Comput. Vis., 2022, pp. 1–18

  21. [29]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” inAdv. Neural Inform. Process. Syst., vol. 30, 2017

  22. [30]

    Swin transformer v2: Scaling up capacity and resolution,

    Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo, “Swin transformer v2: Scaling up capacity and resolution,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 12 009–12 019

  23. [31]

    Transnext: Robust foveal visual perception for vision transformers,

    D. Shi, “Transnext: Robust foveal visual perception for vision transformers,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 17 773–17 783

  24. [32]

    Roformer: Enhanced transformer with rotary position embedding,

    J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,”Neuro- comput., vol. 568, no. C, February 2024

  25. [33]

    The llama 3 herd of models,

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al- Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, and et al., “The llama 3 herd of models,”arXiv:2407.21783, 2024

  26. [34]

    Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow,

    P . Weinzaepfel, T. Lucas, V . Leroy, Y. Cabon, V . Arora, R. Br´egier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud, “Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow,” inInt. Conf. Comput. Vis., 2023, pp. 17 969–17 980

  27. [35]

    Cameras as relative positional encoding,

    R. Li, B. Yi, J. Liu, H. Gao, Y. Ma, and A. Kanazawa, “Cameras as relative positional encoding,”arXiv:2507.10496, 2025

  28. [36]

    Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,

    Z. Li, X. Liu, N. Drenkow, A. Ding, F. X. Creighton, R. H. Taylor, and M. Unberath, “Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers,” inInt. Conf. Comput. Vis., 2021, pp. 6197–6206

  29. [37]

    End-to-end learning of geometry and context for deep stereo regression,

    A. Kendall, H. Martirosyan, S. Dasgupta, P . Henry, R. Kennedy, A. Bachrach, and A. Bry, “End-to-end learning of geometry and context for deep stereo regression,” inInt. Conf. Comput. Vis., 2017, pp. 66–75

  30. [38]

    Pyramid stereo matching network,

    J.-R. Chang and Y.-S. Chen, “Pyramid stereo matching network,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 5410–5418

  31. [39]

    Correlate-and-excite: Real-time stereo matching via guided cost volume excitation,

    A. Bangunharcana, J. W. Cho, S. Lee, I. S. Kweon, K.-S. Kim, and S. Kim, “Correlate-and-excite: Real-time stereo matching via guided cost volume excitation,” inInt. Conf. Intell. Robots Syst., 2021

  32. [40]

    Raft: Recurrent all-pairs field transforms for optical flow,

    Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” inEur. Conf. Comput. Vis.Springer, 2020, pp. 402–419

  33. [41]

    Continuous 3D Label Stereo Matching using Local Expansion Moves,

    T. Taniai, Y. Matsushita, Y. Sato, and T. Naemura, “Continuous 3D Label Stereo Matching using Local Expansion Moves,”IEEE T rans. Pattern Anal. Mach. Intell., vol. 40, no. 11, pp. 2725–2739, 2018

  34. [42]

    Hierarchical belief propagation on image segmentation pyramid,

    T. Yan, X. Yang, G. Yang, and Q. Zhao, “Hierarchical belief propagation on image segmentation pyramid,”IEEE T rans. Image Process., vol. 32, pp. 4432–4442, 2023

  35. [43]

    Iterative geometry encod- ing volume for stereo matching,

    G. Xu, X. Wang, X. Ding, and X. Yang, “Iterative geometry encod- ing volume for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 21 919–21 928

  36. [44]

    Selective-stereo: Adaptive frequency information selection for stereo matching,

    X. Wang, G. Xu, H. Jia, and X. Yang, “Selective-stereo: Adaptive frequency information selection for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 701–19 710

  37. [45]

    All-in-one: Transferring vision foundation models into stereo matching,

    J. Zhou, H. Zhang, J. Yuan, P . Ye, T. Chen, H. Jiang, M. Chen, and Y. Zhang, “All-in-one: Transferring vision foundation models into stereo matching,” inAssoc. Advancement Artif. Intell., February 2025

  38. [46]

    DINOv2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V . Vo, M. Szafraniec, V . Khalidov, P . Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P .-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P . Labatut...

  39. [47]

    Segment anything,

    A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P . Dollar, and R. Gir- shick, “Segment anything,” inInt. Conf. Comput. Vis., October 2023, pp. 4015–4026

  40. [48]

    Depth anything v2,

    L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” inAdv. Neural Inform. Process. Syst., 2024

  41. [49]

    Vision transformers for dense prediction,

    R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inInt. Conf. Comput. Vis., October 2021, pp. 12 179–12 188

  42. [50]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” inInt. Conf. Learn. Represent., 2020

  43. [51]

    Trans- formers as support vector machines,

    D. A. Tarzanagh, Y. Li, C. Thrampoulidis, and S. Oymak, “Trans- formers as support vector machines,” inAdv. Neural Inform. Pro- cess. Syst., 2023

  44. [52]

    Metaformer baselines for vision,

    W. Yu, C. Si, P . Zhou, M. Luo, Y. Zhou, J. Feng, S. Yan, and X. Wang, “Metaformer baselines for vision,”IEEE T rans. Pattern Anal. Mach. Intell., vol. 46, no. 2, pp. 896–912, 2024

  45. [53]

    Gaussian error linear units (GELUs),

    D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELUs),”arXiv:1606.08415, 2023

  46. [54]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P . Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” 2015, pp. 234–241

  47. [55]

    Conditional positional encodings for vision transformers,

    X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, “Conditional positional encodings for vision transformers,” inInt. Conf. Learn. Represent., 2023

  48. [56]

    Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,

    S. Elfwing, E. Uchibe, and K. Doya, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,”Neural Netw., vol. 107, pp. 3–11, 2018

  49. [57]

    Adaptive multi-modal cross-entropy loss for stereo matching,

    P . Xu, Z. Xiang, C. Qiao, J. Fu, and X. Zhao, “Adaptive multi-modal cross-entropy loss for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024

  50. [58]

    Decoupled weight decay regulariza- tion,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regulariza- tion,” inInt. Conf. Learn. Represent., 2019

  51. [59]

    A naturalistic open source movie for optical flow evaluation,

    D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black, “A naturalistic open source movie for optical flow evaluation,” inEur. Conf. Comput. Vis.Springer, 2012, pp. 611–625

  52. [60]

    Practical stereo matching via cascaded recurrent network with adaptive correlation,

    J. Li, P . Wang, P . Xiong, T. Cai, Z. Yan, L. Yang, J. Liu, H. Fan, and S. Liu, “Practical stereo matching via cascaded recurrent network with adaptive correlation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 16 263–16 272

  53. [61]

    Falling things: A synthetic dataset for 3d object detection and pose estimation,

    J. Tremblay, T. To, and S. Birchfield, “Falling things: A synthetic dataset for 3d object detection and pose estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 2038–2041

  54. [62]

    In- stereo2k: a large real dataset for stereo matching in indoor scenes,

    W. Bao, W. Wang, Y. Xu, Y. Guo, S. Hong, and X. Zhang, “In- stereo2k: a large real dataset for stereo matching in indoor scenes,” Sci. China Inf. Sci., vol. 63, pp. 1–11, 2020

  55. [63]

    Hierarchical deep stereo matching on high-resolution images,

    G. Yang, J. Manela, M. Happold, and D. Ramanan, “Hierarchical deep stereo matching on high-resolution images,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 5515–5524

  56. [64]

    Virtual kitti 2,

    Y. Cabon, N. Murray, and M. Humenberger, “Virtual kitti 2,” arXiv:2001.10773, 2020

  57. [65]

    Flownet: Learning optical flow with convolutional networks,

    A. Dosovitskiy, P . Fischer, E. Ilg, P . H¨ausser, C. Hazırbas ¸, V . Golkov, P . v. d. Smagt, D. Cremers, and T. Brox, “Flownet: Learning optical flow with convolutional networks,” inInt. Conf. Comput. Vis., 2015, pp. 2758–2766

  58. [66]

    The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driv- ing,

    D. Kondermann, R. Nair, K. Honauer, K. Krispin, J. Andrulis, A. Brock, B. G ¨ussefeld, M. Rahimimoghaddam, S. Hofmann, C. Brenner, and B. J ¨ahne, “The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driv- ing,” inIEEE Conf. Comput. Vi...

  59. [67]

    Defom-stereo: Depth foundation model based stereo matching,

    H. Jiang, Z. Lou, L. Ding, R. Xu, M. Tan, W. Jiang, and R. Huang, “Defom-stereo: Depth foundation model based stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025, pp. 21 857– 21 867

  60. [68]

    Depth pro: Sharp monocular metric depth in less than a second,

    A. Bochkovskiy, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, and V . Koltun, “Depth pro: Sharp monocular metric depth in less than a second,” inInt. Conf. Learn. Represent., April 2025

  61. [69]

    Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,

    D. Sun, X. Yang, M.-Y. Liu, and J. Kautz, “Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 8934–8943

  62. [70]

    Learning to estimate hidden motions with global motion aggregation,

    S. Jiang, D. Campbell, Y. Lu, H. Li, and R. Hartley, “Learning to estimate hidden motions with global motion aggregation,” inInt. Conf. Comput. Vis., October 2021, pp. 9772–9781. 16

  63. [71]

    Skflow: Learning optical flow with super kernels,

    S. Sun, Y. Chen, Y. Zhu, G. Guo, and G. Li, “Skflow: Learning optical flow with super kernels,”Adv. Neural Inform. Process. Syst., vol. 35, pp. 11 313–11 326, 2022

  64. [72]

    Flowformer: A transformer architecture for optical flow,

    Z. Huang, X. Shi, C. Zhang, Q. Wang, K. C. Cheung, H. Qin, J. Dai, and H. Li, “Flowformer: A transformer architecture for optical flow,” inEur. Conf. Comput. Vis.Springer, 2022, pp. 668–685

  65. [73]

    Dip: Deep inverse patchmatch for high-resolution optical flow,

    Z. Zheng, N. Nie, Z. Ling, P . Xiong, J. Liu, H. Wang, and J. Li, “Dip: Deep inverse patchmatch for high-resolution optical flow,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 8925–8934

  66. [74]

    Explicit motion disentangling for efficient optical flow estimation,

    C. Deng, A. Luo, H. Huang, S. Ma, J. Liu, and S. Liu, “Explicit motion disentangling for efficient optical flow estimation,” inInt. Conf. Comput. Vis., 2023, pp. 9521–9530

  67. [75]

    Craft: Cross-attentional flow transformer for robust optical flow,

    X. Sui, S. Li, X. Geng, Y. Wu, X. Xu, Y. Liu, R. Goh, and H. Zhu, “Craft: Cross-attentional flow transformer for robust optical flow,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 17 602–17 611

  68. [76]

    Recurrent partial kernel network for efficient optical flow estimation,

    H. Morimitsu, X. Zhu, X. Ji, and X.-C. Yin, “Recurrent partial kernel network for efficient optical flow estimation,”Assoc. Ad- vancement Artif. Intell., vol. 38, no. 5, pp. 4278–4286, March 2024

  69. [77]

    Global matching with overlapping attention for optical flow estimation,

    S. Zhao, L. Zhao, Z. Zhang, E. Zhou, and D. Metaxas, “Global matching with overlapping attention for optical flow estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 17 592–17 601

  70. [78]

    Sea-raft: Simple, efficient, accurate raft for optical flow,

    Y. Wang, L. Lipson, and J. Deng, “Sea-raft: Simple, efficient, accurate raft for optical flow,” inEur. Conf. Comput. Vis., 2024, pp. 36–54

  71. [79]

    Unifying flow, stereo and depth estimation,

    H. Xu, J. Zhang, J. Cai, H. Rezatofighi, F. Yu, D. Tao, and A. Geiger, “Unifying flow, stereo and depth estimation,”IEEE T rans. Pattern Anal. Mach. Intell., 2023

  72. [80]

    Memflow: Optical flow estimation and prediction with memory,

    Q. Dong and Y. Fu, “Memflow: Optical flow estimation and prediction with memory,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 19 068–19 078

  73. [81]

    Domain-invariant stereo matching networks,

    F. Zhang, X. Qi, R. Yang, V . Prisacariu, B. Wah, and P . Torr, “Domain-invariant stereo matching networks,” inEur. Conf. Com- put. Vis.Springer, 2020, pp. 420–439

  74. [82]

    Cfnet: Cascade and fused cost volume for robust stereo matching,

    Z. Shen, Y. Dai, and Z. Rao, “Cfnet: Cascade and fused cost volume for robust stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 13 906–13 915

  75. [83]

    Graftnet: Towards domain general- ized stereo matching with a broad-spectrum and task-oriented feature,

    B. Liu, H. Yu, and G. Qi, “Graftnet: Towards domain general- ized stereo matching with a broad-spectrum and task-oriented feature,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 13 012–13 021

  76. [84]

    Itsa: An information-theoretic approach to automatic shortcut avoidance and domain generalization in stereo matching networks,

    W. Chuah, R. Tennakoon, R. Hoseinnezhad, A. Bab-Hadiashar, and D. Suter, “Itsa: An information-theoretic approach to automatic shortcut avoidance and domain generalization in stereo matching networks,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 13 022–13 032

  77. [85]

    Domain generalized stereo matching via hierarchical visual transformation,

    T. Chang, X. Yang, T. Zhang, and M. Wang, “Domain generalized stereo matching via hierarchical visual transformation,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 9559–9568

  78. [86]

    Neural markov random field for stereo matching,

    T. Guan, C. Wang, and Y.-H. Liu, “Neural markov random field for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2024, pp. 5459–5469

  79. [87]

    Mocha-stereo: Motif channel attention network for stereo matching,

    Z. Chen, W. Long, H. Yao, Y. Zhang, B. Wang, Y. Qin, and J. Wu, “Mocha-stereo: Motif channel attention network for stereo matching,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024

  80. [88]

    Stereo any- where: Robust zero-shot deep stereo matching even where either stereo or mono fail,

    L. Bartolomei, F. Tosi, M. Poggi, and S. Mattoccia, “Stereo any- where: Robust zero-shot deep stereo matching even where either stereo or mono fail,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025, pp. 1013–1027. Tingman Yanreceived the B.S. degree from Beihang University...

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.