REVIEW 3 major objections 3 minor 2 cited by
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This survey claims that sensing modality, not algorithm family, should be the organizing principle of feature matching—a lens that unifies RGB, 3D, medical, and vision-language methods.
desk verdict A useful modality-aware survey whose comparative benchmark tables need a protocol audit before the numbers can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a modality-driven taxonomy that treats sensing modality as a first-class factor defining what matchability means and which priors are exploitable. The taxonomy distinguishes geometric correspondence (RGB, 3D, medical registration) from semantic cross-modal alignment (vision-language), and within each branch traces the shift from handcrafted pipelines to learned, often detector-free, architectures. The survey uses this lens to separate transferable principles—attention-based matching, contrastive objectives, coarse-to-fine refinement—from modality-specific modeling such as geometric invariance in LiDAR, intensity-robust similarity in medical images, and semantic grounding in vision-language.
What would settle it
Re-running the compiled comparisons under a single protocol—for instance, evaluating all methods in Table 12 on the same COCO 5K split with identical fine-tuning settings, and all 3D methods on the same 3DMatch sampling—would settle whether the reported ordering reflects modality-driven progress or protocol artifacts.
Extended reading notes
Core claim
The paper's core claim is that feature matching methods are best understood through the lens of data modality rather than algorithm family alone. It organizes the field into single-modality and cross-modality branches: RGB and 3D data form single-modality geometric correspondence, medical imaging sits at the intersection with both same- and cross-modality registration, and vision-language matching forms a semantic alignment branch. Within each branch it traces an evolution from handcrafted pipelines to learned and detector-free approaches, and it reads this evolution as evidence that modality-specific assumptions—appearance, geometry, deformation, semantics—drive the design of detectors, descriptors, and matchers.
Load-bearing premise
The survey's conclusions stand on the assumption that the benchmark numbers it compiles from external papers are accurate and comparable across protocols, since it does not re-run experiments.
Editorial extensions
If this is right
- If the modality-aware taxonomy is correct, future systems can expect shared principles such as attention, contrastive learning, and coarse-to-fine matching to transfer across RGB, 3D, and vision-language, while intensity handling and deformation modeling remain domain-specific problems.
- The explicit cross-modality thread implies that medical registration and vision-language alignment are not separate fields but instances of the same matching problem with different similarity definitions and supervision signals.
- The survey's comparative tables suggest that detector-free transformer matchers and learned dense correspondence have displaced handcrafted pipelines as the default for RGB and 3D, with clinical and open-vocabulary applications following the same arc.
- A unified multi-modal, multi-task framework that handles retrieval, registration, and captioning within one system is the paper's stated trajectory, made plausible by the common mechanisms the taxonomy identifies.
- Benchmark practice would benefit from evaluating methods under modality-specific protocols rather than a single generic metric, because matchability itself is modality-dependent.
Reading between the lines
- Beyond the paper: the taxonomy predicts that a modality spectrum from photometric (RGB) to geometric (LiDAR) to semantic (text) could guide descriptor design, with hybrid representations at the boundaries—such as RGB-D features—outperforming single-modality ones.
- Beyond the paper: if modality is truly first-class, benchmark design should stratify by acquisition sensor and protocol, so that the survey's compiled numbers (for example, Table 12 mixing zero-shot and fine-tuned results) become a testable hypothesis rather than a fixed ranking.
- Beyond the paper: the identified transferable principles suggest a concrete experiment—training one attention-based matching backbone on paired RGB and LiDAR data and measuring whether gains transfer across sensors—which the survey itself does not run.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The survey proposes a modality-aware perspective on feature matching, arguing that sensing modality should be a first-class factor determining what "matchability" means and which priors are exploitable. It reviews handcrafted and deep learning methods across RGB images, 3D/RGB-D/LiDAR data, medical images, and vision-language tasks, and it positions cross-modality matching (medical registration, vision-language alignment) as a unifying thread. The paper states three contributions: a modality-driven taxonomy, an explicit treatment of modality gaps, and a comparative synthesis with transferable insights. The body is organized by modality, with benchmark tables for HPatches, 3DMatch, Learn2Reg, and MSCOCO image-text retrieval, plus dataset/protocol summaries and appendices on embodied AI, VQA, and future directions.
Significance. If its benchmark evidence is made reliable, this would be a useful synthesis: it connects four areas usually surveyed separately, and it identifies shared principles (contrastive objectives, cross-attention, geometric-invariance requirements) as well as modality-specific constraints. Strengths include explicit differentiation from three prior surveys, broad and current coverage including 2024-2025 works, and characterizations of major methods that are consistent with the literature I can verify. There are no derived parameters or machine-checked proofs to assess; this is a literature review. The main weakness is not the taxonomy but the quantitative layer: the comparative claims rest on tables whose evaluation protocols are not comparable as presented, and one of the three stated contributions is explicitly comparative.
major comments (3)
- [§5.5, Table 12] The table caption states that all rows are COCO fine-tuned on the 5K protocol except CLIP, which is zero-shot, but the table then ranks CLIP in the same list as the fine-tuned methods. This makes the numeric ordering (e.g., VSE++ 41.3 vs. CLIP 58.4 I2T R@1) reflect an evaluation-protocol difference rather than a method-family difference, directly weakening contribution 3. Please split zero-shot and fine-tuned results into separate tables or add an explicit protocol column, and revise any prose that invites cross-protocol ranking.
- [§3.4, Table 6] The caption asserts one standard protocol (f=5000, tau1=10cm, tau2=5%), but the rows are taken from different source papers whose sampling schemes, RANSAC settings, and validation details are known to vary. With adjacent methods separated by roughly one FMR point (CoFiNet 98.1 vs. HST 98.8), even a small undisclosed protocol difference can change the ordering. Please provide row-level protocol provenance (source table, split, parameter choices) or explicitly state which rows are directly comparable and which are indicative only.
- [§4.2-4.3, Table 9] Table 9 reports mean Dice for "Brain MRI inter-patient registration on Learn2Reg", but the cited works evaluate on different Learn2Reg tasks or splits: SynthMorph and VoxelMorph report on their own validation settings, while Zhang et al. is a challenge submission, and the table does not indicate which task or split each value came from. The ranking TransMorph > VoxelMorph may therefore reflect evaluation setup rather than method quality. Please add per-row test-set/split information or move these numbers to a clearly labeled table of heterogeneous settings.
minor comments (3)
- [Abstract] The metadata abstract calls SuperPoint a detector-free strategy, whereas §2.2 correctly describes it as a self-supervised CNN detector-descriptor and positions SuperGlue as the learned matcher; align the abstract with the body text.
- [§5/Appendix A] The vision-language dataset table appears twice (Table 10 in §5 and Table 13 in Appendix A) with largely overlapping entries; consolidate to avoid duplication.
- [Tables 9 and 12] Several rows cite venue/year information that is inconsistent with the arXiv status at submission time (e.g., D2S-VSE listed as ICCV 2025 and MaxMatch as ACL 2025); please verify against published versions or mark them as preprints.
Circularity Check
No significant circularity: the survey is a literature synthesis with no fitted parameters, derivations, or self-referential claims that reduce to their own inputs.
full rationale
This paper is a comprehensive survey, not a derivation or prediction pipeline. Its stated contributions are a modality-driven taxonomy, an explicit treatment of cross-modality gaps, and comparative synthesis of existing methods. None of these contributions is obtained by fitting data, by defining a quantity in terms of the quantity it claims to explain, or by invoking a self-citation chain to forbid alternatives. The benchmark tables (Tables 3, 6, 9, and 12) compile numbers from external papers without re-running experiments; however, compilation of external results is not circular reasoning, and the paper transparently states the source of each table, including the explicit note in Table 12 that CLIP is reported in the zero-shot setting while other rows are COCO fine-tuned. That protocol inconsistency is a correctness/comparability concern, not a circularity concern. The authors do cite their own prior works (e.g., Refs. [70], [128]-[130], [134], [213], [220], [238], [239]), but these citations are used only to point to specific methods or datasets within the surveyed literature; they are not load-bearing premises for the survey's organizational claims, and the survey does not rely on any uniqueness theorem or ansatz imported from the authors' own prior work. The central 'modality-aware perspective' is an explicitly adopted analytical lens, stated as such rather than derived, so it cannot be circular. Accordingly, the appropriate finding is no significant circularity.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey." pith.science (2026). https://pith.science/paper/UOSV5OJE
@misc{pith2026250722791,
author = {Pith},
title = {Pith review of: Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/UOSV5OJE}},
note = {Machine review of arXiv:2507.22791}
}
read the original abstract
Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matching, exploring traditional handcrafted methods and emphasizing contemporary deep learning approaches across various modalities, including RGB images, depth images, 3D point clouds, LiDAR scans, medical images, and vision-language interactions. Traditional methods, leveraging detectors like Harris corners and descriptors such as SIFT and ORB, demonstrate robustness under moderate intra-modality variations but struggle with significant modality gaps. Contemporary deep learning-based methods, exemplified by detector-free strategies like CNN-based SuperPoint and transformer-based LoFTR, substantially improve robustness and adaptability across modalities. We highlight modality-aware advancements, such as geometric and depth-specific descriptors for depth images, sparse and dense learning methods for 3D point clouds, attention-enhanced neural networks for LiDAR scans, and specialized solutions like the MIND descriptor for complex medical image matching. Cross-modal applications, particularly in medical image registration and vision-language tasks, underscore the evolution of feature matching to handle increasingly diverse data interactions.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
A weakly supervised 3D point cloud segmentation method that back-projects Semantic-SAM 2D masks into 3D, propagates sparse labels inside masks, and uses reliability-filtered pseudo labels, reporting state-of-the-art m...
-
Uncertainty Awareness on Unsupervised Domain Adaptation for Time Series Data
A UDA framework with multi-scale input mixing and Dirichlet-prior uncertainty estimation improves F1 and calibration on five time-series benchmarks.
Reference graph
Works this paper leans on
-
[1]
Alexandre Alahi, Raphael Ortiz, and Pierre Vandergheynst. 2012. Freak: Fast retina keypoint. In2012 IEEE conference on computer vision and pattern recognition. Ieee, 510–517
2012
-
[2]
Hani Alomari, Anushka Sivakumar, Andrew Zhang, and Chris Thomas. 2025. Maximal Matching Matters: Preventing Representation Collapse for Robust Cross-Modal Retrieval.arXiv preprint arXiv:2506.21538(2025)
work page Pith review arXiv 2025
-
[3]
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. InProceedings of the IEEE conference on computer vision and pattern recognition. 6077–6086
2018
-
[4]
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. 2018. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. InProceedings of the IEEE conference on computer vision and pattern recognition. 3674–3683
2018
-
[5]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. InProceedings of the IEEE international conference on computer vision. 2425–2433
2015
-
[6]
Sheng Ao, Qingyong Hu, Bo Yang, Andrew Markham, and Yulan Guo. 2021. Spinnet: Learning a general surface descriptor for 3d point cloud registration. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11753–11762
2021
-
[7]
Brian B Avants, Charles L Epstein, Murray Grossman, and James C Gee. 2008. Symmetric diffeomorphic image registration with cross-correlation: evaluating automated labeling of elderly and neurodegenerative brain.Medical image analysis12, 1 (2008), 26–41
2008
-
[8]
Dzmitry Bahdanau, Harm de Vries, Timothy J O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron Courville. 2019. Closure: Assessing systematic generalization of clevr models.arXiv preprint arXiv:1912.05783 (2019)
arXiv 2019
Show all 254 references
-
[9]
Xuyang Bai, Zixin Luo, Lei Zhou, Hongbo Fu, Long Quan, and Chiew-Lan Tai. 2020. D3feat: Joint learning of dense detection and description of 3d local features. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 6359–6367
2020
-
[10]
Guha Balakrishnan, Amy Zhao, Mert R Sabuncu, John Guttag, and Adrian V Dalca. 2019. Voxelmorph: a learning framework for deformable medical image registration.IEEE transactions on medical imaging38, 8 (2019), 1788–1800
2019
-
[11]
Vassileios Balntas, Karel Lenc, Andrea Vedaldi, and Krystian Mikolajczyk. 2017. HPatches: A benchmark and evaluation of handcrafted and learned local descriptors. InProceedings of the IEEE conference on computer vision and pattern recognition. 5173–5182
2017
-
[12]
Vassileios Balntas, Edgar Riba, Daniel Ponsa, and Krystian Mikolajczyk. 2016. Learning local feature descriptors with triplets and shallow convolutional neural networks.. InBmvc, Vol. 1. 3
2016
-
[13]
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran. 2018. Zero-shot object detection. InProceedings of the European conference on computer vision (ECCV). 384–400
2018
-
[14]
Daniel Barath and Jiří Matas. 2018. Graph-cut RANSAC. InProceedings of the IEEE conference on computer vision and pattern recognition. 6733–6741
2018
-
[15]
Axel Barroso-Laguna, Sowmya Munukutla, Victor Adrian Prisacariu, and Eric Brachmann. 2024. Matching 2d images in 3d: Metric relative pose from metric correspondences. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4852–4863
2024
-
[16]
Matteo Bastico, Etienne Decencière, Laurent Corté, Yannick Tillier, and David Ryckelynck. 2024. Coupled Laplacian Eigenmaps for Locally-Aware 3D Rigid Point Cloud Matching. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3447–3458
2024
-
[17]
Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. 2006. Surf: Speeded up robust features. InComputer Vision–ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006. Proceedings, Part I 9. Springer, 404–417
2006
-
[18]
Neslihan Bayramoglu and A Aydin Alatan. 2010. Shape index SIFT: Range image recognition using local features. In 2010 20th International Conference on Pattern Recognition. IEEE, 352–355
2010
-
[19]
M Faisal Beg, Michael I Miller, Alain Trouvé, and Laurent Younes. 2005. Computing large deformation metric mappings via geodesic flows of diffeomorphisms.International journal of computer vision61 (2005), 139–157. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: J...
2005
-
[20]
Paul J Besl and Neil D McKay. 1992. Method for registration of 3-D shapes.Sensor fusion IV: control paradigms and data structures1611, 586–606
1992
-
[21]
Biomedical Image Analysis Group, Imperial College London. 2025. IXI Dataset
2025
-
[22]
Bookstein
Fred L. Bookstein. 1989. Principal warps: Thin-plate splines and the decomposition of deformations.IEEE Transactions on pattern analysis and machine intelligence11, 6 (1989), 567–585
1989
-
[23]
Gary Bradski, Adrian Kaehler, et al. 2000. OpenCV.Dr. Dobb’s journal of software tools3, 2 (2000)
2000
-
[24]
Lisa Gottesfeld Brown. 1992. A survey of image registration techniques.ACM computing surveys (CSUR)24, 4 (1992), 325–376
1992
-
[25]
Matthew Brown, Gang Hua, and Simon Winder. 2010. Discriminative Learning of Local Image Descriptors. In European Conference on Computer Vision (ECCV) (LNCS, Vol. 6313). Springer, 677–691
2010
-
[26]
Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Pérez. 2019. Zero-shot semantic segmentation.Advances in Neural Information Processing Systems32 (2019)
2019
-
[27]
Michael Calonder, Vincent Lepetit, Christoph Strecha, and Pascal Fua. 2010. Brief: Binary robust independent elementary features. InComputer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part IV 11. ...
2010
-
[28]
Xiaohuan Cao, Jianhua Yang, Yaozong Gao, Yanrong Guo, Guorong Wu, and Dinggang Shen. 2017. Dual-core steered non-rigid registration for multi-modal images via bi-directional image synthesis.Medical image analysis41 (2017), 18–31
2017
-
[29]
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. 2017. Matterport3d: Learning from rgb-d data in indoor environments.arXiv preprint arXiv:1709.06158
2017 arXiv
-
[30]
Hongkai Chen, Zixin Luo, Lei Zhou, Yurun Tian, Mingmin Zhen, Tian Fang, David Mckinnon, Yanghai Tsin, and Long Quan. 2022. Aspanformer: Detector-free image matching with adaptive span transformer. InEuropean Conference on Computer Vision. Springer, 20–36
2022
-
[31]
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2019. Touchdown: Natural language navigation and spatial reasoning in visual street environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12538–12547
2019
-
[32]
Junyu Chen, Eric C Frey, Yufan He, William P Segars, Ye Li, and Yong Du. 2022. Transmorph: Transformer for unsupervised medical image registration.Medical image analysis82 (2022), 102615
2022
-
[33]
Junyu Chen, Yufan He, Eric C Frey, Ye Li, and Yong Du. 2021. Vit-v-net: Vision transformer for unsupervised volumetric medical image registration.arXiv preprint arXiv:2104.06468
2021 arXiv
-
[34]
Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. 2021. History aware multimodal transformer for vision-and-language navigation.Advances in neural information processing systems34 (2021), 5834–5847
2021
-
[35]
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick
-
[36]
Xiang Chen, Renjiu Hu, Jinwei Zhang, Yuxi Zhang, Xinyao Yue, Min Liu, Yaonan Wang, and Hang Zhang. 2025. Encoder-Only Image Registration.arXiv preprint arXiv:2509.00451(2025)
2025
-
[37]
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2022. Pali: A jointly-scaled multilingual language-image model.arXiv preprint arXiv:2209.06794(2022)
2022 arXiv
-
[38]
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. InEuropean conference on computer vision. Springer, 104–120
2020
-
[39]
Minsu Cho, Jungmin Lee, and Kyoung Mu Lee. 2010. Reweighted random walks for graph matching. InComputer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part V 11. Springer, 492–505
2010
-
[40]
Sungjoon Choi, Qian-Yi Zhou, and Vladlen Koltun. 2015. Robust reconstruction of indoor scenes. InProceedings of the IEEE conference on computer vision and pattern recognition. 5556–5565
2015
-
[41]
Christopher Choy, Jaesik Park, and Vladlen Koltun. 2019. Fully convolutional geometric features. InProceedings of the IEEE/CVF international conference on computer vision. 8958–8966
2019
-
[42]
Ondrej Chum and Jiri Matas. 2005. Matching with PROSAC-progressive sample consensus. In2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), Vol. 1. IEEE, 220–226
2005
-
[43]
Ondřej Chum, Jiří Matas, and Josef Kittler. 2003. Locally optimized RANSAC. InJoint pattern recognition symposium. Springer, 236–243
2003
-
[44]
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. 2017. Scannet: Richly-annotated 3d reconstructions of indoor scenes. InProceedings of the IEEE conference on computer vision and pattern recognition. 5828–5839. ACM Comput. Surv.,...
2017
-
[45]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in neural information processing systems36 (2023), 49250–49267
2023
-
[46]
Adrian V Dalca, Guha Balakrishnan, John Guttag, and Mert R Sabuncu. 2019. Unsupervised learning of probabilistic diffeomorphic registration for images and surfaces.Medical image analysis57 (2019), 226–236
2019
-
[47]
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra
-
[48]
Haowen Deng, Tolga Birdal, and Slobodan Ilic. 2018. Ppf-foldnet: Unsupervised learning of rotation invariant 3d local descriptors. InProceedings of the European conference on computer vision (ECCV). 602–618
2018
-
[49]
Haowen Deng, Tolga Birdal, and Slobodan Ilic. 2018. Ppfnet: Global context aware local features for robust 3d point matching. InProceedings of the IEEE conference on computer vision and pattern recognition. 195–205
2018
-
[50]
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li. 2021. Transvg: End-to-end visual grounding with transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision. 1769–1779
2021
-
[51]
Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. 2018. Superpoint: Self-supervised interest point detection and description. InProceedings of the IEEE conference on computer vision and pattern recognition workshops. 224–236
2018
-
[52]
Wangbin Ding, Lei Li, Xiahai Zhuang, and Liqin Huang. 2022. Cross-modality multi-atlas segmentation via deep registration and label fusion.IEEE Journal of Biomedical and Health Informatics26, 7 (2022), 3104–3115
2022
-
[53]
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. 2023. Palm-e: An embodied multimodal language model.arXiv preprint arXiv:2303.03378(2023)
2023 arXiv
-
[54]
Mihai Dusmanu, Ignacio Rocco, Tomas Pajdla, Marc Pollefeys, Josef Sivic, Akihiko Torii, and Torsten Sattler. 2019. D2-net: A trainable cnn for joint description and detection of local features. InProceedings of the ieee/cvf conference on computer vision and pattern recognition...
2019
-
[55]
Johan Edstedt, Ioannis Athanasiadis, Mårten Wadenbäck, and Michael Felsberg. 2023. DKM: Dense kernelized feature matching for geometry estimation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 17765–17775
2023
-
[56]
Mohamed S Elmahdy, Jelmer M Wolterink, Hessam Sokooti, Ivana Išgum, and Marius Staring. 2019. Adversarial optimization for joint registration and segmentation in prostate CT radiotherapy. InMedical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd Internatio...
2019
-
[57]
Koen AJ Eppenhof, Maxime W Lafarge, Mitko Veta, and Josien PW Pluim. 2019. Progressively trained convolutional neural networks for deformable image registration.IEEE transactions on medical imaging39, 5 (2019), 1594–1604
2019
-
[58]
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse++: Improving visual-semantic embeddings with hard negatives.arXiv preprint arXiv:1707.05612(2017)
2017 arXiv
-
[59]
Jingfan Fan, Xiaohuan Cao, Qian Wang, Pew-Thian Yap, and Dinggang Shen. 2019. Adversarial learning for mono-or multi-modal registration.Medical image analysis58 (2019), 101545
2019
-
[60]
Jingfan Fan, Xiaohuan Cao, Zhong Xue, Pew-Thian Yap, and Dinggang Shen. 2018. Adversarial similarity network for evaluating image alignment in deep learning based registration. InMedical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Confere...
2018
-
[61]
Philipp Fischer, Alexey Dosovitskiy, and Thomas Brox. 2014. Descriptor matching with convolutional neural networks: a comparison to sift.arXiv preprint arXiv:1405.5769(2014)
2014 arXiv
-
[62]
Martin A Fischler and Robert C Bolles. 1981. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography.Commun. ACM24, 6 (1981), 381–395
1981
-
[63]
Victor Fragoso, Pradeep Sen, Sergio Rodriguez, and Matthew Turk. 2013. EVSAC: accelerating hypotheses generation by modeling matching scores with extreme value theory. InProceedings of the IEEE international conference on computer vision. 2472–2479
2013
-
[64]
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg- Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018. Speaker-follower models for vision-and-language navigation.Advances in neural information processing syste...
2018
-
[65]
Andrea Frome, Daniel Huber, Ravi Kolluri, Thomas Bülow, and Jitendra Malik. 2004. Recognizing objects in range data using regional point descriptors. InComputer Vision-ECCV 2004: 8th European Conference on Computer Vision, Prague, Czech Republic, May 11-14, 2004. Proceedings, ...
2004
-
[66]
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016. Multimodal compact bilinear pooling for visual question answering and visual grounding.arXiv preprint arXiv:1606.01847(2016). ACM Comput. Surv., Vol. 1, No. 1, Article . Publicat...
2016 arXiv
-
[67]
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. 2013. Vision meets robotics: The kitti dataset. The international journal of robotics research32, 11 (2013), 1231–1237
2013
-
[68]
Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. 2022. Scaling open-vocabulary image segmentation with image-level labels. InEuropean conference on computer vision. Springer, 540–557
2022
-
[69]
Zan Gojcic, Caifa Zhou, Jan D Wegner, and Andreas Wieser. 2019. The perfect match: 3d point cloud matching with smoothed densities. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5545–5554
2019
-
[70]
Rui Gong, Weide Liu, Zaiwang Gu, Xulei Yang, and Jun Cheng. 2024. Learning Intra-view and Cross-view Geometric Knowledge for Stereo Matching. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
2024
-
[71]
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. InProceedings of the IEEE conference on computer vision and pattern recognition. 6904–6913
2017
-
[72]
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui. 2021. Open-vocabulary object detection via vision and language knowledge distillation.arXiv preprint arXiv:2104.13921(2021)
2021 arXiv
-
[73]
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Min Lu, and Jianwei Wan. 2013. Rotational projection statistics for 3D local surface description and object recognition.International journal of computer vision105 (2013), 63–86
2013
-
[74]
Yulan Guo, Ferdous Sohel, Mohammed Bennamoun, Jianwei Wan, and Min Lu. 2015. A novel local surface feature for 3D object recognition under clutter and occlusion.Information Sciences293 (2015), 196–213
2015
-
[75]
Agrim Gupta, Piotr Dollar, and Ross Girshick. 2019. Lvis: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5356–5364
2019
-
[76]
Kun Han, Shanlin Sun, Xiangyi Yan, Chenyu You, Hao Tang, Junayed Naushad, Haoyu Ma, Deying Kong, and Xiaohui Xie. 2023. Diffeomorphic image registration with neural velocity field. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1869–1879
2023
-
[77]
Xufeng Han, Thomas Leung, Yangqing Jia, Rahul Sukthankar, and Alexander C Berg. 2015. Matchnet: Unifying feature and metric learning for patch-based matching. InProceedings of the IEEE conference on computer vision and pattern recognition. 3279–3286
2015
-
[78]
Ankur Handa, Thomas Whelan, John McDonald, and Andrew J Davison. 2014. A benchmark for RGB-D visual odometry, 3D reconstruction and SLAM. In2014 IEEE international conference on Robotics and automation (ICRA). IEEE, 1524–1531
2014
-
[79]
Chris Harris, Mike Stephens, et al. 1988. A combined corner and edge detector. InAlvey vision conference, Vol. 15. Citeseer, 10–5244
1988
-
[80]
Kun He, Yan Lu, and Stan Sclaroff. 2018. Local descriptors optimized for average precision. InProceedings of the IEEE conference on computer vision and pattern recognition. 596–605
2018
-
[81]
Mattias P Heinrich, Mark Jenkinson, Manav Bhushan, Tahreema Matin, Fergus V Gleeson, Michael Brady, and Julia A Schnabel. 2012. MIND: Modality independent neighbourhood descriptor for multi-modal deformable registration. Medical image analysis16, 7 (2012), 1423–1435
2012
-
[82]
Alessa Hering, Lasse Hansen, Tony CW Mok, Albert CS Chung, Hanna Siebert, Stephanie Häger, Annkristin Lange, Sven Kuckertz, Stefan Heldmann, Wei Shao, et al . 2022. Learn2Reg: comprehensive multi-task medical image registration challenge, dataset and evaluation in the era of d...
2022
-
[83]
Derek LG Hill, Philipp G Batchelor, Mark Holden, and David J Hawkes. 2001. Medical image registration.Physics in medicine & biology46, 3 (2001), R1
2001
-
[84]
Malte Hoffmann, Benjamin Billot, Douglas N Greve, Juan Eugenio Iglesias, Bruce Fischl, and Adrian V Dalca. 2021. SynthMorph: learning contrast-invariant registration without acquired images.IEEE transactions on medical imaging 41, 3 (2021), 543–558
2021
-
[85]
Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. 2016. Segmentation from natural language expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14. Springer, 108–124
2016
-
[86]
Qian Huang, Xiaotong Guo, Yiming Wang, Huashan Sun, and Lijie Yang. 2024. A survey of feature matching methods. IET Image Processing18, 6 (2024), 1385–1410
2024
-
[87]
Renlang Huang, Yufan Tang, Jiming Chen, and Liang Li. 2024. A consistency-aware spot-guided transformer for versatile and hierarchical point cloud registration.Advances in Neural Information Processing Systems37 (2024), 70230–70258
2024
-
[88]
Shengyu Huang, Zan Gojcic, Mikhail Usvyatsov, Andreas Wieser, and Konrad Schindler. 2021. Predator: Registration of 3d point clouds with low overlap. InProceedings of the IEEE/CVF Conference on computer vision and pattern recognition. 4267–4276
2021
-
[89]
Drew A Hudson and Christopher D Manning. 2019. Gqa: A new dataset for real-world visual reasoning and compositional question answering. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication ...
2019
-
[90]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning. PMLR, 4904–4916
2021
-
[91]
Haobo Jiang, Mathieu Salzmann, Zheng Dang, Jin Xie, and Jian Yang. 2023. Se (3) diffusion model-based point cloud registration for robust 6d object pose estimation.Advances in Neural Information Processing Systems36 (2023), 21285–21297
2023
-
[92]
Jue Jiang and Harini Veeraraghavan. 2022. One shot PACS: Patient specific Anatomic Context and Shape prior aware recurrent registration-segmentation of longitudinal thoracic cone beam CTs.IEEE transactions on medical imaging41, 8 (2022), 2021–2032
2022
-
[93]
Wei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi, and Kwang Moo Yi. 2021. Cotr: Correspondence transformer for matching across images. InProceedings of the IEEE/CVF international conference on computer vision. 6207–6217
2021
-
[94]
Yuhe Jin, Dmytro Mishkin, Anastasiia Mishchuk, Jiri Matas, Pascal Fua, Kwang Moo Yi, and Eduard Trulls. 2021. Image matching across wide baselines: From paper to practice.International Journal of Computer Vision129, 2 (2021), 517–547
2021
-
[95]
Andrew E Johnson and Martial Hebert. 1999. Using spin images for efficient object recognition in cluttered 3D scenes. IEEE Transactions on pattern analysis and machine intelligence21, 5 (1999), 433–449
1999
-
[97]
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. InProceedings of the IEEE conference on computer vision and pattern recog...
2017
-
[98]
Timor Kadir and Michael Brady. 2001. Saliency, scale and image description.International Journal of Computer Vision 45 (2001), 83–105
2001
-
[99]
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. 2021. Mdetr- modulated detection for end-to-end multi-modal understanding. InProceedings of the IEEE/CVF international conference on computer vision. 1780–1790
2021
-
[100]
Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 3128–3137
2015
-
[101]
Yan Ke and Rahul Sukthankar. 2004. PCA-SIFT: A more distinctive representation for local image descriptors. In Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004., Vol. 2. IEEE, II–II
2004
-
[102]
Marc Khoury, Qian-Yi Zhou, and Vladlen Koltun. 2017. Learning compact geometric features. InProceedings of the IEEE international conference on computer vision. 153–161
2017
-
[103]
Boah Kim, Inhwa Han, and Jong Chul Ye. 2022. Diffusemorph: Unsupervised deformable image registration using diffusion model. InEuropean conference on computer vision. Springer, 347–364
2022
-
[104]
Dahun Kim, Anelia Angelova, and Weicheng Kuo. 2023. Region-aware pretraining for open-vocabulary object detection with vision transformers. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11144–11154
2023
-
[105]
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018. Bilinear attention networks.Advances in neural information processing systems31 (2018)
2018
-
[106]
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. InInternational conference on machine learning. PMLR, 5583–5594
2021
-
[107]
Stefan Klein, Marius Staring, Keelin Murphy, Max A Viergever, and Josien PW Pluim. 2009. Elastix: a toolbox for intensity-based medical image registration.IEEE transactions on medical imaging29, 1 (2009), 196–205
2009
-
[108]
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, et al . 2017. Ai2-thor: An interactive 3d environment for visual ai.arXiv preprint arXiv:1712.05474(2017)
2017 arXiv
-
[109]
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. 2020. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding.arXiv preprint arXiv:2010.07954(2020)
2020 arXiv
-
[110]
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling. 2009. Learning to detect unseen object classes by between-class attribute transfer. In2009 IEEE conference on computer vision and pattern recognition. IEEE, 951–958
2009
-
[111]
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. 2018. Stacked cross attention for image-text matching. InProceedings of the European conference on computer vision (ECCV). 201–216. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: June 2026. 38 Liu et al
2018
-
[112]
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is more: Clipbert for video-and-language learning via sparse sampling. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 7331–7341
2021
-
[113]
Marius Leordeanu and Martial Hebert. 2005. A spectral technique for correspondence problems using pairwise constraints. InTenth IEEE International Conference on Computer Vision (ICCV’05) Volume 1, Vol. 2. IEEE, 1482–1489
2005
-
[114]
Vincent Leroy, Yohann Cabon, and Jérôme Revaud. 2024. Grounding image matching in 3d with mast3r. InEuropean Conference on Computer Vision. Springer, 71–91
2024
-
[115]
Stefan Leutenegger, Margarita Chli, and Roland Y Siegwart. 2011. BRISK: Binary robust invariant scalable keypoints. In2011 International conference on computer vision. Ieee, 2548–2555
2011
-
[116]
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and René Ranftl. 2022. Language-driven semantic segmentation.arXiv preprint arXiv:2201.03546(2022)
2022 arXiv
-
[117]
Jiaxin Li and Gim Hee Lee. 2019. Usip: Unsupervised stable interest point detection from 3d point clouds. InProceedings of the IEEE/CVF international conference on computer vision. 361–370
2019
-
[118]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. (2023), 19730–19742
2023
-
[119]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InInternational conference on machine learning. PMLR, 12888–12900
2022
-
[120]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation.Advances in neural information processing systems34 (2021), 9694–9705
2021
-
[121]
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. 2019. Visual semantic reasoning for image-text matching. InProceedings of the IEEE/CVF international conference on computer vision. 4654–4662
2019
-
[122]
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. 2022. Grounded language-image pre-training. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 1...
2022
-
[123]
Muchen Li and Leonid Sigal. 2021. Referring transformer: A one-step approach to multi-task visual grounding. Advances in neural information processing systems34, 19652–19664
2021
-
[124]
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. InComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 2...
2020
-
[125]
Lawrence Zitnick
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. InEuropean Conference on Computer Vision (ECCV). 740–755
2014
-
[126]
Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Pollefeys. 2023. Lightglue: Local feature matching at light speed. InProceedings of the IEEE/CVF International Conference on Computer Vision. 17627–17638
2023
-
[127]
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. 2024. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European conference on computer vision. Springer, 38–55
2024
-
[128]
Weide Liu, Jieming Lou, Xingxing Wang, Wei Zhou, Jun Cheng, and Xulei Yang. 2025. Physically-guided open vocabulary segmentation with weighted patched alignment loss.Neurocomputing614 (2025), 128788
2025
-
[129]
Weide Liu, Chi Zhang, Henghui Ding, Tzu-Yi Hung, and Guosheng Lin. 2022. Few-shot segmentation with optimal transport matching and message flow.IEEE Transactions on Multimedia25 (2022), 5130–5141
2022
-
[130]
Weide Liu, Chi Zhang, Guosheng Lin, and Fayao Liu. 2020. Crnet: Cross-reference networks for few-shot segmentation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4165–4173
2020
-
[131]
Yang Liu, Wentao Feng, Zhuoyao Liu, Shudong Huang, and Jiancheng Lv. 2025. Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching.arXiv preprint arXiv:2503.14953(2025)
2025 arXiv
-
[132]
Tsz-Wai Rachel Lo and J Paul Siebert. 2009. Local feature extraction and matching on range images: 2.5 D SIFT. Computer Vision and Image Understanding113, 12 (2009), 1235–1250
2009
-
[133]
Zijun Long, George Killick, Richard McCreadie, and Gerardo Aragon Camarasa. 2024. Multiway-adapter: Adapting multimodal large language models for scalable image-text retrieval. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)...
2024
-
[134]
Jieming Lou, Weide Liu, Zhuo Chen, Fayao Liu, and Jun Cheng. 2023. ELFNet: Evidential Local-global Fusion for Stereo Matching. InProceedings of the IEEE/CVF Conference on International Conference on Computer Vision
2023
-
[135]
David G Lowe. 2004. Distinctive image features from scale-invariant keypoints.International journal of computer vision60 (2004), 91–110. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: June 2026. Modality-Aware Feature Matching in Visual and Vision-Language Appli...
2004
-
[136]
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, et al. 2024. Deepseek-vl: towards real-world vision-language understanding.arXiv preprint arXiv:2403.05525 (2024)
2024 arXiv
-
[137]
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic represen- tations for vision-and-language tasks.Advances in neural information processing systems32
2019
-
[138]
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical question-image co-attention for visual question answering.Advances in neural information processing systems29
2016
-
[139]
Timo Lüddecke and Alexander Ecker. 2022. Image segmentation using text and image prompts. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 7086–7096
2022
-
[140]
Zixin Luo, Tianwei Shen, Lei Zhou, Jiahui Zhang, Yao Yao, Shiwei Li, Tian Fang, and Long Quan. 2019. Contextdesc: Local descriptor augmentation with cross-modality context. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2527–2536
2019
-
[141]
Zixin Luo, Tianwei Shen, Lei Zhou, Siyu Zhu, Runze Zhang, Yao Yao, Tian Fang, and Long Quan. 2018. Geodesc: Learning local descriptors by integrating geometry constraints. InProceedings of the European conference on computer vision (ECCV). 168–183
2018
-
[142]
Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. 2019. The regretful agent: Heuristic-aided navigation through progress estimation. InProceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition. 6732–6740
2019
-
[143]
Jiayi Ma, Xingyu Jiang, Aoxiang Fan, Junjun Jiang, and Junchi Yan. 2021. Image matching from handcrafted to deep features: A survey.International Journal of Computer Vision129, 1 (2021), 23–79
2021
-
[144]
Frederik Maes, Andre Collignon, Dirk Vandermeulen, Guy Marchal, and Paul Suetens. 1997. Multimodality image registration by maximization of mutual information.IEEE transactions on Medical Imaging16, 2 (1997), 187–198
1997
-
[145]
Dwarikanath Mahapatra and Zongyuan Ge. 2020. Training data independent image registration using generative adversarial networks and domain adaptation.Pattern Recognition100 (2020), 107109
2020
-
[146]
Dwarikanath Mahapatra, Zongyuan Ge, Suman Sedai, and Rajib Chakravorty. 2018. Joint registration and segmentation of xray images using generative adversarial networks. InMachine Learning in Medical Imaging: 9th International Workshop, MLMI 2018, Held in Conjunction with MICCAI...
2018
-
[147]
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. InProceedings of the IEEE conference on computer vision and pattern recognition. 11–20
2016
-
[148]
Jiri Matas, Ondrej Chum, Martin Urban, and Tomás Pajdla. 2004. Robust wide-baseline stereo from maximally stable extremal regions.Image and vision computing22, 10 (2004), 761–767
2004
-
[149]
Bjoern H. Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, Levente Lanczi, Elizabeth Gerstner, Marc-Andre Weber, Tal Arbel, Brian B. Avants, Nicholas Ayache, Patricia Buend...
-
[150]
Krystian Mikolajczyk and Cordelia Schmid. 2005. A performance evaluation of local descriptors.IEEE transactions on pattern analysis and machine intelligence27, 10 (2005), 1615–1630
2005
-
[151]
Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, Frederik Schaffalitzky, Timor Kadir, and L Van Gool. 2005. A comparison of affine region detectors.International journal of computer vision 65 (2005), 43–72
2005
-
[152]
The Multimodal Brain Tumor Image Segmentation Benchmark (BRATS).IEEE Transactions on Medical Imaging 34, 10 (2015), 1993–2024
2015
-
[153]
Matthias Minderer, Alexey Gritsenko, and Neil Houlsby. 2023. Scaling open-vocabulary object detection.Advances in Neural Information Processing Systems36 (2023), 72983–73007
2023
-
[154]
Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, et al. 2022. Simple open-vocabulary object detection. InEuropean conference on computer vision. Springer, 7...
2022
-
[155]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space.arXiv preprint arXiv:1301.3781(2013)
2013 arXiv
-
[156]
Tony CW Mok and Albert CS Chung. 2020. Large deformation diffeomorphic image registration with laplacian pyramid networks. InMedical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Proceedings, Part ...
2020
-
[157]
Kai Ni, Hailin Jin, and Frank Dellaert. 2009. GroupSAC: Efficient consensus in the presence of groupings. In2009 IEEE 12th International Conference on Computer Vision. IEEE, 2193–2200
2009
-
[158]
Anastasiia Mishchuk, Dmytro Mishkin, Filip Radenovic, and Jiri Matas. 2017. Working hard to know your neighbor’s margins: Local descriptor learning loss.Advances in neural information processing systems30 (2017)
2017
-
[159]
Yuki Ono, Eduard Trulls, Pascal Fua, and Kwang Moo Yi. 2018. LF-Net: Learning local features from images.Advances in neural information processing systems31 (2018)
2018
-
[160]
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. InProceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 1532–1543
2014
-
[161]
Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S Cor- rado, and Jeffrey Dean. 2013. Zero-shot learning by convex combination of semantic embeddings.arXiv preprint arXiv:1312.5650(2013)
2013 arXiv
-
[162]
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. InProceedings of the IEEE conference on computer vision and pattern recognition. 652–660
2017
-
[163]
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. 2020. Reverie: Remote embodied visual referring expression in real indoor environments. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. ...
2020
-
[164]
François Pomerleau, Ming Liu, Francis Colas, and Roland Siegwart. 2012. Challenging data sets for point cloud registration algorithms.The International Journal of Robotics Research31, 14 (2012), 1705–1711
2012
-
[165]
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, et al. 2024. Grounded sam: Assembling open-world models for diverse visual tasks.arXiv preprint arXiv:2401.14159 (2024)
2024 arXiv
-
[166]
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. InProceedings of the IEEE conference on computer vision and pattern recognition. 7008–7024
2017
-
[167]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...
2021
-
[168]
Ignacio Rocco, Mircea Cimpoi, Relja Arandjelović, Akihiko Torii, Tomas Pajdla, and Josef Sivic. 2018. Neighbourhood consensus networks.Advances in neural information processing systems31 (2018)
2018
-
[169]
Edward Rosten and Tom Drummond. 2006. Machine learning for high-speed corner detection. InComputer Vision– ECCV 2006: 9th European Conference on Computer Vision, Graz, Austria, May 7-13, 2006. Proceedings, Part I 9. Springer, 430–443
2006
-
[170]
Jerome Revaud, Cesar De Souza, Martin Humenberger, and Philippe Weinzaepfel. 2019. R2d2: Reliable and repeatable detector and descriptor.Advances in neural information processing systems32 (2019)
2019
-
[171]
Radu Bogdan Rusu, Nico Blodow, and Michael Beetz. 2009. Fast point feature histograms (FPFH) for 3D registration. In2009 IEEE international conference on robotics and automation. IEEE, 3212–3217
2009
-
[172]
Radu Bogdan Rusu, Nico Blodow, Zoltan Csaba Marton, and Michael Beetz. 2008. Aligning point cloud views using persistent feature histograms. In2008 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 3384–3391
2008
-
[173]
Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. 2011. ORB: An efficient alternative to SIFT or SURF. In2011 International conference on computer vision. Ieee, 2564–2571
2011
-
[174]
Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, et al. 2018. Benchmarking 6dof outdoor visual localization in changing conditions. InProceedings of the IEEE conference on ...
2018
-
[175]
Cordelia Schmid, Roger Mohr, and Christian Bauckhage. 2000. Evaluation of interest point detectors.International Journal of computer vision37, 2 (2000), 151–172
2000
-
[176]
Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. 2020. Superglue: Learning feature matching with graph neural networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4938–4947
2020
-
[177]
Xuelun Shen, Zhipeng Cai, Wei Yin, Matthias Müller, Zijun Li, Kaixuan Wang, Xiaozhi Chen, and Cheng Wang. 2024. Gim: Learning generalizable image matcher from internet videos.arXiv preprint arXiv:2402.11095(2024)
2024 arXiv
-
[178]
Jianbo Shi et al. 1994. Good features to track. In1994 Proceedings of IEEE conference on computer vision and pattern recognition. IEEE, 593–600
1994
-
[179]
Thomas Schops, Johannes L Schonberger, Silvano Galliani, Torsten Sattler, Konrad Schindler, Marc Pollefeys, and Andreas Geiger. 2017. A multi-view stereo benchmark with high-resolution images and multi-camera videos. In Proceedings of the IEEE conference on computer vision and...
2017
-
[180]
Jamie Shotton, Ben Glocker, Christopher Zach, Shahram Izadi, Antonio Criminisi, and Andrew Fitzgibbon. 2013. Scene coordinate regression forests for camera relocalization in RGB-D images. InProceedings of the IEEE conference on computer vision and pattern recognition. 2930–2937
2013
-
[181]
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. InProceedings of the IEEE/CVF conference on computer vision and pat...
2020
-
[182]
Jiacheng Shi, Yuting He, Youyong Kong, Jean-Louis Coatrieux, Huazhong Shu, Guanyu Yang, and Shuo Li. 2022. Xmorpher: Full transformer for deformable medical image registration via cross attention. InInternational Conference on Medical Image Computing and Computer-Assisted Inte...
2022
-
[183]
Martin Simonovsky, Benjamín Gutiérrez-Becker, Diana Mateus, Nassir Navab, and Nikos Komodakis. 2016. A deep metric for multimodal registration. InMedical Image Computing and Computer-Assisted Intervention-MICCAI 2016: 19th International Conference, Athens, Greece, October 17-2...
2016
-
[184]
Ivan Sipiran and Benjamin Bustos. 2011. Harris 3D: a robust extension of the Harris operator for interest point detection on 3D meshes.The Visual Computer27 (2011), 963–976
2011
-
[185]
Edgar Simo-Serra, Eduard Trulls, Luis Ferraz, Iasonas Kokkinos, Pascal Fua, and Francesc Moreno-Noguer. 2015. Discriminative learning of deep convolutional feature point descriptors. InProceedings of the IEEE international conference on computer vision. 118–126
2015
-
[186]
Richard Socher, Milind Ganjoo, Christopher D Manning, and Andrew Ng. 2013. Zero-shot learning through cross- modal transfer.Advances in neural information processing systems26 (2013)
2013
-
[187]
Hessam Sokooti, Bob De Vos, Floris Berendsen, Boudewijn PF Lelieveldt, Ivana Išgum, and Marius Staring. 2017. Nonrigid image registration using multi-scale 3D convolutional neural networks. InMedical Image Computing and Computer Assisted Intervention- MICCAI 2017: 20th Interna...
2017
-
[188]
Stephen M Smith and J Michael Brady. 1997. SUSAN—a new approach to low level image processing.International journal of computer vision23, 1 (1997), 45–78
1997
-
[189]
Xinrui Song, Hanqing Chao, Xuanang Xu, Hengtao Guo, Sheng Xu, Baris Turkbey, Bradford J Wood, Thomas Sanford, Ge Wang, and Pingkun Yan. 2022. Cross-modal attention for multi-modal image registration.Medical Image Analysis 82 (2022), 102612
2022
-
[190]
Christoph Strecha, Wolfgang Von Hansen, Luc Van Gool, Pascal Fua, and Ulrich Thoennessen. 2008. On benchmarking camera calibration and multi-view stereo for high resolution imagery. In2008 IEEE conference on computer vision and pattern recognition. Ieee, 1–8
2008
-
[191]
Edward J Somer, Paul K Marsden, Nigel A Benatar, Joanne Goodey, Michael J O’Doherty, and Michael A Smith. 2003. PET-MR image fusion in soft tissue sarcoma: accuracy, reliability and practicality of interactive point-based and automated mutual information techniques.European jo...
2003
-
[192]
Jürgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. 2012. A benchmark for the evaluation of RGB-D SLAM systems. In2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 573–580
2012
-
[193]
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. 2021. LoFTR: Detector-free local feature matching with transformers. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 8922–8931
2021
-
[194]
Colin Studholme, Derek LG Hill, and David J Hawkes. 1999. An overlap invariant entropy measure of 3D medical image alignment.Pattern recognition32, 1 (1999), 71–86
1999
-
[195]
Shitao Tang, Jiahui Zhang, Siyu Zhu, and Ping Tan. 2022. Quadtree attention for vision transformers.arXiv preprint arXiv:2201.02767(2022)
2022 arXiv
-
[196]
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. 2020. Vision-and-dialog navigation. In Conference on Robot Learning. PMLR, 394–406
2020
-
[197]
Hao Tan and Mohit Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers.arXiv preprint arXiv:1908.07490(2019)
2019 arXiv
-
[198]
Yurun Tian, Bin Fan, and Fuchao Wu. 2017. L2-net: Deep learning of discriminative patch descriptor in euclidean space. InProceedings of the IEEE conference on computer vision and pattern recognition. 661–669. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: June 2...
2017
-
[199]
Engin Tola, Vincent Lepetit, and Pascal Fua. 2009. Daisy: An efficient dense descriptor applied to wide-baseline stereo. IEEE transactions on pattern analysis and machine intelligence32, 5 (2009), 815–830
2009
-
[200]
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. 2016. Yfcc100m: The new data in multimedia research.Commun. ACM59, 2 (2016), 64–73
2016
-
[201]
Federico Tombari, Samuele Salti, and Luigi Di Stefano. 2010. Unique signatures of histograms for local surface description. InComputer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part III 11. Sprin...
2010
-
[202]
Tomasz Trzcinski, Mario Christoudias, and Vincent Lepetit. 2014. Learning image descriptors with boosting.IEEE transactions on pattern analysis and machine intelligence37, 3 (2014), 597–610
2014
-
[203]
Federico Tombari, Samuele Salti, and Luigi Di Stefano. 2010. Unique shape context for 3D data description. In Proceedings of the ACM workshop on 3D object retrieval. 57–62
2010
-
[204]
Yannick Verdie, Kwang Yi, Pascal Fua, and Vincent Lepetit. 2015. Tilde: A temporally invariant learned detector. In Proceedings of the IEEE conference on computer vision and pattern recognition. 5279–5288
2015
-
[205]
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. InProceedings of the IEEE conference on computer vision and pattern recognition. 3156–3164
2015
-
[206]
Tinne Tuytelaars and Luc Van Gool. 2004. Matching widely separated views based on affine invariant regions. International journal of computer vision59 (2004), 61–85
2004
-
[207]
Paul Viola and William M Wells III. 1997. Alignment by maximization of mutual information.International journal of computer vision24, 2 (1997), 137–154
1997
-
[208]
Ching-Wei Wang, Yu-Ching Lee, Muhammad-Adil Khalil, Kuan-Yu Lin, Cheng-Ping Yu, and Huang-Chun Lien. 2022. Fast cross-staining alignment of gigapixel whole slide images with application to prostate cancer and breast cancer analysis.Scientific Reports12, 1 (2022), 11623
2022
-
[209]
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2016. Show and tell: Lessons learned from the 2015 mscoco image captioning challenge.IEEE transactions on pattern analysis and machine intelligence39, 4 (2016), 652–663
2016
-
[210]
Qing Wang, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. 2022. Matchformer: Interleaving attention in transformers for feature matching. InProceedings of the Asian Conference on Computer Vision. 2746–2762
2022
-
[211]
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. 2019. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. InProceedings of the IEEE/CVF conference on com...
2019
-
[212]
Haiping Wang, Yuan Liu, Bing Wang, Yujing Sun, Zhen Dong, Wenping Wang, and Bisheng Yang. 2023. Freereg: Image-to-point cloud registration leveraging pretrained diffusion models and monocular depth estimators.arXiv preprint arXiv:2310.03420(2023)
2023 arXiv
-
[213]
Yu Wang, Xiaoye Wang, Zaiwang Gu, Weide Liu, Wee Siong Ng, Weimin Huang, and Jun Cheng. 2024. Superjunction: Learning-based junction detection for retinal image registration. InProceedings of the AAAI conference on artificial intelligence, Vol. 38. 292–300
2024
-
[214]
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2021. Simvlm: Simple visual language model pretraining with weak supervision.arXiv preprint arXiv:2108.10904(2021)
2021 arXiv
-
[215]
Yifan Wang, Xingyi He, Sida Peng, Dongli Tan, and Xiaowei Zhou. 2024. Efficient LoFTR: Semi-dense local feature matching with sparse-like speed. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 21666–21675
2024
-
[216]
Junshen Xu, Eric Z Chen, Xiao Chen, Terrence Chen, and Shanhui Sun. 2021. Multi-scale neural odes for 3d medical image registration. InMedical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–Octobe...
2021
-
[217]
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. 2022. Groupvit: Semantic segmentation emerges from text supervision. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 18134–18144
2022
-
[218]
Jay West, J Michael Fitzpatrick, Matthew Y Wang, Benoit M Dawant, Calvin R Maurer Jr, Robert M Kessler, Robert J Maciunas, Christian Barillot, Didier Lemoine, Andre Collignon, et al. 1997. Comparison and evaluation of retrospective intermodality brain image registration techni...
1997
-
[219]
Shibiao Xu, Shunpeng Chen, Rongtao Xu, Changwei Wang, Peng Lu, and Li Guo. 2024. Local feature matching using deep learning: A survey.Information Fusion107 (2024), 102344
2024
-
[220]
Jiebin Yan, Jiale Rao, Xuelin Liu, Yuming Fang, Yifan Zuo, and Weide Liu. 2025. Subjective and objective quality assessment of non-uniformly distorted omnidirectional images.IEEE Transactions on Multimedia27 (2025), 2695–2707
2025
-
[221]
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. InInternational conference on machine learning. PMLR, 2048–2057
2015
-
[222]
Qianye Yang, David Atkinson, Yunguan Fu, Tom Syer, Wen Yan, Shonit Punwani, Matthew J Clarkson, Dean C Barratt, Tom Vercauteren, and Yipeng Hu. 2022. Cross-modality image registration using a training-time privileged third modality.IEEE Transactions on Medical Imaging41, 11 (2...
2022
-
[223]
Sibei Yang, Guanbin Li, and Yizhou Yu. 2019. Dynamic graph attention for referring expression comprehension. In Proceedings of the IEEE/CVF international conference on computer vision. 4644–4653
2019
-
[224]
Jiaqi Yang, Qian Zhang, Yang Xiao, and Zhiguo Cao. 2017. TOLDI: An effective and robust approach for 3D local shape description.Pattern Recognition65 (2017), 175–187. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: June 2026. Modality-Aware Feature Matching in Vi...
2017
-
[225]
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016. Stacked attention networks for image question answering. InProceedings of the IEEE conference on computer vision and pattern recognition. 21–29
2016
-
[226]
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu. 2021. Filip: Fine-grained interactive language-image pre-training.arXiv preprint arXiv:2111.07783 (2021)
2021 arXiv
-
[227]
Xiao Yang, Roland Kwitt, Martin Styner, and Marc Niethammer. 2017. Quicksilver: Fast predictive image registration–a deep learning approach.NeuroImage158 (2017), 378–396
2017
-
[228]
Kwang Moo Yi, Eduard Trulls, Vincent Lepetit, and Pascal Fua. 2016. Lift: Learned invariant feature transform. In European conference on computer vision. Springer, 467–483
2016
-
[229]
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.Transactions of the association for computational linguistics2 (2014), 67–78
2014
-
[230]
Zi Jian Yew and Gim Hee Lee. 2018. 3dfeat-net: Weakly supervised local 3d features for point cloud registration. In Proceedings of the European conference on computer vision (ECCV). 607–623
2018
-
[231]
Hao Yu, Zheng Qin, Ji Hou, Mahdi Saleh, Dongsheng Li, Benjamin Busam, and Slobodan Ilic. 2023. Rotation-invariant transformer for point cloud matching. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5384–5393
2023
-
[232]
Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. 2018. Mattnet: Modular attention network for referring expression comprehension. InProceedings of the IEEE conference on computer vision and pattern recognition. 1307–1315
2018
-
[233]
Hao Yu, Fu Li, Mahdi Saleh, Benjamin Busam, and Slobodan Ilic. 2021. Cofinet: Reliable coarse-to-fine correspondences for robust pointcloud registration.Advances in Neural Information Processing Systems34 (2021), 23872–23884
2021
-
[234]
Paul A Yushkevich, Yang Gao, and Guido Gerig. 2016. ITK-SNAP: An interactive tool for semi-automatic segmentation of multi-modality biomedical images. In2016 38th annual international conference of the IEEE engineering in medicine and biology society (EMBC). IEEE, 3342–3345
2016
-
[235]
Sergey Zagoruyko and Nikos Komodakis. 2015. Learning to compare image patches via convolutional neural networks. InProceedings of the IEEE conference on computer vision and pattern recognition. 4353–4361
2015
-
[236]
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg. 2017. A joint speaker-listener-reinforcer model for referring expressions. InProceedings of the IEEE conference on computer vision and pattern recognition. 7282–7290
2017
-
[237]
Andy Zeng, Shuran Song, Matthias Nießner, Matthew Fisher, Jianxiong Xiao, and Thomas Funkhouser. 2017. 3dmatch: Learning local geometric descriptors from rgb-d reconstructions. InProceedings of the IEEE conference on computer vision and pattern recognition. 1802–1811
2017
-
[238]
Jingbo Zeng, Zaiwang Gu, Weide Liu, Lile Cai, and Jun Cheng. 2025. Uncertainty Aware Interest Point Detection and Description.. InW ACV. 2144–2153
2025
-
[239]
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 6720–6731
2019
-
[240]
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer
-
[241]
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao
-
[242]
Qi Zeng, Weide Liu, Bo Li, Ryne Didier, P Ellen Grant, and Davood Karimi. 2025. Towards automatic US-MR fetal brain image registration with learning-based methods.NeuroImage310 (2025), 121104
2025
-
[243]
Yuxi Zhang, Xiang Chen, Jiazheng Wang, Min Liu, Yaonan Wang, Dongdong Liu, Renjiu Hu, and Hang Zhang. 2024. Large Scale Unsupervised Brain MRI Image Registration Solution for Learn2Reg 2024.arXiv preprint arXiv:2409.00917 (2024)
2024 arXiv
-
[244]
Dora Zhao, Angelina Wang, and Olga Russakovsky. 2021. Understanding and evaluating racial biases in image captioning. InProceedings of the IEEE/CVF international conference on computer vision. 14830–14840. ACM Comput. Surv., Vol. 1, No. 1, Article . Publication date: June 2026...
2021
-
[245]
Guiyu Zhao, Zhentao Guo, Zewen Du, and Hongbin Ma. 2025. Cross-PCR: A Robust Cross-Source Point Cloud Registration Framework. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 10403–10411
2025
-
[246]
Shengyu Zhao, Tingfung Lau, Ji Luo, Eric I-Chao Chang, and Yan Xu. 2019. Unsupervised 3D end-to-end medical image registration with volume tweening network.IEEE journal of biomedical and health informatics24, 5 (2019), 1394–1404
2019
-
[247]
Xingwu Zhang, Guanxuan Li, Zhuocheng Zhang, and Zijun Long. 2025. RoboEye: Enhancing 2D Robotic Object Identification with Selective 3D Geometric Keypoint Matching.arXiv preprint arXiv:2509.14966(2025)
2025
-
[248]
Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al. 2022. Regionclip: Region-based language-image pretraining. InProceedings of the IEEE/CVF conference on computer vision and pattern recognit...
2022
-
[249]
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. 2022. Learning to prompt for vision-language models.International Journal of Computer Vision130, 9 (2022), 2337–2348
2022
-
[250]
Xiahai Zhuang, Lei Li, Christian Payer, Darko Štern, Martin Urschler, Mattias P Heinrich, Julien Oster, Chunliang Wang, Örjan Smedby, Cheng Bian, et al. 2019. Evaluation of algorithms for multi-modality whole heart segmentation: an open-access grand challenge.Medical image ana...
2019
-
[252]
Yu Zhong. 2009. Intrinsic shape signatures: A shape descriptor for 3D object recognition. In2009 IEEE 12th international conference on computer vision workshops, ICCV Workshops. IEEE, 689–696
2009
-
[2015]
Microsoft coco captions: Data collection and evaluation server.arXiv preprint arXiv:1504.00325(2015)
2015 arXiv
-
[2017]
InProceedings of the IEEE conference on computer vision and pattern recognition
Visual dialog. InProceedings of the IEEE conference on computer vision and pattern recognition. 326–335
-
[2021]
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
Vinvl: Revisiting visual representations in vision-language models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5579–5588
-
[2022]
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
Lit: Zero-shot transfer with locked-image text tuning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 18123–18133
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.