REVIEW 3 major objections 4 minor 1 cited by
Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Auxiliary information is what moves the needle in text-to-video retrieval
desk verdict A thorough survey with a genuinely useful taxonomy of auxiliary information in T2V retrieval, but its comparative conclusions outrun the evidence — no no-auxiliary baseline, and class membership correlates with model scale. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The taxonomy of auxiliary information is the paper's organizing device. It splits every method by the source of its extra signal—video-extracted or text-extracted—and then by granularity (video-level versus region-level, sentence-level versus word-level) and by whether the signal is observed, from another modality, or generated by a vision-language or large language model. The comparative tables built on this taxonomy, with results separated into zero-shot and fine-tuned settings, are what carry the argument that some auxiliary classes work better than others.
What would settle it
Take one pre-training corpus and one video/text backbone, then train four copies of the model that differ only in auxiliary information: no extra signal, generated video captions (GVC), generated multimodal captions (GMC), and text refinement (GTC). If the GVC/GMC copies do not beat the text-refinement copy on the same benchmark, the paper's central comparative finding would fail; if the no-extra-signal copy matches the others, the whole premise that auxiliary information helps would be in doubt.
Extended reading notes
Core claim
The paper claims that the semantic gap between video and text is best bridged by auxiliary information rather than by stronger alignment of the two raw modalities alone. It classifies 81 methods: video-extracted information divides into visual info (video-level expert features and temporal context; region-level object boxes and spatial-temporal dependencies), modality info (video tags, audio, and extra sensors), and generated info (generated video captions, generated multimodal captions, and question-answer systems); text-extracted information divides into textual info (temporal dependencies, multiple captions, text experts, part-of-speech tags, POS-derived captions, hierarchy dependencies) and generated info (generated text captions and multilingual translations). Comparing Recall@1, nDCG, and mAP across MSR-VTT, MSVD, ActivityNet Captions, DiDeMo, LSMDC, YouCook2, VATEX, and EPIC-KITCHENS-100, the paper finds that Q&A-based methods lead zero-shot retrieval, that generated video captions and multimodal alignment usually beat text-only refinement, that audio matters on domain-specific datasets such as cooking and movies, and that datasets with multiple captions per clip especially reward text-side methods.
Load-bearing premise
The comparison assumes recall numbers reported by different papers are directly comparable once split into zero-shot and fine-tuned, even though the methods differ in pre-training data, model size, and training budget.
Editorial extensions
If this is right
- Retrieval systems should treat auxiliary signals as part of the input design, not as an optional trick: generated captions and multimodal alignment are the classes that most often top the leaderboards.
- Zero-shot text-to-video retrieval can be led by interactive question-answer pipelines, which refine the query before matching.
- Datasets that supply multiple captions per clip or multilingual captions unlock the strongest text-side methods, so data collection should prioritize caption multiplicity.
- Audio is not a marginal extra: on domain-specific datasets like cooking and movies it is frequently the difference between top and mid-tier retrieval.
- For untrimmed videos, temporal context from neighbouring clips or object tracks is a reliable source of gain.
Reading between the lines
- The paper's comparisons do not hold model scale, pre-training corpus, or training budget fixed; a fair reading is that the ranking of auxiliary classes is entangled with the size of the models that use them. A controlled study with one backbone and one pre-training set, varying only the auxiliary signal, would isolate the true contribution of each class.
- The success of generated captions suggests that auxiliary information acts partly as data augmentation: the same video can be described many ways, and each description teaches the aligner something new. If so, purely scaling caption diversity could bring gains without changing architecture.
- Q&A methods may owe their zero-shot edge to query expansion rather than to richer video understanding; an ablation that feeds the original query into the same model would test this.
- The dataset analysis implies designers of new benchmarks should include audio and multiple human captions by default, since those are the signals most consistently tied to gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews 81 text-to-video retrieval methods that incorporate auxiliary information beyond the video-text pair. The authors propose a taxonomy that divides auxiliary information into video-extracted (visual info, modality info, generated info) and text-extracted (textual info, generated info) categories, with subclasses such as Video Experts, Regions of Interest, Audio, Generated Video Captions, and Question-Answering. They also inventory the auxiliary information available in pre-training and downstream datasets (Table 3) and compare reported Recall@1, mAP, and nDCG results across eight benchmarks in zero-shot and fine-tuned settings (Section 4.2 and Appendix B). The paper concludes that generated and multimodal auxiliary information generally outperforms text-refinement approaches, that Q&A methods dominate zero-shot settings, and that auxiliary information 'has proven to be a powerful tool' in this field (Section 6).
Significance. If the comparative claims were reliable, the survey would fill a genuine gap: no prior survey focuses specifically on auxiliary information in text-to-video retrieval. The taxonomy is clearly presented and largely internally consistent, and the dataset inventory in Table 3 is a useful reference contribution. The appendix tables are a strength: they provide full metric tables for many methods, and I verified that key numbers such as MQVR's R@1 of 73.1 on MSR-VTT and 61.7 on MSVD appear consistently in both the main text and the appendix. The main weakness is that the central comparative conclusion is not supported by the reported evidence as presented, because the comparison across auxiliary-information classes is confounded with model scale, pre-training data, and inference protocol. The survey is still potentially valuable as a structured reference, but the comparative claims need substantial qualification or additional controlled analysis.
major comments (3)
- [Sec. 4.2.1 and Tables 4-11] The comparison treats R@1 differences between auxiliary-information classes as evidence about the value of the auxiliary information, but class membership is confounded with pre-training data, model scale, and training budget. The zero-shot leaders in GMC/MM (InternVideo2, VAST, GRAM, LanguageBind) are foundation-scale models trained on large multimodal corpora, whereas the GTC/POS methods they are compared against (e.g., HowToCaption, GQE, TACo) are mostly built on smaller CLIP/HowTo100M-scale bases; the same pattern appears in the fine-tuned tables. The statement in Sec. 4.2.1 that the zero-shot/fine-tuned split 'allows a fairer comparison' is therefore not supported, and the conclusion in Sec. 6 that auxiliary information 'has proven to be a powerful tool' cannot be tested from these tables because no no-auxiliary baseline (e.g., plain CLIP or a single-encoder model without auxiliary inputs) is reported anywhere. I recommend adding at least one no-auxiliary baseline per dataset and setting, adding explicit caveats about pre-training and scale, or reframing the Sec. 4.2.3 and Sec. 6 conclusions as trends among the surveyed methods rather than causal statements about auxiliary-information type.
- [Sec. 2.3 and Tables 5-9] The Q&A methods (IVR-QA, MERLIN) are iterative question-generation/rerank pipelines that generate candidate-specific questions and answers at inference, as described in Sec. 2.3, whereas the other zero-shot entries are single-query embedding models. Reporting both under the same 'Zero-Shot' heading (Sec. 4.2.2, e.g., MSR-VTT R@1 67.9 for IVR-QA vs 55.9 for InternVideo2) conflates retrieval performance with a different inference protocol and a larger compute budget. The claim that 'Q&A-based approaches achieve the best performance' in zero-shot across five datasets is a statement about the pipeline, not about the auxiliary-information type. To make the comparison fair, Q&A methods should be either excluded from the per-class ranking or clearly reported as a separate protocol, with the additional inference cost stated.
- [Sec. 4.2.3] The conclusion that 'GVC- and GMC-based methods generally outperform text-refinement approaches (GTC)' is based on a hand-count of per-dataset winners (three datasets for GVC vs two for GTC), not on any controlled comparison. Within each dataset cell the number of methods per class is small, and the winner is often the only representative of a class (e.g., GTC on LSMDC is represented largely by GQE). Such counts are not a reliable basis for the general claim; they should be softened to observations about specific methods, or supported by a more systematic analysis that controls for backbone, pre-training data, and training budget.
minor comments (4)
- [Sec. 2.2] The first paragraph of Sec. 2.2 contains a duplicated reference: 'Multi-Modalities (MM) [30, 115, 115]'; the second [115] should be removed or replaced.
- [Table 3] The table uses the same check symbol for both video- and text-extracted information, despite the caption saying '✓ and ✓ indicate the video- and text-extracted information available'; using distinct symbols would avoid ambiguity.
- [Sec. 4.1] The class is defined as 'Temporal Dependency (TD)' in Table 2 but referred to as 'Textual Dependencies (TD)' in the Sec. 4.1 text; use one name consistently throughout.
- [Sec. 2.3 and Table 1] The term 'Questions&Answers' is unconventional; consider 'Question-Answering (Q&A)' for consistency with standard terminology.
Circularity Check
No circularity: the review's conclusions rest on externally reported benchmark scores; self-citations are non-load-bearing.
full rationale
This is a survey paper, not a derivation. Its taxonomy (Tables 1-2) is an organizational scheme applied to 81 externally published methods, and the comparative findings in Sections 4.2 and 6 are empirical summaries of Recall@1, mAP, and nDCG values reported by those external papers. No equation is fitted, no quantity is predicted from another quantity, and no result is defined in terms of the conclusion, so the self-definitional and fitted-input-as-prediction patterns do not apply. The authors' own works appear (ConTra, Ego4D, EPIC-KITCHENS, JPoSE, EgoClip), but only as ordinary methods or datasets: ConTra is listed in the YouCook2 fine-tuned table with R@1 = 16.7, a mid-table score, and Ego4D is cited for its publicly documented auxiliary annotations, not as evidence for the review's central comparative claims. No uniqueness theorem, ansatz, or load-bearing argument is imported from the authors' prior work. The skeptical concern that R@1 gaps between auxiliary-information classes may be confounded by model scale, pretraining corpora, or Q&A rerank inference protocols is a valid evidentiary limitation, and the absence of a no-auxiliary baseline weakens the causal phrasing in the conclusion; however, a confound is not circularity. The central content of the review is self-contained in the sense that it compiles and compares external benchmark numbers, and the self-citations do not carry the argument. Score 0 accordingly.
Assumptions & free parameters
assumptions (4)
- domain assumption The 81 reviewed papers are representative of the field of T2V retrieval with auxiliary information.
- domain assumption Reported metrics from the original papers are accurate and mutually consistent.
- ad hoc to paper Auxiliary information type is a meaningful axis for comparing method performance.
- domain assumption Excluding methods that do not report R@1 (VATT, LINAS) or that report only averaged metrics (HierVL, EgoInstructor) does not bias the quantitative conclusions.
Cite this review
Pith. "Pith review of Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review." pith.science (2026). https://pith.science/paper/QMT57ARQ
@misc{pith2026250523952,
author = {Pith},
title = {Pith review of: Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/QMT57ARQ}},
note = {Machine review of arXiv:2505.23952}
}
read the original abstract
Text-to-Video (T2V) retrieval aims to identify the most relevant item from a gallery of videos based on a user's text query. Traditional methods rely solely on aligning video and text modalities to compute the similarity and retrieve relevant items. However, recent advancements emphasise incorporating auxiliary information extracted from video and text modalities to improve retrieval performance and bridge the semantic gap between these modalities. Auxiliary information can include visual attributes, such as objects; temporal and spatial context; and textual descriptions, such as speech and rephrased captions. This survey comprehensively reviews 81 research papers on Text-to-Video retrieval that utilise such auxiliary information. It provides a detailed analysis of their methodologies; highlights state-of-the-art results on benchmark datasets; and discusses available datasets and their auxiliary information. Additionally, it proposes promising directions for future research, focusing on different ways to further enhance retrieval performance using this information.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
A camera-trap TVR benchmark of 135 ethology queries plus an interpretable SALMA-to-JSON plus constrained-LLM-parser pipeline yields 34% set F1, beating zero-shot VLMs at 18%.
Reference graph
Works this paper leans on
-
[1]
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021. VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text. InConference on Neural Information Processing Systems (NeurIPS)
2021
-
[2]
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman. 2023. HierVL: Learning Hierarchical Video-Language Embeddings. In Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[3]
Zechen Bai, Tianjun Xiao, Tong He, Pichao Wang, Zheng Zhang, Thomas Brox, and Mike Zheng Shou. 2025. Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach. arXiv (2025). Accepted at ICLR 2025
2025
-
[4]
Min Cao, Shiping Li, Juntao Li, Liqiang Nie, and Min Zhang. 2022. Image-text Retrieval: A Survey on Recent Research and Development. In International Joint Conference on Artificial Intelligence (IJCAI)
2022
-
[5]
Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David Harwath, James R
Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogério Feris, David Harwath, James R. Glass, Michael Picheny, and Shih-Fu Chang. 2021. Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos. In International Conference on Computer Vision (ICCV)
2021
-
[6]
Chen and William B
David L. Chen and William B. Dolan. 2011. Collecting Highly Parallel Data for Paraphrase Evaluation. In Annual Meeting of the Association for Computational Linguistics (ACL)
2011
-
[7]
Lei Chen, Zhen Deng, Libo Liu, and Shibai Yin. 2024. Multilevel Semantic Interaction Alignment for Video–Text Cross-Modal Retrieval. Transactions on Circuits and Systems for Video Technology (TCSVT) (2024)
2024
-
[8]
Sihan Chen, Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Xinxin Zhu, and Jing Liu. 2023. VAST: A Vision- Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset. In Conference on Neural Information Processing Systems (NeurIPS
2023
Show all 120 references
-
[9]
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu. 2020. Fine-Grained Video-Text Retrieval With Hierarchical Graph Reasoning. In Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[10]
Yizhen Chen, Jie Wang, Lijian Lin, Zhongang Qi, Jin Ma, and Ying Shan. 2023. Tagging before Alignment: Integrating Multi-Modal Tags for Video-Text Retrieval. InConference on Artificial Intelligence (AAAI)
2023
-
[11]
Xing Cheng, Hezheng Lin, Xiangyu Wu, Fan Yang, and Dong Shen. 2021. Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss. CoRR abs/2109.04290 (2021)
2021 arXiv
-
[12]
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. 2024. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in , Vol. 1, No. 1, Article . Publication date: August 2025. 24 Frag...
2024 arXiv
-
[13]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-...
2023
-
[14]
Giordano Cicchetti, Eleonora Grassucci, Luigi Sigillo, and Danilo Comminiello. 2025. Gramian Multimodal Represen- tation Learning and Alignment. In International Conference on Learning Representations (ICLR)
2025
-
[15]
Begum Citamak, Ozan Caglayan, Menekse Kuyu, Erkut Erdem, Aykut Erdem, Pranava Madhyastha, and Lucia Specia
-
[16]
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, and Yang Liu. 2021. TeachText: CrossModal Generalized Distillation for Text-Video Retrieval. InInternational Conference on Computer Vision (ICCV)
2021
-
[17]
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018. Scaling Egocentric Vision: The EPIC-KITCHENS Dataset. In European Conference on Computer...
2018
-
[18]
Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang. 2022. Dual Encoding for Video Retrieval by Text. Transactions on Pattern Analysis and Machine Intelligence (TPAMI) (2022)
2022
-
[19]
Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang, and Qingpei Guo. 2024. M2-RAAP: A Multi- Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval. In Conference on Research and Development in Info...
2024
-
[20]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. CoRR abs/2407.21783 (2024)
2024 arXiv
-
[21]
Maksim Dzabraev, Maksim Kalashnikov, Stepan Komkov, and Aleksandr Petiushko. 2021. MDMMT: Multidomain Multimodal Transformer for Video Retrieval. In Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[22]
Sheng Fang, Shuhui Wang, Junbao Zhuo, Xinzhe Han, and Qingming Huang. 2022. Learning Linguistic Association Towards Efficient Text-Video Retrieval. InEuropean Conference on Computer Vision (ECCV)
2022
-
[23]
Zerun Feng, Zhimin Zeng, Caili Guo, and Zheng Li. 2020. Exploiting Visual Semantic Reasoning for Video-Text Retrieval. In International Joint Conference on Artificial Intelligence (IJCAI)
2020
-
[24]
Adriano Fragomeni, Michael Wray, and Dima Damen. 2022. ConTra: (Con)text (Tra)nsformer for Cross-Modal Video Retrieval. In Asian Conference on Computer Vision (ACCV)
2022
-
[25]
Valentin Gabeur, Arsha Nagrani, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2022. Masking Modalities for Cross-modal Video Retrieval. In Winter Conference on Applications of Computer Vision (W ACV)
2022
-
[26]
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2020. Multi-modal Transformer for Video Retrieval. In European Conference on Computer Vision (ECCV)
2020
-
[27]
Damianos Galanopoulos and Vasileios Mezaris. 2022. Are All Combinations Equal? Combining Textual and Visual Features with Multiple Space Learning for Text-Based Video Retrieval. In European Conference on Computer Vision (ECCV)
2022
-
[28]
Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li, Ying Shan, Xiaohu Qie, and Ping Luo. 2022. Bridging Video-text Retrieval with Multiple Choice Questions. In Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[29]
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, and Thomas Brox. 2020. COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning. InConference on Neural Information Processing Systems (NeurIPS)
2020
-
[30]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind One Embedding Space to Bind Them All. In Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[31]
Satya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. 2022. X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval. In Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[32]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengme...
2025
-
[33]
Donghoon Han, Eunhwan Park, Gisang Lee, Adam Lee, and Nojun Kwak. 2024. MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline. In Conference on Empirical Methods in Natural Language Processing (EMNLP)
2024
-
[34]
Ning Han, Jingjing Chen, Guangyi Xiao, Hao Zhang, Yawen Zeng, and Hao Chen. 2021. Fine-grained Cross-modal Alignment Network for Text-Video Retrieval. In Multimedia Conference (MC)
2021
-
[35]
Ning Han, Yawen Zeng, Chuhao Shi, Guangyi Xiao, Hao Chen, and Jingjing Chen. 2024. BiC-Net: Learning Efficient Spatio-temporal Relation for Text-Video Retrieval. Transactions on Multimedia Computing, Communications, and Applications (TOMM) (2024)
2024
-
[36]
Zhichao Han, Azreen Azman, Mas Rina Mustaffa, and Fatimah Khalid. 2024. Cross-Modal Retrieval: A Review of Methodologies, Datasets, and Future Perspectives. IEEE Access (2024)
2024
-
[37]
Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li, and Weiping Wang. 2021. Multi-Feature Graph Attention Network for Cross-Modal Video-Text Retrieval. InInternational Conference on Multimedia Retrieval (ICMR)
2021
-
[38]
Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li, Weiping Wang, and Dan Meng. 2021. What Matters: Attentive and Relational Feature Aggregation Network for Video-Text Retrieval. In International Conference on Multimedia & Expo (ICME)
2021
-
[39]
Feng He, Qi Wang, Zhifan Feng, Wenbin Jiang, Yajuan Lü, Yong Zhu, and Xiao Tan. 2021. Improving Video Retrieval by Adaptive Margin. In Conference on Research and Development in Information Retrieval (SIGIR)
2021
-
[40]
Willy Fitra Hendria. 2023. MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian. CoRR abs/2306.11341 (2023)
2023 arXiv
-
[41]
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2017. Localizing Moments in Video with Natural Language. In International Conference on Computer Vision (ICCV)
2017
-
[42]
Fan Hu, Aozhu Chen, Ziyue Wang, Fangming Zhou, Jianfeng Dong, and Xirong Li. 2022. Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval. In European Conference on Computer Vision (ECCV)
2022
-
[43]
Poyao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, and Alex Hauptmann. 2021. Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models. In Conference of the North American Chapter of the Association for Computationa...
2021
-
[44]
Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg, Ashutosh Sanan, and Mohamed Omar. 2023. Audio- Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment. InInternational Conference on Computer Vision (ICCV)
2023
-
[45]
Weike Jin, Zhou Zhao, Pengcheng Zhang, Jieming Zhu, Xiuqiang He, and Yueting Zhuang. 2021. Hierarchical Cross-Modal Graph Consistency Learning for Video-Text Retrieval. InConference on Research and Development in Information Retrieval (SIGIR)
2021
-
[46]
Parminder Kaur, Husanbir Singh Pannu, and Avleen Kaur Malhi. 2021. Comparative analysis on cross-modal information retrieval: A review. Computer Science Review (CSR) (2021)
2021
-
[47]
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017. Dense-Captioning Events in Videos. In International Conference on Computer Vision (ICCV)
2017
-
[48]
Alexander Kunitsyn, Maksim Kalashnikov, Maksim Dzabraev, and Andrei Ivaniuta. 2022. MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization. CoRR abs/2203.07086 (2022)
2022 arXiv
-
[49]
Huy Le, Tung Kieu, Anh Nguyen, and Ngan Le. 2024. WAVER: Writing-Style Agnostic Text-Video Retrieval Via Distilling Vision-Language Models Through Open-Vocabulary Knowledge. In International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2024
-
[50]
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven C. H. Hoi. 2022. Align and Prompt: Video-and- Language Pre-training with Entity Prompts. In Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[51]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning (ICML)
2023
-
[52]
Tieying Li, Lingdu Kong, Xiaochun Yang, Bin Wang, and Jiaxing Xu. 2024. Bridging Modalities: A Survey of Cross-Modal Image-Text Retrieval. Chinese Journal of Information Fusion (CJIF) (2024)
2024
-
[53]
Xirong Li, Fangming Zhou, Chaoxi Xu, Jiaqi Ji, and Gang Yang. 2021. SEA: Sentence Encoder Assembly for Video Retrieval by Textual Queries. Transactions on Multimedia (TM) (2021)
2021
-
[54]
Yili Li, Jing Yu, Keke Gai, Bang Liu, Gang Xiong, and Qi Wu. 2024. T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval. In International Conference on Multimedia (ICM)
2024
-
[55]
Kaiqu Liang and Samuel Albanie. 2023. Simple Baselines for Interactive Video Retrieval with Questions and Answers. In International Conference on Computer Vision (ICCV) . , Vol. 1, No. 1, Article . Publication date: August 2025. 26 Fragomeni et al
2023
-
[56]
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rong- Cheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu, and Mike Zheng Shou. 2022. Egocentric Video-Language Pretraining. In ...
2022
-
[57]
Yan-Bo Lin, Jie Lei, Mohit Bansal, and Gedas Bertasius. 2022. EclipSE: Efficient Long-Range Video Retrieval Using Sight and Sound. In European Conference on Computer Vision (ECCV)
2022
-
[58]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning. In Conference on Neural Information Processing Systems (NeurIPS)
2023
-
[59]
Song Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen, Wenkui Ding, and Zhongyuan Wang. 2021. HiT: Hierarchical Transformer with Momentum Contrast for Video-Text Retrieval. In International Conference on Computer Vision (ICCV)
2021
-
[60]
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. 2019. Use What You Have: Video retrieval using representations from collaborative experts. In British Machine Vision Conference (BMVC)
2019
-
[61]
Gang Lv, Yining Sun, and Fudong Nian. 2024. Video-text retrieval via multi-modal masked transformer and adaptive attribute-aware graph convolutional network. Multim. Syst. (2024)
2024
-
[62]
Rawat, and Thomas B
Neelu Madan, Andreas Møgelmose, Rajat Modi, Yogesh S. Rawat, and Thomas B. Moeslund. 2024. Foundation Models for Video Understanding: A Survey. CoRR abs/2405.03770 (2024)
2024 arXiv
-
[63]
Avinash Madasu, Estelle Aflalo, Gabriela Ben Melech Stan, Shachar Rosenman, Shao-Yen Tseng, Gedas Bertasius, and Vasudev Lal. 2023. MuMUR: Multilingual Multimodal Universal Retrieval. Inf. Retr. J. (2023)
2023
-
[64]
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020. End-to- End Learning of Visual Representations From Uncurated Instructional Videos. In Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[65]
Antoine Miech, Ivan Laptev, and Josef Sivic. 2018. Learning a Text-Video Embedding from Incomplete and Heteroge- neous Data. CoRR abs/1804.02516 (2018)
2018 arXiv
-
[66]
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. InInternational Conference on Computer Vision (ICCV)
2019
-
[67]
Arsha Nagrani, Paul Hongsuck Seo, Bryan Seybold, Anja Hauth, Santiago Manen, Chen Sun, and Cordelia Schmid
-
[68]
Thong Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy Nguyen, See-Kiong Ng, and Anh Tuan Luu. 2024. Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives. In Annual Meeting of the Association for Com...
2024
-
[69]
B. V. Patel and B. B. Meshram. 2012. Content based video retrieval systems. CoRR abs/1205.1641 (2012)
2012 arXiv
-
[70]
Hauptmann, João F
Mandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze, Alexander G. Hauptmann, João F. Henriques, and Andrea Vedaldi. 2021. Support-set bottlenecks for video-text representation learning. In International Conference on Learning Representations (ICLR)
2021
-
[71]
Yuxin Peng, Xin Huang, and Yunzhen Zhao. 2018. An Overview of Cross-Media Retrieval: Concepts, Methodologies, Benchmarks, and Challenges. Transactions on Circuits and Systems for Video Technology (TCSVT) (2018)
2018
-
[72]
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. 2015. A dataset for Movie Description. In Conference on Computer Vision and Pattern Recognition (CVPR)
2015
-
[73]
Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Au- dhkhasi, Hilde Kuehne, Rameswar Panda, Rogério Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, and James R. Glass. 2021. AVLnet: Learning Audio-Visual L...
2021
-
[74]
Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogério Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, and James R. Glass. 2023. C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval. InIntern...
2023
-
[75]
Glass, and Hilde Kuehne
Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogério Feris, David Harwath, James R. Glass, and Hilde Kuehne. 2022. Everything at Once - Multi-modal Fusion Transformer for Video Retrieval. In Conference on Computer Vision and Pattern Recognit...
2022
-
[76]
Nina Shvetsova, Anna Kukleva, Xudong Hong, Christian Rupprecht, Bernt Schiele, and Hilde Kuehne. 2024. How- ToCaption: Prompting LLMs to Transform Video Annotations at Scale. In European Conference on Computer Vision (ECCV)
2024
-
[77]
Nina Shvetsova, Anna Kukleva, Bernt Schiele, and Hilde Kuehne. 2023. In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval. In International Conference on Computer Vision (ICCV) . , Vol. 1, No. 1, Article . Publication date: August 2025. Lev...
2023
-
[78]
Xue Song, Jingjing Chen, and Yu-Gang Jiang. 2023. Relation Triplet Construction for Cross-modal Text-to-Video Retrieval. In International Conference on Multimedia (ICM)
2023
-
[79]
Xue Song, Jingjing Chen, Zuxuan Wu, and Yu-Gang Jiang. 2022. Spatial-Temporal Graphs for Cross-Modal Text2Video Retrieval. Transactions on Multimedia (TM) (2022)
2022
-
[80]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. CoRR abs/2302.13971 (2023)
2023 arXiv
-
[81]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. CoRR abs/2307.09288 (2023)
2023 arXiv
-
[82]
Alex Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan, Xudong Lin, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2022. Object-aware Video-language Pre-training for Retrieval. In Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[83]
Haoran Wang, Di Xu, Dongliang He, Fu Li, Zhong Ji, Jungong Han, and Errui Ding. 2022. Boosting Video-Text Retrieval with Explicit High-Level Semantics. In International Conference on Multimedia (ICM)
2022
-
[84]
Jinpeng Wang, Bin Chen, Dongliang Liao, Ziyun Zeng, Gongfu Li, Shu-Tao Xia, and Jin Xu. 2022. Hybrid Contrastive Quantization for Efficient Cross-View Video Retrieval. In Web Conference (WWW)
2022
-
[85]
Kaiye Wang, Qiyue Yin, Wei Wang, Shu Wu, and Liang Wang. 2016. A Comprehensive Survey on Cross-modal Retrieval. CoRR abs/1607.06215 (2016)
2016 arXiv
-
[86]
Wei Wang, Junyu Gao, Xiaoshan Yang, and Changsheng Xu. 2021. Learning Coarse-to-Fine Graph Neural Networks for Video-Text Retrieval. Transactions on Multimedia (TM) (2021)
2021
-
[87]
Wenzhe Wang, Mengdan Zhang, Runnan Chen, Guanyu Cai, Penghao Zhou, Pai Peng, Xiaowei Guo, Jian Wu, and Xing Sun. 2021. Dig into Multi-modal Cues for Video Retrieval with Hierarchical Alignment. In International Joint Conference on Artificial Intelligence (IJCAI)
2021
-
[88]
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019. VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research. In International Conference on Computer Vision (ICCV)
2019
-
[89]
Xiaohan Wang, Linchao Zhu, and Yi Yang. 2021. T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval. In Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[90]
Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Jilan Xu, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, and Limin Wang. 2024. InternVideo2: Scaling Foundation Models for Multimodal ...
2024
-
[91]
Ziyue Wang, Aozhu Chen, Fan Hu, and Xirong Li. 2022. Learn to Understand Negation in Video Retrieval. In International Conference on Multimedia (ICM)
2022
-
[92]
Zeyu Wang, Yu Wu, Karthik Narasimhan, and Olga Russakovsky. 2022. Multi-query Video Retrieval. InEuropean Conference on Computer Vision (ECCV)
2022
-
[93]
Michael Wray, Gabriela Csurka, Diane Larlus, and Dima Damen. 2019. Fine-Grained Action Retrieval Through Multiple Parts-of-Speech Embeddings. In International Conference on Computer Vision (ICCV)
2019
-
[94]
Peng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv, and Jing Liu. 2021. HANet: Hierarchical Alignment Networks for Video-Text Retrieval. In Multimedia Conference (MC)
2021
-
[95]
Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang. 2023. Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?. InConference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[96]
Xiaoyu Wu, Jiayao Qian, and Tiantian Wang. 2023. Text-video retrieval method based on enhanced self-attention and multi-task learning. Multimedia Tools and Applications (MTA) (2023)
2023
-
[97]
Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song, Sipeng Zheng, and Qin Jin. 2025. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?. In International Conference on Learning Representations (ICLR)
2025
-
[98]
Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. 2024. Retrieval-Augmented Egocentric Video Captioning. In Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[99]
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. MSR-VTT: A Large Video Description Dataset for Bridging Video and Language. In Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[100]
Ran Xu, Caiming Xiong, Wei Chen, and Jason J. Corso. 2015. Jointly Modeling Deep Video and Compositional Text to Bridge Vision and Language in a Unified Framework. In Conference on Artificial Intelligence (AAAI)
2015
-
[101]
Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo. 2023. CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Alignment. In International Conference on Learning Representations (ICLR). , Vol. 1, No. 1, Article . Publication da...
2023
-
[102]
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. 2021. TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment. In International Conference on Computer Vision (ICCV)
2021
-
[103]
Xun Yang, Jianfeng Dong, Yixin Cao, Xun Wang, Meng Wang, and Tat-Seng Chua. 2020. Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval. In Conference on Research and Development in Information Retrieval (SIGIR)
2020
-
[104]
Xuzheng Yu, Chen Jiang, Xingning Dong, Tian Gan, Ming Yang, and Qingpei Guo. 2024. SHE-Net: Syntax-Hierarchy- Enhanced Text-Video Retrieval. CoRR abs/2404.14066 (2024)
2024 arXiv
-
[105]
Youngjae Yu, Jongseok Kim, and Gunhee Kim. 2018. A Joint Sequence Fusion Model for Video Question Answering and Retrieval. In European Conference on Computer Vision (ECCV)
2018
-
[106]
Youngjae Yu, Hyungjin Ko, Jongwook Choi, and Gunhee Kim. 2017. End-to-End Concept Word Detection for Video Captioning, Retrieval, and Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR)
2017
-
[107]
Bowen Zhang, Hexiang Hu, and Fei Sha. 2018. Cross-Modal and Hierarchical Modeling of Video and Text. InEuropean Conference on Computer Vision (ECCV)
2018
-
[108]
Chuhan Zhang, Ankush Gupta, and Andrew Zisserman. 2023. Helping Hands: An Object-Aware Ego-Centric Video Recognition Model. In International Conference on Computer Vision (ICCV)
2023
-
[110]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. In Conference on Empirical Methods in Natural Language Processing (EMNLP)
2023
-
[111]
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar. 2023. Learning Video Representations from Large Language Models. In Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[112]
Yue Zhao, Long Zhao, Xingyi Zhou, Jialin Wu, Chun-Te Chu, Hui Miao, Florian Schroff, Hartwig Adam, Ting Liu, Boqing Gong, Philipp Krähenbühl, and Liangzhe Yuan. 2024. Distilling Vision-Language Models on Millions of Videos. In Conference on Computer Vision and Pattern Recognit...
2024
-
[113]
Kun Zhou, Fadratul Hafinaz Hassan, and Gan Keng Hoon. 2023. The State of the Art for Cross-Modal Retrieval: A Survey. IEEE Access (2023)
2023
-
[114]
Luowei Zhou, Chenliang Xu, and Jason J. Corso. 2018. Towards Automatic Learning of Procedures From Web Instructional Videos. In Conference on Innovative Applications of Artificial Intelligence (IAAI)
2018
-
[115]
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, Hongfa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Caiwan Zhang, Zhifeng Li, Wei Liu, and Li Yuan. 2024. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment. In I...
2024
-
[116]
Cunjuan Zhu, Qi Jia, Wei Chen, Yanming Guo, and Yu Liu. 2023. Deep learning for video-text retrieval: a review. International Journal of Multimedia Information Retrieval (IJMIR) (2023)
2023
-
[117]
Lei Zhu, Tianshi Wang, Fengling Li, Jingjing Li, Zheng Zhang, and Heng Tao Shen. 2023. Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions. CoRR abs/2308.14263 (2023)
2023 arXiv
-
[118]
Linchao Zhu and Yi Yang. 2020. ActBERT: Learning Global-Local Video-Text Representations. In Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[119]
Lei Zhu, Chaoqun Zheng, Weili Guan, Jingjing Li, Yang Yang, and Heng Tao Shen. 2024. Multi-Modal Hashing for Efficient Multimedia Retrieval: A Survey. Transactions on Knowledge and Data Engineering (TKDE) (2024). , Vol. 1, No. 1, Article . Publication date: August 2025. Levera...
2024
-
[2021]
Machine Translation (MT) (2021)
MSVD-Turkish: a comprehensive multimodal video dataset for integrated vision and language research in Turkish. Machine Translation (MT) (2021)
2021
-
[2022]
In European Conference on Computer Vision (ECCV)
Learning Audio-Video Modalities from Image Captions. In European Conference on Computer Vision (ECCV)
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.