REVIEW 5 major objections 7 minor 98 references
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
T0 review · 5 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A survey maps a decade of visual question answering through an extractive-versus-abstractive lens.
desk verdict Frequent citation errors and a systematic Visual/Video drift make this VQA survey unreliable as a reference, despite a sensible structure and a useful extractive/abstractive framing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central organizing device is the extractive-versus-abstractive paradigm distinction, applied section by section: extractive VQA treats answering as selecting or grounding a predefined answer, while abstractive VQA treats it as generating a free-form, context-aware response. Carrying the argument alongside this axis are the technical mechanisms the survey identifies as milestone drivers: attention in its stacked, co-attention, and bottom-up/top-down forms; compositional reasoning modules such as neural module networks, scene graphs, and MAC; and transformer-based vision-language pretraining represented by LXMERT, ViLBERT, UNITER, OSCAR, CLIP, BLIP-2, and Flamingo. The survey uses these mechanisms as the stages of its chronological narrative and as the basis for its tabulated summaries of models and methods.
What would settle it
Look up any sample of the paper's in-text citations and check whether the cited document actually makes the claim attributed to it; for example, the text describes (Chen et al., 2017) as a Spatial Memory Network for VQA, while the reference list entry points to a paper on spatial memory for object detection, so that single check can directly test the survey's central reliability claim.
Extended reading notes
Core claim
The paper claims that the entire trajectory of VQA can be understood as movement along two axes: a chronological sequence of technical leaps, from deep CNN-LSTM fusion and bilinear pooling through attention and compositional reasoning to transformer-based vision-language pretraining, and a persistent paradigm split between extractive answer retrieval and abstractive free-form generation. It argues that transformers and large-scale multimodal pretraining have been the decisive drivers of recent progress, and that the same extractive/abstractive split reappears in domain-specific VQA for medicine and entertainment. On the paper's own terms, the central discovery is not a new model but an organizing narrative: one framework under which the field's major benchmarks, architectures, and open problems can be laid out coherently.
Load-bearing premise
The survey's value depends on its literature summaries being accurate and on its terminology staying stable; the manuscript contains multiple citation mismatches and at least one passage that says 'Video Question Answering' where 'Visual Question Answering' is meant, so if these errors are representative, the map it draws is unreliable.
Editorial extensions
If this is right
- A newcomer to VQA can use the survey as a staged reading list: CNN-LSTM fusion, bilinear pooling, attention, compositional reasoning, vision-language pretraining, then large multimodal models.
- The extractive/abstractive split gives a vocabulary for comparing models across eras, so that a 2016 attention model and a 2023 multimodal LLM can be discussed in the same terms.
- Domain-specific VQA in medical imaging, movie understanding, fashion, and scientific figures inherits the same paradigm split and the same dependence on benchmark datasets.
- The paper's stated open challenges, including dataset bias, interpretability, and the need for common-sense and external knowledge, define the next targets for VQA research.
- If the narrative is right, transformer-based vision-language pretraining, rather than task-specific fusion architectures, is what drove the largest accuracy gains in VQA's recent history.
Reading between the lines
- The extractive/abstractive axis could be sharpened into a testable design spectrum: a model's position on it predicts whether its failures show up as wrong label choices or as ungrounded fluent text, which would give evaluators a cheap diagnostic.
- The framework suggests a concrete next benchmark: hold the image constant while varying only the extractive-versus-abstractive demand of the question, isolating what each paradigm contributes.
- If the survey's historical narrative is right, future VQA progress will come less from new fusion tricks and more from pretraining data scale and external-knowledge integration, since those are the levers the surveyed history shows moving performance.
- Readers should anchor on the abstract and early sections when using the framework, because some later passages speak of 'Video Question Answering' where 'Visual Question Answering' is meant.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of Visual Question Answering (VQA), organized by a claimed dichotomy between extractive and abstractive approaches. It spans early CNN-LSTM models, attention mechanisms, compositional reasoning, vision-language pre-training, domain-specific applications, and future directions, and it includes a large summary table. The abstract and opening sections describe VQA as starting in 2015 and claim the survey is comprehensive.
Significance. If the survey were reliable, it would serve as a useful entry point for newcomers to VQA. The chronological organization and the table of models are potentially valuable. The paper also attempts to cover a wide range of subareas, including medical VQA and recent large multimodal models. However, the manuscript currently contains numerous citation errors, internal inconsistencies, and a systematic conflation of image-based VQA with video question answering, which invalidate its central claim of being a trustworthy, comprehensive overview. The strengths—a broad structure and an extensive table—are outweighed by the fact that the details are frequently incorrect.
major comments (5)
- [Section 5 and Table 1] The paper attributes ViLT to (Radford et al., 2021), the same citation used for CLIP; ViLT is by Kim et al. (2021) and does not appear in the reference list at all. Section 5 states 'Models such as CLIP (Radford et al., 2021) and ViLT (Radford et al., 2021)', which is a direct misattribution of a different architecture. Similarly, Section 2.4 attributes the DAQUAR dataset to (Malinowski et al., 2014), but the reference given is 'Multimodal Learning with Deep Convolutional Neural Networks', not the DAQUAR dataset paper (Malinowski and Fritz, 2014). These are not isolated typos; they affect foundational works and make the survey unreliable as a literature map.
- [Sections 7, 8, and 9] The survey repeatedly redefines VQA as 'Video Question Answering'. The opening sentences of Sections 7, 8, and 9 all use the definition 'Video Question Answering (VQA)', whereas Sections 1–6 are about image-based VQA. This is not a harmless abbreviation: the challenges and future directions discuss temporal reasoning, video datasets, and video-specific models (VideoBERT, Frozen in Time, VideoDistill) as if they were part of the image-VQA story, without acknowledging that the scope has shifted. The paper's own conclusion then frames the entire survey as an introduction to video question answering, contradicting the abstract and the earlier sections.
- [Section 3.3 and Table 1] The central extractive/abstractive dichotomy is applied inconsistently. For example, Stacked Attention Networks are described in Section 3.3 as an extractive method, but Table 1 classifies 'Stacked Attention Networks for Image Question Answering' as abstractive. Multimodal Compact Bilinear pooling is described in Section 3.2 under the extractive paradigm, yet Table 1 lists 'Multimodal Compact Bilinear Pooling' as abstractive. The paper never defines the criteria for labeling a model extractive or abstractive, so the framework cannot be applied in a principled way and its use throughout the survey is not reliable.
- [Section 6.1] The paper attributes GPT-4V to 'Radford and Narasimhan (2018)', which is the reference for the original GPT paper, despite also citing (Li et al., 2024c) for GPT-4V in the same paragraph. Section 3.3 also misattributes the 'Knowing When to Look' image-captioning paper (Lu et al., 2017) to a VQA mixed-attention mechanism. Such errors are frequent enough that the survey cannot be used as a reliable secondary source, and they undercut the paper's claim of a comprehensive overview.
- [Table 1 and Sections 1–2] The comprehensiveness claim is undercut by both omissions and irrelevant entries. The survey omits several influential VQA systems, such as LLaVA, InstructBLIP, and other recent large multimodal models, despite discussing such models in Section 8. Conversely, Table 1 includes entries that are not VQA works, such as 'Finding Structure in Time' (Elman, 1990) and 'ImageNet Classification' (2012), and it contains a duplicate VisualBERT row. The table therefore does not support the abstract's assertion that this is a comprehensive overview of VQA's evolution.
minor comments (7)
- [Section 1.2] The paper writes 'introduced VisualQA'; the dataset is 'VQA' or 'Visual Question Answering', not 'VisualQA'.
- [Section 1.4] The sentence 'Each section is divided into two main paradigms: extractive and abstractive.the Section 2 reviews...' has a missing space, a lowercase 't', and an extraneous 'the' before 'Section 2'.
- [Section 5] The opening line of Section 5 says 'Vision-Question Answering (VQA)' instead of 'Visual Question Answering'.
- [Sections 4.4 and 5] The paper inconsistently spells 'VilBERT' and 'ViLBERT', and the reference list contains two entries for the same ViLBERT paper (Lu et al., 2019a and 2019b).
- [Table 1] The VisualBERT row appears twice with identical wording, and several rows list 'NaN' in the 'Datasets Used' column (e.g., 'Learning Transferable Visual Models'), indicating the table data was not carefully curated.
- [References] Several works cited in the text are missing from the reference list, including the actual ViLT paper (Kim et al., 2021), the DAQUAR paper (Malinowski and Fritz, 2014), and a correct UNITER citation; the existing UNITER entry (Chen et al., 2020) has an author list that does not match the published paper.
- [Table 1] The year for Visual Genome is listed as 2017 in the table but as 2016 in the references; please make these entries consistent.
Circularity Check
No circularity: the survey makes no derivation-based claims, and its self-citations are not load-bearing.
full rationale
This is a survey paper rather than a derivation or empirical study, so the standard circularity patterns do not apply. The central claim is that the paper offers a structured overview of the evolution of Visual Question Answering, organizing prior work under extractive and abstractive paradigms. There is no fitted parameter that is later relabeled as a prediction, no equation that reduces to its own input, and no uniqueness theorem imported from the authors' prior work to force a conclusion. The paper's value depends on accurate citation and attribution, and the skeptic's findings identify genuine correctness problems, such as ViLT being attributed to Radford et al. (2021), DAQUAR being tied to a paper that is not the DAQUAR dataset paper, the Spatial Memory Network being credited to a non-VQA object-detection paper, and the systematic replacement of 'Visual' with 'Video' in later sections. These are reliability and scholarship concerns, not circularity: they do not make the survey's descriptive claims equivalent to their inputs by construction. One author, Asif Ekbal, is a coauthor on some cited works, including the code-mixed VQA system by Khan et al. (2021), but those citations are used as ordinary literature references in the future-directions discussion and are not load-bearing for the survey's organization or any derived conclusion. Machine-checked or externally verifiable support is not needed for a narrative survey, and no self-citation chain is invoked to forbid alternative viewpoints. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The cited references accurately represent the models and datasets they are claimed to describe.
- ad hoc to paper The extractive versus abstractive paradigm dichotomy can meaningfully classify every VQA model discussed.
- domain assumption The paper's narrative accurately reflects the historical sequence and significance of VQA milestones.
Cite this review
Pith. "Pith review of The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering." pith.science (2026). https://pith.science/paper/W6VQWM6M
@misc{pith2026250107109,
author = {Pith},
title = {Pith review of: The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/W6VQWM6M}},
note = {Machine review of arXiv:2501.07109}
}
read the original abstract
Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about images. Since its inception in 2015, VQA has rapidly evolved, driven by advances in deep learning, attention mechanisms, and transformer-based models. This survey traces the journey of VQA from its early days, through major breakthroughs, such as attention mechanisms, compositional reasoning, and the rise of vision-language pre-training methods. We highlight key models, datasets, and techniques that shaped the development of VQA systems, emphasizing the pivotal role of transformer architectures and multimodal pre-training in driving recent progress. Additionally, we explore specialized applications of VQA in domains like healthcare and discuss ongoing challenges, such as dataset bias, model interpretability, and the need for common-sense reasoning. Lastly, we discuss the emerging trends in large multimodal language models and the integration of external knowledge, offering insights into the future directions of VQA. This paper aims to provide a comprehensive overview of the evolution of VQA, highlighting both its current state and potential advancements.
Reference graph
Works this paper leans on
-
[1]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, SM Iftekhar Uddin, Shayekh Bin Islam, et al. 2024. Maya: An instruction finetuned multilingual multimodal model. arXiv e-prints, pages arXiv--2412
2024
-
[4]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...
arXiv 2022
-
[5]
Anderson, X
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR
2018
-
[6]
Andreas, M
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein. 2016. Neural module networks. In ICML
2016
-
[7]
Antol, A
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, and C. L. Lawrence Zitnick. 2015 a . Vqa: Visual question answering. In ICCV
2015
-
[8]
Lawrence Zitnick, and Devi Parikh
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 b . http://arxiv.org/abs/1505.00468 VQA: visual question answering . CoRR, abs/1505.00468
arXiv 2015
Show all 98 references
-
[9]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2016. http://arxiv.org/abs/1409.0473 Neural machine translation by jointly learning to align and translate
2016 arXiv
-
[10]
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2022. http://arxiv.org/abs/2104.00650 Frozen in time: A joint video and image encoder for end-to-end retrieval
2022 arXiv
-
[11]
Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Laila Bashmal, and Mansour Zuair. 2023. https://doi.org/10.3390/bioengineering10030380 Vision–language model for visual question answering in medical imagery . Bioengineering, 10(3)
2023 doi
-
[12]
Hedi Ben-younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017. http://arxiv.org/abs/1705.06676 Mutan: Multimodal tucker fusion for visual question answering
2017 arXiv
-
[13]
Sarath Chandar, Sungjin Ahn, Hugo Larochelle, Pascal Vincent, Gerald Tesauro, and Yoshua Bengio. 2016. http://arxiv.org/abs/1605.07427 Hierarchical memory networks
2016 arXiv
-
[14]
L. Chen, H. Kornblith, M. Swersky, and M. Norouzi. 2020. Uniter: Universal image-text representation learning. In ECCV
2020
-
[15]
Long Chen, Hanwang Jiang, Jin-Hwa Xiao, Shih-Fu Shi, and Shuicheng Chen. 2017. Spatial memory for context reasoning in object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), pages 4086--4096
2017
-
[16]
Marks, and Jonathan Le Roux
Anoop Cherian, Chiori Hori, Tim K. Marks, and Jonathan Le Roux. 2022. http://arxiv.org/abs/2202.09277 (2.5+1)d spatio-temporal scene graphs for video question answering
2022 arXiv
-
[17]
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. http://arxiv.org/abs/1406.1078 Learning phrase representations using rnn encoder-decoder for statistical machine translation
2014 arXiv
-
[18]
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra. 2017. http://arxiv.org/abs/1611.08669 Visual dialog
2017 arXiv
-
[19]
J. Deng, W. Dong, R. Socher, L. Li, and F. Li. 2009. Imagenet: A large-scale hierarchical image database. In CVPR 2009
2009
-
[20]
J. L. Elman. 1990. Finding structure in time. Cognitive Science
1990
-
[21]
Everingham, L
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. 2010. The pascal visual object classes (voc) challenge. In IJCV 2010
2010
-
[22]
Fukui, D
A. Fukui, D. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach. 2016. Multimodal compact bilinear pooling for visual question answering and visual grounding. In CVPR
2016
-
[23]
Kunihiko Fukushima. 1980. https://doi.org/10.1007/BF00344251 Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position . Biological Cybernetics, 36(4):193--202
1980 doi
-
[24]
H. Gao, J. Mao, J. Zhou, T. Huang, L. Xu, and Y. Wang. 2015. Are you talking to a machine? dataset and methods for multilingual image question answering. In CVPR
2015
-
[25]
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes. 2015. https://doi.org/10.1073/pnas.1422953112 Visual turing test for computer vision systems . Proceedings of the National Academy of Sciences, 112(12):3618--3623
2015 doi
-
[26]
Girshick, J
R. Girshick, J. Donahue, T. Darrell, and J. Malik. 2014 a . Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR 2014
2014
-
[27]
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014 b . http://arxiv.org/abs/1311.2524 Rich feature hierarchies for accurate object detection and semantic segmentation
2014 arXiv
-
[28]
Goyal, T
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. 2017. Visual question answering in the wild. In CVPR
2017
-
[29]
Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P
Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. http://arxiv.org/abs/1802.08218 Vizwiz grand challenge: Answering visual questions from blind people
2018 arXiv
-
[30]
Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller
Sadid A. Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller. 2018. Overview of the ImageCLEF 2018 medical domain visual question answering task. In CLEF2018 Working Notes, CEUR Workshop Proceedings, Avignon, France. CEUR-WS.org < http://ceur-ws.org >
2018
-
[31]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. http://arxiv.org/abs/1512.03385 Deep residual learning for image recognition
2015 arXiv
-
[33]
Sepp Hochreiter and Jürgen Schmidhuber. 1997 b . https://doi.org/10.1162/neco.1997.9.8.1735 Long short-term memory . Neural computation, 9:1735--80
1997 doi
-
[34]
D. A. Hudson and C. D. Manning. 2018. Compositional attention networks for machine reasoning. In NeurIPS
2018
-
[35]
Hudson and Christopher D
Drew A. Hudson and Christopher D. Manning. 2019. http://arxiv.org/abs/1902.09506 Gqa: A new dataset for real-world visual reasoning and compositional question answering
2019 arXiv
-
[36]
Johnson, B
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick. 2017 a . Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR 2017
2017
-
[37]
Johnson, B
J. Johnson, B. Hariharan, L. van der Maaten, J. Hoffman, L. Fei-Fei, and C. L. Zitnick. 2017 b . Inferring scene structure and generating stories from images. In CVPR
2017
-
[38]
Kushal Kafle and Christopher Kanan. 2017. http://arxiv.org/abs/1703.09684 An analysis of visual question answering algorithms
2017 arXiv
-
[39]
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos Kadar, Adam Trischler, and Yoshua Bengio. 2018. http://arxiv.org/abs/1710.07300 Figureqa: An annotated figure dataset for visual reasoning
2018 arXiv
-
[40]
Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. Computer Vision and Pattern Recognition (CVPR)
2015
-
[41]
Aisha Urooj Khan, Amir Mazaheri, Niels da Vitoria Lobo, and Mubarak Shah. 2020. http://arxiv.org/abs/2010.14095 Mmft-bert: Multimodal fusion transformer with bert encodings for visual question answering
2020 arXiv
-
[42]
Humair Raj Khan, Deepak Gupta, and Asif Ekbal. 2021. http://arxiv.org/abs/2109.04653 Towards developing a multilingual and code-mixed visual question answering system by knowledge distillation
2021 arXiv
- [43]
-
[44]
J. Kim, Y. Jun, H. Zhang, and J. Kim. 2017. Bilinear attention networks for multimodal learning. In CVPR
2017
-
[45]
J. Kim, Y. Jun, H. Zhang, and J. Kim. 2018. Learning to answer visual questions with attention on attention. In CVPR
2018
-
[46]
Krishna, Y
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, ..., and L. Fei-Fei. 2017. Visual genome: Connecting language and vision using crowdsourced dense annotations. In IJCV, volume 123, pages 32--73
2017
-
[47]
Shamma, Michael S
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. 2016. http://arxiv.org/abs/1602.07332 Visual genome: Connecting language and vision using cr...
2016 arXiv
-
[48]
Krizhevsky, I
A. Krizhevsky, I. Sutskever, and G. E. Hinton. 2012 a . Imagenet classification with deep convolutional neural networks. In NeurIPS 2012
2012
-
[49]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 b . https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf Imagenet classification with deep convolutional neural networks . In Advances in Neural Information Processing...
2012
-
[50]
Jason Joseph Lau, Soumya Gayen, Dina Demner, and Asma Ben Abacha. 2018. https://doi.org/10.17605/OSF.IO/89KPS Visual question answering in radiology (vqa-rad) . Open Science Framework. A dataset of clinically generated visual questions and answers about radiology images
2018 doi
-
[51]
Haopeng Li, Andong Deng, Qiuhong Ke, Jun Liu, Hossein Rahmani, Yulan Guo, Bernt Schiele, and Chen Chen. 2024 a . http://arxiv.org/abs/2401.01505 Sports-qa: A large-scale video question answering benchmark for complex and professional sports
2024
-
[52]
J. Li, H. Tan, L. Wang, M. Yang, and S. C. H. Hoi. 2019. Visualbert: A simple and performant visual language model. In NeurIPS
2019
-
[53]
Jian Li, Weiheng Lu, Hao Fei, Meng Luo, Ming Dai, Min Xia, Yizhang Jin, Zhenye Gan, Ding Qi, Chaoyou Fu, Ying Tai, Wankou Yang, Yabiao Wang, and Chengjie Wang. 2024 b . http://arxiv.org/abs/2408.08632 A survey on benchmarks of multimodal large language models
2024 arXiv
-
[54]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. http://arxiv.org/abs/2301.12597 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
2023 arXiv
-
[55]
X. Li, H. Tan, L. Wang, M. Yang, and S. C. H. Hoi. 2020. Oscar: Object-semantics aware pre-training for vision-language tasks. In CVPR
2020
-
[56]
Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lei Wang, Lingqiao Liu, Leyang Cui, Zhaopeng Tu, Longyue Wang, and Luping Zhou. 2024 c . http://arxiv.org/abs/2310.20381 A systematic evaluation of gpt-4v's multimodal capability for medical image analysis
2024 arXiv
-
[57]
T. Y. Lin, M. Maire, S. Belongie, et al. 2014. Microsoft coco: Common objects in context. In ECCV 2014
2014
-
[58]
Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. https://doi.org/10.1109/CVPR.2015.7298965 Fully convolutional networks for semantic segmentation . In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3431--3440
2015
-
[59]
J. Lu, Z. Yang, H. Mobahi, and D. Parikh. 2019 a . Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of NeurIPS 2019
2019
-
[60]
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 b . http://arxiv.org/abs/1908.02265 Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
2019 arXiv
-
[61]
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017. http://arxiv.org/abs/1612.01887 Knowing when to look: Adaptive attention via a visual sentinel for image captioning
2017 arXiv
-
[62]
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016. Hierarchical question-image co-attention for visual question answering. In Advances in neural information processing systems (NeurIPS), pages 289--297
2016
-
[63]
Sandeep Maddu and Viziananda Row Sanapala. 2024. https://doi.org/10.1145/3695766 A survey on nlp tasks, resources and techniques for low-resource telugu-english code-mixed text . ACM Trans. Asian Low-Resour. Lang. Inf. Process. Just Accepted
2024 doi
-
[64]
Malinowski, M
M. Malinowski, M. Rohrbach, and A. Vedaldi. 2014. Multimodal learning with deep convolutional neural networks. In ICCV
2014
-
[65]
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz. 1993. Building a large annotated corpus of english: The penn treebank. In ACL 1993
1993
-
[66]
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. http://arxiv.org/abs/1906.00067 Ok-vqa: A visual question answering benchmark requiring external knowledge
2019 arXiv
-
[67]
Khapra, and Pratyush Kumar
Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. 2020. http://arxiv.org/abs/1909.00997 Plotqa: Reasoning over scientific plots
2020 arXiv
-
[68]
Mikolov, K
T. Mikolov, K. Chen, G. Corrado, and J. Dean. 2013 a . Efficient estimation of word representations in vector space. In NeurIPS 2013
2013
-
[69]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 b . http://arxiv.org/abs/1301.3781 Efficient estimation of word representations in vector space
2013 arXiv
-
[70]
Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim, and Edward Choi. 2022. https://doi.org/10.1109/jbhi.2022.3207502 Multi-modal understanding and generation for medical images and text via vision-language pre-training . IEEE Journal of Biomedical and Health Informatic...
2022
-
[71]
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual attention networks for multimodal reasoning and matching. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pages 299--307
2017
-
[72]
Narasimhan, P
K. Narasimhan, P. Dixit, and A. Gupta. 2022. Knowledge-augmented neural networks for visual question answering. arXiv preprint arXiv:2203.13843
2022 arXiv
-
[73]
Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. https://doi.org/10.18653/v1/2021.mrl-1.11 Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages . In Proceedings of the 1st Workshop on Multilingual Representation ...
2021 doi
-
[74]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, O. Agrawal, and I. Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of CVPR 2021
2021
-
[75]
Alec Radford and Karthik Narasimhan. 2018. https://api.semanticscholar.org/CorpusID:49313245 Improving language understanding by generative pre-training
2018
-
[76]
Dai, Nissan Hajaj, Michaela Hardt, Peter J
Alvin Rajkomar, Eyal Oren, Kai Chen, Andrew M. Dai, Nissan Hajaj, Michaela Hardt, Peter J. Liu, Xiaobing Liu, Jake Marcus, Mimi Sun, Patrik Sundberg, Hector Yee, Kun Zhang, Yi Zhang, Gerardo Flores, Gavin E Duggan, Jamie Irvine, Quoc V. Le, Kurt Litsch, Alexander Mossin, Justi...
2018
-
[77]
Karen Simonyan and Andrew Zisserman. 2015. http://arxiv.org/abs/1409.1556 Very deep convolutional networks for large-scale image recognition
2015 arXiv
-
[78]
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. http://arxiv.org/abs/1904.08920 Towards vqa models that can read
2019 arXiv
-
[79]
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019. http://arxiv.org/abs/1811.00491 A corpus for reasoning about natural language grounded in photographs
2019 arXiv
-
[80]
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. http://arxiv.org/abs/1904.01766 Videobert: A joint model for video and language representation learning
2019 arXiv
-
[81]
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. http://arxiv.org/abs/1409.3215 Sequence to sequence learning with neural networks
2014 arXiv
-
[82]
Tan and M
H. Tan and M. Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. In Proceedings of IJCNLP 2019
2019
-
[83]
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016. http://arxiv.org/abs/1512.02902 Movieqa: Understanding stories in movies through question-answering
2016 arXiv
-
[84]
Teney, L
D. Teney, L. Shen, L. Demszky, and K. Saenko. 2017. Graph neural networks for visual question answering. In CVPR
2017
-
[85]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. http://arxiv.org/abs/1706.03762 Attention is all you need
2023 arXiv
-
[86]
Vinyals, A
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. 2015. Show and tell: A neural image caption generator. In CVPR 2015
2015
-
[87]
Viola and M
P. Viola and M. Jones. 2001. https://doi.org/10.1109/CVPR.2001.990517 Rapid object detection using a boosted cascade of simple features . In Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. CVPR 2001, volume 1, pages I--I
2001
-
[88]
Min Wang, Ata Mahjoubfar, and Anupama Joshi. 2022. http://arxiv.org/abs/2208.11253 Fashionvqa: A domain-specific visual question answering system
2022 arXiv
-
[89]
Yizhou Wang, Ruiyi Zhang, Haoliang Wang, Uttaran Bhattacharya, Yun Fu, and Gang Wu. 2023. http://arxiv.org/abs/2312.02310 Vaquita: Enhancing alignment in llm-assisted video understanding
2023 arXiv
-
[90]
J. Xiao, J. Hays, K. Ehinger, A. Oliva, and A. Torralba. 2010. Sun database: Large-scale scene recognition from abbey to zoo. In CVPR 2010
2010
-
[91]
Zemel, and Yoshua Bengio
Kelvin Xu, Jimmy Lei Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015. Show, attend and tell: neural image caption generation with visual attention. In Proceedings of the 32nd International Conference on Internatio...
2015
-
[92]
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola. 2016. Stacked attention networks for image question answering. In CVPR
2016
-
[93]
Z. Yang, J. Mao, T. Huang, L. Xu, and Y. Wang. 2018. Dynamic scene graph for visual question answering. In CVPR
2018
-
[94]
K. Yi, J. Shen, Z. Zhang, H. Zhang, and A. Hengel. 2018. Neural-symbolic visual reasoning and generation. In ECCV
2018
-
[95]
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019. http://arxiv.org/abs/1906.10770 Deep modular co-attention networks for visual question answering
2019 arXiv
-
[96]
Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency. 2019. Social-iq: A question answering benchmark for artificial social intelligence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[97]
Zhong et al
Z. Zhong et al. 2022. https://arxiv.org/abs/2203.01225 Videoqa: A comprehensive survey of datasets and methods . arXiv preprint arXiv:2203.01225
2022 arXiv
-
[98]
Corso, and Jianfeng Gao
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2019. http://arxiv.org/abs/1909.11059 Unified vision-language pre-training for image captioning and vqa
2019 arXiv
-
[99]
Bo Zou, Chao Yang, Yu Qiao, Chengbin Quan, and Youjian Zhao. 2024. http://arxiv.org/abs/2404.00973 Videodistill: Language-aware vision distillation for video question answering
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.