REVIEW 5 major objections 6 minor 5 cited by
Visual question answering: from early developments to recent advances -- a survey
T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper's central claim is that a single taxonomy of VQA architectures—organized by vision encoder, language encoder, fusion machine, and answer decoder—can structure the whole field and expose its shift from simple fusion to large…
desk verdict A useful but sloppy VQA survey: structure and LVLM coverage are fine, but citation and transcription errors in the dataset and model tables break its value as a reference until fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the taxonomy shown in Figure 3, which decomposes every VQA system into the same four functional blocks and then subdivides each block into the technique families used in the literature. It does the organizing work of the survey: every model reviewed in the paper is placed in this grid, and Table 3 and Figure 4 translate the taxonomy into quantitative comparisons by reporting each model's accuracy on benchmarks such as VQA v1.0, VQA v2.0, DAQUAR, Visual7W, COCO-QA, CLEVR, and GQA. The taxonomy is what allows the paper to draw timeline conclusions, such as object-based vision encoders giving way to ViT patch encoders and transformer-based fusion displacing earlier attention designs.
What would settle it
Select a random sample of rows from Table 3, locate the cited original papers, and verify the reported accuracy values and dataset names; if a substantial share of entries differ from the originals, the paper's claim to be a reliable VQA resource fails.
Extended reading notes
Core claim
The paper's central claim is that the whole of VQA architecture design can be organized by four components: vision encoder (grid-based, object-based, or ViT patch-based), language encoder (bag-of-words, RNN, CNN, or transformer), fusion machine (simple fusion, attention, bilinear pooling, relation networks, neural module networks, or large visual-language models), and answer decoder (closed-vocabulary or open-vocabulary). Within this taxonomy, the survey charts a historical progression: early models used CNNs and LSTMs with simple fusion, attention-based and bilinear pooling methods then raised accuracy, and after 2019 LVLMs pretrained on image-text data became the dominant approach, with accuracy on VQA v2.0 climbing from about 52 percent in 2015 to above 84 percent by 2022. The authors' stated intent is that this structured framework makes comparative analysis and evaluation straightforward and gives researchers and practitioners a single resource covering methods, datasets, metrics, and applications.
Load-bearing premise
The survey's usefulness as a reference depends on the accuracy numbers in Table 3 and the dataset attributions in Table 2 being transcribed correctly from the cited papers; if those transcriptions are unreliable, the comparison tables cannot be trusted.
Editorial extensions
If this is right
- Researchers can locate any VQA model's design choices on the taxonomy's four axes and compare models that differ in only one component, isolating which part drives accuracy.
- The four-period timeline implies that future advances will come from LVLMs, with ViT-based patch encoders and pretrained large language model backbones dominating the field.
- The accuracy tables imply that transformer-based fusion, especially self-attention and cross-attention, was the strongest fusion family before the LVLM era, with MCAN and its extensions setting the pre-LVLM ceiling on VQA v2.0.
- Because the survey maps datasets by domain, it implies that progress in medical, remote sensing, education, cultural heritage, and advertising VQA tracks dataset creation as much as model design.
- For closed-vocabulary VQA tasks, the decoder taxonomy suggests that open-ended and multiple-choice setups should be treated as distinct problems with distinct evaluation metrics.
Reading between the lines
- If the taxonomy is taken as the field's organizing scheme, an implicit testable prediction is that future state-of-the-art VQA models will be describable as LVLMs with ViT encoders and LLM backbones, making the older fusion categories mostly historical.
- The survey's own section on question relevance reports that a large share of human questions are unrelated to their images; a production VQA system could therefore benefit from an explicit relevance-detection gate before answering, a design the paper does not develop.
- A consequence the authors leave implicit is that Table 3's benchmark rankings conflate architecture choice with model scale; separating those two effects would require controlled comparisons the table does not provide.
- One testable extension would be to use the taxonomy's four axes as features for predicting which new VQA datasets a model will transfer to, treating the survey's classification as input to a prediction study.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes a taxonomy of VQA architectures organized by vision encoder, language encoder, fusion method, and answer decoder, and uses it to structure a review of deep-learning VQA methods, LVLMs, datasets, metrics, applications, and future directions. The paper includes a large table of datasets (Table 2) and a large table of model accuracies (Table 3), and it describes applications in medicine, accessibility, remote sensing, cultural heritage, advertising, and education. The stated contribution is to serve as a comprehensive, up-to-date reference resource for VQA researchers and practitioners as of early 2025.
Significance. If the factual content were reliable, this would be a genuinely useful synthesis: the taxonomy in Figure 3 is reasonable, the coverage of historical periods from 2015 to 2024 is broad, and the treatment of LVLMs as the current dominant paradigm is appropriate. The paper also covers multiple applied domains and gives a clear account of the standard VQA pipeline. However, the survey's main value as a reference depends on the accuracy of its tables and citations, and the manuscript currently contains several concrete traceability failures. Those failures are not cosmetic; they directly undermine the paper's central claim to be a comprehensive resource for lookup and comparison. I therefore evaluate the manuscript as promising but not yet publishable in its present form.
major comments (5)
- [Table 2, rows 37–38] RSVQA-low and RSVQA-high are attributed to Gurari et al. [2018a], but Section 7.3 correctly attributes RSVQA to Lobry et al. [2019] and Lobry et al. [2020]. A reader using Table 2 as a lookup cannot trace the dataset to its primary source, so the table fails its stated reference function.
- [Table 3, Qwen-VL 7B row] The Qwen-VL 7B row cites Masry et al. [2022b] for both the image encoder and the language encoder, but Qwen-VL is introduced in Bai et al. [2023]; Masry et al. is the ChartQA paper. The 78.80 accuracy should be verified against the Qwen-VL report, and the row should cite the correct source.
- [Section 7.4] The sentence 'To overcome the need for expert-generated descriptions, Bongini Brown et al. [2020] used GPT-3 to automatically generate detailed descriptions of artworks' cites a work that does not appear in the bibliography. The closest entry, Brown et al. [2020] 'Language models are few-shot learners', describes no such artwork study, so this claim is currently unsupported and needs a correct reference or removal.
- [Table 2, row 10] AI2D is cited to Sheng et al. [2016], the same source used for Art-VQA in row 9, but the actual AI2D source is Kembhavi et al. [2016], which is cited in the text. This is another instance in which the dataset table provides incorrect provenance for an entry.
- [Section 8.3, Period III] The text states that VQA v2.0 accuracy during Period III ranged from 71% to 78%, but Table 3 lists Florence (80.16), SimVLM (80.34), and VLMo (82.78) with 2021 dates inside that period. The text and table contradict each other; the range should be corrected or the period boundaries clarified.
minor comments (6)
- [Section 3.3.3] The example question is 'What is the object on the refrigerator?', but the described reasoning step says 'identifying the object on the table'; this internal inconsistency should be fixed.
- [Section 3.3.1] The sentence 'simple fusion techniques need to provide insight into how the visual and textual inputs are combined' appears to mean the opposite; it should say 'fail to provide insight' or 'provide no insight'.
- [Figure 4] The axis labels are garbled, for example 'VI515/0(20 - 08)017/2 7/08(201 - 8)19/020 /089201( - )1/08202 1/08(202 - )222/120'; the figure should be regenerated with legible period labels.
- [Section 4.2] The text says questions 'must begin with one of six letters: what, where, how, why, who, and when', but these are interrogative words rather than letters; the wording should be corrected to 'six question categories' or similar.
- [References] The bibliography contains duplicate entries (e.g., Antol et al. 2015a/2015b, Devlin et al. 2018/2019, Liu et al. 2019a/2019b) and a malformed author string in the Masry et al. [2022b] entry ('andbai2023qwen Villavicencio'); these need cleanup.
- [Section 5] The final sentence about the SimpsonsVQA dataset appears without citation or connection to the surrounding discussion of question relevance; it should be integrated or removed.
Circularity Check
No significant circularity: the survey's contributions are descriptive attributions to external prior work, with no derivation or fitted prediction that reduces to its own inputs.
full rationale
This is a literature survey, not a derivation-driven paper. Its central claim is that it introduces a taxonomy of VQA architectures and organizes existing datasets, models, and metrics. A taxonomy is a categorization scheme, not a result derived from equations or fitted parameters; the survey does not claim to predict any quantity from its own assumptions. All technical claims are attributed to external references, and I found no self-citation chain that is load-bearing: the authors do not appear to rely on their own prior work to justify the taxonomy or any other central premise. The paper contains factual traceability problems that a reader should weigh under correctness risk, not circularity: Table 2 attributes RSVQA to Gurari et al. rather than Lobry et al., Section 7.4 cites a non-existent 'Bongini Brown et al.' for a GPT-3 artwork-description study, and Table 3's Qwen-VL row cites the ChartQA paper as the model source instead of the Qwen technical report. These are citation/attribution errors that reduce the survey's usefulness as a reference, but they do not make any claim definitionally equivalent to its input. Because the survey is self-contained as a review and makes no circular derivation, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The cited papers are accurately represented in the survey text and tables
- domain assumption The selected papers form a representative and complete sample of the VQA literature as of early 2025
Cite this review
Pith. "Pith review of Visual question answering: from early developments to recent advances -- a survey." pith.science (2026). https://pith.science/paper/QRQR2TJZ
@misc{pith2026250103939,
author = {Pith},
title = {Pith review of: Visual question answering: from early developments to recent advances -- a survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/QRQR2TJZ}},
note = {Machine review of arXiv:2501.03939}
}
read the original abstract
Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text embedding, natural language understanding, and language generation. With the growth of multimodal data research, VQA has gained significant attention due to its broad applications, including interactive educational tools, medical image diagnosis, customer service, entertainment, and social media captioning. Additionally, VQA plays a vital role in assisting visually impaired individuals by generating descriptive content from images. This survey introduces a taxonomy of VQA architectures, categorizing them based on design choices and key components to facilitate comparative analysis and evaluation. We review major VQA approaches, focusing on deep learning-based methods, and explore the emerging field of Large Visual Language Models (LVLMs) that have demonstrated success in multimodal tasks like VQA. The paper further examines available datasets and evaluation metrics essential for measuring VQA system performance, followed by an exploration of real-world VQA applications. Finally, we highlight ongoing challenges and future directions in VQA research, presenting open questions and potential areas for further development. This survey serves as a comprehensive resource for researchers and practitioners interested in the latest advancements and future
Figures
Forward citations
Cited by 5 Pith papers
-
Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models
Training-free SpecFlow condenses VLM visual tokens via kNN heat diffusion, adaptive quadtree budgets, and coreset sinks, retaining 95.6% LLaVA-1.5 performance after pruning 88.9% of tokens.
-
Affordance Benchmark for MLLMs
A new 2,000-question benchmark finds multimodal AI models recognize object affordances far worse than humans, with top model Gemini-2.0-Pro at 18.05% versus 85.34% human best.
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.
-
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
Four open-source VQA models reach moderate accuracy on a new classroom video dataset, with yes/no questions easiest and counting/reasoning hardest.
-
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
A staged knowledge-prompting method (KAAR) improves LLM test accuracy on ARC by about 5 absolute points over repeated-sampling plan-guided code generation, reaching 35% with GPT-o3-mini.
Reference graph
Works this paper leans on
-
[1]
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41 0 (2): 0 423--443, 2019. doi:10.1109/TPAMI.2018.2798607
arXiv 2019
-
[2]
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi. Attention-based multimodal fusion for video description. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 4203--4212, 2017. doi:10.1109/ICCV.2017.450
-
[3]
Multimodal language analysis in the wild: CMU - MOSEI dataset and interpretable dynamic fusion graph
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. Multimodal language analysis in the wild: CMU - MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236--2246, Melbourne, Australia, July...
-
[4]
Merlot reserve: Neural script knowledge through vision and language and sound
Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. Merlot reserve: Neural script knowledge through vision and language and sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16375--16387, June 2022
2022
-
[5]
Multimodal learning with graphs
Yasha Ektefaie, George Dasoulas, Ayush Noori, Maha Farhat, and Marinka Zitnik. Multimodal learning with graphs. Nature Machine Intelligence, pages 1--11, 2023
2023
-
[6]
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904, 2021
arXiv 2021
-
[7]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022
arXiv 2022
-
[8]
Visual language integration: A survey and open challenges
Sang-Min Park and Young-Gab Kim. Visual language integration: A survey and open challenges. Computer Science Review, 48: 0 100548, 2023. ISSN 1574-0137. doi:https://doi.org/10.1016/j.cosrev.2023.100548. URL https://www.sciencedirect.com/science/article/pii/S1574013723000151
arXiv 2023
Show all 264 references
-
[9]
Lawrence Zitnick, and Devi Parikh
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), December 2015 a
2015
-
[10]
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel. Exploring models and data for image question answering. Advances in neural information processing systems, 28, 2015 a
2015
-
[11]
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Yin and yang: Balancing and answering binary visual questions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016
2016
-
[12]
Learning to answer questions from image using convolutional neural network
Lin Ma, Zhengdong Lu, and Hang Li. Learning to answer questions from image using convolutional neural network. In Thirtieth AAAI Conference on Artificial Intelligence, 2016
2016
-
[13]
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017 a
2017
-
[14]
Lawrence Zitnick, and Devi Parikh
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA : V isual Q uestion A nswering. In International Conference on Computer Vision (ICCV), 2015 b
2015
-
[15]
Shamma, Michael S
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li - Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei - Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations...
2016 arXiv
-
[16]
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2: 0 67--78, 2014. doi:10.1162/tacl...
2014 doi
-
[17]
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR...
2017 doi
-
[18]
Medical visual question answering: A survey
Zhihong Lin, Donghao Zhang, Qingyi Tac, Danli Shi, Gholamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. Medical visual question answering: A survey. arXiv preprint arXiv:2111.10056, 2021
2021 arXiv
-
[19]
Cross-modal self-attention with multi-task pre-training for medical visual question answering
Haifan Gong, Guanqi Chen, Sishuo Liu, Yizhou Yu, and Guanbin Li. Cross-modal self-attention with multi-task pre-training for medical visual question answering. In Proceedings of the 2021 International Conference on Multimedia Retrieval, ICMR '21, page 456–460, New York, NY, US...
2021
-
[20]
A sequence-to-sequence model approach for imageclef 2018 medical domain visual question answering
Rahul Ambati and Chakravardhan Reddy Dudyala. A sequence-to-sequence model approach for imageclef 2018 medical domain visual question answering. In 2018 15th IEEE India Council International Conference (INDICON), pages 1--6, 2018. doi:10.1109/INDICON45594.2018.8987108
2018
-
[21]
An encoder-decoder model for visual question answering in the medical domain
Imane Allaouzi, Mohamed Ben Ahmed, and Badr Benamrou. An encoder-decoder model for visual question answering in the medical domain. In Conference and Labs of the Evaluation Forum, 2019
2019
-
[22]
Sysu-hcp at vqa-med 2021: A data-centric model with efficient training methodology for medical visual question answering
Haifan Gong, Ricong Huang, Guanqi Chen, and Guanbin Li. Sysu-hcp at vqa-med 2021: A data-centric model with efficient training methodology for medical visual question answering. In CLEF, 2021 b
2021
-
[23]
Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller
Sadid A. Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller. Overview of the ImageCLEF 2018 medical domain visual question answering task. In CLEF2018 Working Notes, CEUR Workshop Proceedings, Avignon, France, September 10-14 2018 a . CEUR-WS.or...
2018
-
[24]
Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman
Jason J. Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific Data, 5, 2018 a
2018
-
[25]
Hasan, and Henning M\"uller
Asma Ben Abacha , Mourad Sarrouti, Dina Demner-Fushman, Sadid A. Hasan, and Henning M\"uller. Overview of the vqa-med task at imageclef 2021: Visual question answering and generation in the medical domain. In CLEF 2021 Working Notes, CEUR Workshop Proceedings, Bucharest, Roman...
2021
-
[26]
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. Iqa: Visual question answering in interactive environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018
2018
-
[27]
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015
2015
-
[28]
Yuille, and Kevin Murphy
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016
2016
-
[29]
Transvg: End-to-end visual grounding with transformers
Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li. Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1769--1779, October 2021
2021
-
[30]
Hierarchical lstms with adaptive attention for visual captioning
Lianli Gao, Xiangpeng Li, Jingkuan Song, and Heng Tao Shen. Hierarchical lstms with adaptive attention for visual captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42 0 (5): 0 1112--1131, 2020. doi:10.1109/TPAMI.2019.2894139
2020
-
[31]
Semantic compositional networks for visual captioning
Zhe Gan, Chuang Gan, Xiaodong He, Yunchen Pu, Kenneth Tran, Jianfeng Gao, Lawrence Carin, and Li Deng. Semantic compositional networks for visual captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017
2017
-
[32]
Information fusion in visual question answering: A survey
Dongxiang Zhang, Rui Cao, and Sai Wu. Information fusion in visual question answering: A survey. Information Fusion, 52: 0 268--280, 2019
2019
-
[33]
Visual question answering: methodologies and challenges
Liyana Sahir Kallooriyakath, MV Jithin, PV Bindu, and PP Adith. Visual question answering: methodologies and challenges. In 2020 International Conference on Smart Technologies in Computing, Electrical and Electronics (ICSTCEE), pages 402--407. IEEE, 2020
2020
-
[34]
Visual question answering: a state-of-the-art review
Sruthy Manmadhan and Binsu C Kovoor. Visual question answering: a state-of-the-art review. Artificial Intelligence Review, 53: 0 5705--5745, 2020
2020
-
[35]
A survey on vqa: Datasets and approaches
Yeyun Zou and Qiyu Xie. A survey on vqa: Datasets and approaches. In 2020 2nd International Conference on Information Technology and Computer Application (ITCA), pages 289--297. IEEE, 2020
2020
-
[36]
A survey of methods, datasets and evaluation metrics for visual question answering
Himanshu Sharma and Anand Singh Jalal. A survey of methods, datasets and evaluation metrics for visual question answering. Image and Vision Computing, 116: 0 104327, 2021
2021
-
[37]
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019
2019
-
[38]
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019 a
1908 arXiv
-
[39]
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530, 2019
1908 arXiv
-
[40]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317--8326, 2019
2019
-
[41]
Tag: Boosting text-vqa via text-aware visual question-answer generation
Jun Wang, Mingfei Gao, Yuqian Hu, Ramprasaath R Selvaraju, Chetan Ramaiah, Ran Xu, Joseph F JaJa, and Larry S Davis. Tag: Boosting text-vqa via text-aware visual question-answer generation. arXiv preprint arXiv:2208.01813, 2022 a
2022 arXiv
-
[42]
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 947--952, 2019. doi:10.1109/ICDAR.2019.00156
2019
-
[43]
Visualmrc: Machine reading comprehension on document images
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. Visualmrc: Machine reading comprehension on document images. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13878--13888, 2021
2021
-
[44]
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Mar c al Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. Scene text visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4291--4301, 2019
2019
-
[45]
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200--2209, 2021
2021
-
[46]
Explicit knowledge-based reasoning for visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick. Explicit knowledge-based reasoning for visual question answering. arXiv preprint arXiv:1511.02570, 2015
2015 arXiv
-
[47]
Ask me anything: Free-form visual question answering based on knowledge from external sources
Qi Wu, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. Ask me anything: Free-form visual question answering based on knowledge from external sources. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4622--4630, 2016
2016
-
[48]
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pages 3195--3204, 2019
2019
-
[49]
A-okvqa: A benchmark for visual question answering using world knowledge
Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-okvqa: A benchmark for visual question answering using world knowledge. In Computer Vision--ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23--27, 2022, Proceedings, P...
2022
-
[50]
Grounding answers for visual questions asked by visually impaired people
Chongyan Chen, Samreen Anjum, and Danna Gurari. Grounding answers for visual questions asked by visually impaired people. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19098--19107, 2022 a
2022
-
[51]
WSDM Cup 2023 Challenge on Visual Question Answering
Dmitry Ustalov, Nikita Pavlichenko, Daniil Likhobaba, and Alisa Smirnova. WSDM Cup 2023 Challenge on Visual Question Answering . In Proceedings of the 4th Crowd Science Workshop on Collaboration of Humans and Learning Algorithms for Data Labeling, pages 1--7, Singapore, 2023. ...
2023
-
[52]
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. Flava: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15638--1...
2022
-
[53]
Masked vision and language modeling for multi-modal representation learning
Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, and Stefano Soatto. Masked vision and language modeling for multi-modal representation learning. arXiv preprint arXiv:2208.02131, 2022
2022 arXiv
-
[54]
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
2023
-
[55]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll \'a r, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000--16009, 2022
2022
-
[56]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...
2021
-
[57]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904--4916...
2021
-
[58]
Sigmoid loss for language image pre-training, 2023
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training, 2023
2023
-
[59]
Chameleon: Mixed-modal early-fusion foundation models, 2024
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models, 2024. URL https://arxiv.org/abs/2405.09818
2024 arXiv
-
[60]
Gpt-4 technical report, 2023
OpenAI. Gpt-4 technical report, 2023
2023
-
[61]
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning...
2022
-
[62]
Visual instruction tuning, 2023 a
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023 a
2023
-
[63]
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models, 2021. URL https://arxiv.org/abs/2106.13884
2021 arXiv
-
[64]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023 a
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023 a . URL https://arxiv.org/abs/2301.12597
2023 arXiv
-
[65]
Qwen technical report, 2023
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Si...
2023 arXiv
-
[66]
A simple neural network module for relational reasoning
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. A simple neural network module for relational reasoning. Advances in neural information processing systems, 30, 2017
2017
-
[67]
Graph-structured representations for visual question answering
Damien Teney, Lingqiao Liu, and Anton van Den Hengel. Graph-structured representations for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1--9, 2017
2017
-
[68]
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv preprint arXiv:1606.01847, 2016
2016 arXiv
-
[69]
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, R \'e mi Cadene, Matthieu Cord, and Nicolas Thome. Mutan: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2612--2620, 2017
2017
-
[70]
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural module networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 39--48, 2016
2016
-
[71]
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. Learning to reason: End-to-end module networks for visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 804--813, 2017
2017
-
[72]
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6281--6290, 2019
2019
-
[73]
An improved attention for visual question answering
Tanzila Rahman, Shih-Han Chou, Leonid Sigal, and Giuseppe Carenini. An improved attention for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1653--1662, 2021
2021
-
[74]
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems, 29, 2016
2016
-
[75]
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. Dual attention networks for multimodal reasoning and matching. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 299--307, 2017
2017
-
[76]
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016
2016
-
[77]
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 60...
2018
-
[78]
Ask your neurons: A neural-based approach to answering questions about images
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz. Ask your neurons: A neural-based approach to answering questions about images. In Proceedings of the IEEE international conference on computer vision, pages 1--9, 2015
2015
-
[79]
BERT : Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Lan...
2019 doi
-
[81]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023 a
2023 arXiv
-
[82]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90\ quality, March 2023. URL https://lmsys.org/blog/2023...
2023
-
[83]
Long short-term memory
J \"u rgen Schmidhuber, Sepp Hochreiter, et al. Long short-term memory. Neural Comput, 9 0 (8): 0 1735--1780, 1997
1997
-
[84]
Schuster and K.K
M. Schuster and K.K. Paliwal. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45 0 (11): 0 2673--2681, 1997. doi:10.1109/78.650093
1997 doi
-
[85]
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, C aglar G \" u l c ehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. CoRR, abs/1406.1078, 2014. URL http://arxiv.org/abs/1406.1078
2014 arXiv
-
[86]
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013
2013 arXiv
-
[87]
G lo V e: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. G lo V e: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 1532--1543, Doha, Qatar, October 2014. Association for Com...
2014 doi
-
[88]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv...
2010 arXiv
-
[89]
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016
2016
-
[90]
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran ...
2015
-
[91]
Spatial pyramid pooling in deep convolutional networks for visual recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37 0 (9): 0 1904--1916, 2015. doi:10.1109/TPAMI.2015.2389824
1904
-
[92]
LeCun, B
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1 0 (4): 0 541--551, 1989. doi:10.1162/neco.1989.1.4.541
1989 doi
-
[93]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248--255, 2009. doi:10.1109/CVPR.2009.5206848
2009
-
[94]
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedin...
2015 arXiv
-
[95]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770--778, 2016. doi:10.1109/CVPR.2016.90
2016 doi
-
[96]
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1--9, 2015. doi:1...
2015
-
[97]
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen. In defense of grid features for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[98]
W ea QA : Weak supervision via captions for visual question answering
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, and Chitta Baral. W ea QA : Weak supervision via captions for visual question answering. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3420--3435, Online, August 2021. Association for Computat...
2021 doi
-
[99]
Question-guided feature pyramid network for medical visual question answering
Yonglin Yu, Haifeng Li, Hanrong Shi, Lin Li, and Jun Xiao. Question-guided feature pyramid network for medical visual question answering. Expert Systems with Applications, 214: 0 119148, 2023 a . ISSN 0957-4174. doi:https://doi.org/10.1016/j.eswa.2022.119148. URL https://www.s...
2023
-
[100]
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and koray kavukcuoglu. Spatial transformer networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL https...
2015
-
[101]
Principe
Ryan Burt, Mihael Cudic, and Jose C. Principe. Fusing attention with visual question answering. In 2017 International Joint Conference on Neural Networks (IJCNN), pages 949--953, 2017. doi:10.1109/IJCNN.2017.7965954
2017
-
[102]
From images to textual prompts: Zero-shot VQA with frozen large language models, 2023
Jiaxian Guo, Junnan Li, Dongxu Li, Anthony Tiong, Boyang Li, Dacheng Tao, and Steven HOI. From images to textual prompts: Zero-shot VQA with frozen large language models, 2023. URL https://openreview.net/forum?id=Ck1UtnVukP8
2023
-
[103]
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct 2017
2017
-
[104]
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583--5594. PMLR, 2021
2021
-
[105]
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, and Furu Wei. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. arXiv preprint arXiv:2111.02358, 2021
2021 arXiv
-
[106]
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li. Multi-grained vision language pre-training: Aligning texts with visual concepts. Proceedings of The 33rd International Conference on Machine Learning, 2022 a
2022
-
[107]
X2-vlm: All-in-one pre-trained model for vision-language tasks
Yan Zeng, Xinsong Zhang, Hang Li, Jiawei Wang, Jipeng Zhang, and Wangchunshu Zhou. X2-vlm: All-in-one pre-trained model for vision-language tasks. arXiv preprint arXiv:2211.12402, 2022 b
2022 arXiv
-
[108]
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34: 0 9694--9705, 2021
2021
-
[109]
Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision ...
2023
-
[110]
Midas v3
Reiner Birkl, D Wofk, and M M \"u ller. Midas v3. 1--a model zoo for robust monocular relative depth estimation. arxiv. arXiv preprint arXiv:2307.14460, 2023
2023 arXiv
-
[111]
Moai: Mixture of all intelligence for large language and vision models
Byung-Kwan Lee, Beomchan Park, Chae Won Kim, and Yong Man Ro. Moai: Mixture of all intelligence for large language and vision models. arXiv preprint arXiv:2403.07508, 2024
2024 arXiv
-
[112]
Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
2024
-
[113]
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Skip-thought vectors. Advances in neural information processing systems, 28, 2015
2015
-
[114]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017
2017
-
[115]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018
2018 arXiv
-
[116]
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019 b
1907 arXiv
-
[117]
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020
2010 arXiv
-
[118]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023 b
2023 arXiv
-
[119]
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten. Revisiting visual question answering baselines. In European conference on computer vision, pages 727--739. Springer, 2016
2016
-
[120]
Dualnet: Domain-invariant network for visual question answering
Kuniaki Saito, Andrew Shin, Yoshitaka Ushiku, and Tatsuya Harada. Dualnet: Domain-invariant network for visual question answering. In 2017 IEEE International Conference on Multimedia and Expo (ICME), pages 829--834. IEEE, 2017
2017
-
[121]
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. Are you talking to a machine? dataset and methods for multilingual image question. Advances in neural information processing systems, 28, 2015
2015
-
[122]
Answer-type prediction for visual question answering
Kushal Kafle and Christopher Kanan. Answer-type prediction for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4976--4984, 2016
2016
-
[123]
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014
2014 arXiv
-
[124]
Multimodal residual learning for visual qa
Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Multimodal residual learning for visual qa. Advances in neural information processing systems, 29, 2016 a
2016
-
[125]
End-to-end memory networks
Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al. End-to-end memory networks. Advances in neural information processing systems, 28, 2015
2015
-
[126]
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Huijuan Xu and Kate Saenko. Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. In European conference on computer vision, pages 451--466. Springer, 2016
2016
-
[127]
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher. Dynamic memory networks for visual and textual question answering. In Maria Florina Balcan and Kilian Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Ma...
2016
-
[128]
Ask me anything: Dynamic memory networks for natural language processing
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. Ask me anything: Dynamic memory networks for natural language processing. In International conference on machine learning, pages 1378--1387....
2016
-
[129]
Where to look: Focus regions for visual question answering
Kevin J Shih, Saurabh Singh, and Derek Hoiem. Where to look: Focus regions for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4613--4621, 2016
2016
-
[130]
A focused dynamic attention model for visual question answering
Ilija Ilievski, Shuicheng Yan, and Jiashi Feng. A focused dynamic attention model for visual question answering. arXiv preprint arXiv:1604.01485, 2016
2016 arXiv
-
[131]
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton Van Den Hengel. Tips and tricks for visual question answering: Learnings from the 2017 challenge. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4223--4232, 2018
2017
-
[132]
From pixels to objects: Cubic visual attention for visual question answering
Jingkuan Song, Pengpeng Zeng, Lianli Gao, and Heng Tao Shen. From pixels to objects: Cubic visual attention for visual question answering. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18 , pages 906--912. International J...
2018 doi
-
[133]
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. Bilinear attention networks. Advances in neural information processing systems, 31, 2018
2018
-
[134]
Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering
Duy-Kien Nguyen and Takayuki Okatani. Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6087--6096, 2018
2018
-
[135]
Attention on attention for image captioning
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. Attention on attention for image captioning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4634--4643, 2019
2019
-
[136]
Trar: Routing the attention spans in transformer for visual question answering
Yiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu, and Rongrong Ji. Trar: Routing the attention spans in transformer for visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20...
2021
-
[137]
Txt: Crossmodal end-to-end learning with transformers
Jan-Martin O Steitz, Jonas Pfeiffer, Iryna Gurevych, and Stefan Roth. Txt: Crossmodal end-to-end learning with transformers. In Pattern Recognition: 43rd DAGM German Conference, DAGM GCPR 2021, Bonn, Germany, September 28--October 1, 2021, Proceedings, pages 405--420. Springer, 2022
2021
-
[138]
Inferring and executing programs for visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Inferring and executing programs for visual reasoning. In Proceedings of the IEEE international conference on computer vision, pages 2989--2998, 2017 a
2017
-
[139]
Explainable neural computation via stack neural module networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko. Explainable neural computation via stack neural module networks. In Proceedings of the European conference on computer vision (ECCV), pages 53--69, 2018
2018
-
[140]
Neural-symbolic vqa: Disentangling reasoning from vision and language understanding
Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. Neural-symbolic vqa: Disentangling reasoning from vision and language understanding. Advances in neural information processing systems, 31, 2018
2018
-
[141]
Probabilistic neural symbolic models for interpretable visual question answering
Ramakrishna Vedantam, Karan Desai, Stefan Lee, Marcus Rohrbach, Dhruv Batra, and Devi Parikh. Probabilistic neural symbolic models for interpretable visual question answering. In International Conference on Machine Learning, pages 6428--6437. PMLR, 2019
2019
-
[142]
Finding frequent items in data streams
Moses Charikar, Kevin Chen, and Martin Farach-Colton. Finding frequent items in data streams. In International Colloquium on Automata, Languages, and Programming, pages 693--703. Springer, 2002
2002
-
[143]
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Hadamard product for low-rank bilinear pooling. arXiv preprint arXiv:1610.04325, 2016 b
2016 arXiv
-
[144]
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 1821--1830, 2017
2017
-
[145]
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018
2018
-
[146]
Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection
Hedi Ben-Younes, Remi Cadene, Nicolas Thome, and Matthieu Cord. Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8102--8109, 2019
2019
-
[147]
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and Fujie Huang. A tutorial on energy-based learning. Predicting structured data, 1 0 (0), 2006
2006
-
[148]
Slip: Self-supervision meets language-image pre-training
Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie. Slip: Self-supervision meets language-image pre-training. In European conference on computer vision, pages 529--544. Springer, 2022
2022
-
[149]
Modeling caption diversity in contrastive vision-language pretraining, 2024
Samuel Lavoie, Polina Kirichenko, Mark Ibrahim, Mahmoud Assran, Andrew Gordon Wilson, Aaron Courville, and Nicolas Ballas. Modeling caption diversity in contrastive vision-language pretraining, 2024
2024
-
[150]
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019
1908 arXiv
-
[151]
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022
2022 arXiv
-
[152]
Fast high-resolution image synthesis with latent adversarial diffusion distillation
Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high-resolution image synthesis with latent adversarial diffusion distillation. arXiv preprint arXiv:2403.12015, 2024
2024 arXiv
-
[153]
Stable diffusion 3: Research paper
Stability.ai. Stable diffusion 3: Research paper. https://stability.ai/news/stable-diffusion-3-research-paper, 2024
2024
-
[154]
meta-llama-3, 2023
Meta. meta-llama-3, 2023. URL https://ai.meta.com/blog/meta-llama-3
2023
-
[155]
Falcon 2
TII. Falcon 2. In European conference on computer vision, 2024. URL https://falconllm.tii.ae/
2024
-
[156]
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a . URL https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
-
[157]
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023
2023 arXiv
-
[158]
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer ...
2017
-
[159]
A dataset for multimodal question answering in the cultural heritage domain
Shurong Sheng, Luc Van Gool, and Marie-Francine Moens. A dataset for multimodal question answering in the cultural heritage domain. In Proceedings of the COLING 2016 Workshop on Language Technology Resources and Tools for Digital Humanities (LT4DH), pages 10--17. ACL, 2016
2016
-
[160]
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. Fvqa: Fact-based visual question answering. IEEE transactions on pattern analysis and machine intelligence, 40 0 (10): 0 2413--2427, 2017
2017
-
[161]
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz. A multi-world approach to question answering about real-world scenes based on uncertain input. Advances in neural information processing systems, 27, 2014
2014
-
[162]
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. Visual7w: Grounded question answering in images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4995--5004, 2016
2016
-
[163]
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904--6913, 2017 b
2017
-
[164]
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognitio...
2017
-
[165]
Don't just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. Don't just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4971--4980, 2018
2018
-
[166]
Automatic understanding of image and video advertisements
Zaeem Hussain, Mingda Zhang, Xiaozhong Zhang, Keren Ye, Christopher Thomas, Zuha Agha, Nathan Ong, and Adriana Kovashka. Automatic understanding of image and video advertisements. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1705--1715, 2017
2017
-
[167]
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern rec...
2017
-
[168]
Figureqa: An annotated figure dataset for visual reasoning
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, \'A kos K \'a d \'a r, Adam Trischler, and Yoshua Bengio. Figureqa: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300, 2017
2017 arXiv
-
[169]
Overview of imageclef 2018 medical domain visual question answering task
Sadid A Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Henning M \"u ller, and Matthew P Lungren. Overview of imageclef 2018 medical domain visual question answering task. In CLEF (Working Notes), 2018 b
2018
-
[170]
Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P
Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3608--3617, 201...
2018
-
[171]
Tallyqa: Answering complex counting questions
Manoj Acharya, Kushal Kafle, and Christopher Kanan. Tallyqa: Answering complex counting questions. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8076--8084, 2019
2019
-
[172]
Dvqa: Understanding data visualizations via question answering
Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. Dvqa: Understanding data visualizations via question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5648--5656, 2018
2018
-
[173]
Vqa-med: Overview of the medical visual question answering task at imageclef 2019
Asma Ben Abacha, Sadid A Hasan, Vivek V Datla, Joey Liu, Dina Demner-Fushman, and Henning M \"u ller. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. CLEF (working notes), 2 0 (6), 2019
2019
-
[174]
On the general value of evidence, and bilingual scene-text visual question answering
Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng, Canjie Luo, Lianwen Jin, Chee Seng Chan, Anton van den Hengel, and Liangwei Wang. On the general value of evidence, and bilingual scene-text visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vi...
2020
-
[175]
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700--6709, 2019 a
2019
-
[176]
Leaf-qa: Locate, encode & attend for figure question answering
Ritwick Chaudhry, Sumit Shekhar, Utkarsh Gupta, Pranav Maneriker, Prann Bansal, and Ajay Joshi. Leaf-qa: Locate, encode & attend for figure question answering. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3512--3521, 2020
2020
-
[177]
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019
2019
-
[178]
A dataset and baselines for visual question answering on art
Noa Garcia, Chentao Ye, Zihua Liu, Qingtao Hu, Mayu Otani, Chenhui Chu, Yuta Nakashima, and Teruko Mitamura. A dataset and baselines for visual question answering on art. In Computer Vision--ECCV 2020 Workshops: Glasgow, UK, August 23--28, 2020, Proceedings, Part II 16, pages ...
2020
-
[179]
Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain
Asma Ben Abacha, Vivek V Datla, Sadid A Hasan, Dina Demner-Fushman, and Henning M \"u ller. Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain. In CLEF (Working Notes), 2020
2020
-
[180]
Towards visual dialog for radiology
Olga Kovaleva, Chaitanya Shivade, Satyananda Kashyap, Karina Kanjaria, Joy Wu, Deddeh Ballah, Adam Coy, Alexandros Karargyris, Yufan Guo, David Beymer Beymer, et al. Towards visual dialog for radiology. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Proce...
2020
-
[181]
Pathvqa: 30000+ questions for medical visual question answering
Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286, 2020 a
2003 arXiv
-
[182]
Khapra, and Pratyush Kumar
Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. Plotqa: Reasoning over scientific plots. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), March 2020
2020
-
[183]
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in neural information processing systems, 33: 0 2611--2624, 2020
2020
-
[184]
Slake: a semantically-labeled knowledge-enhanced dataset for medical visual question answering
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: a semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1650--1654. IEEE, 2021 a
2021
-
[185]
Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning
Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric P Xing, and Liang Lin. Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning. arXiv preprint arXiv:2105.14517, 2021
2021 arXiv
-
[186]
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Pan Lu, Liang Qiu, Jiaqi Chen, Tony Xia, Yizhou Zhao, Wei Zhang, Zhou Yu, Xiaodan Liang, and Song-Chun Zhu. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. In The 35th Conference on Neural Information Processing Systems (NeurIPS) Track...
2021
-
[187]
Clevr-math: A dataset for compositional language, visual and mathematical reasoning
Adam Dahlgren Lindstr \"o m and Savitha Sam Abraham. Clevr-math: A dataset for compositional language, visual and mathematical reasoning. arXiv preprint arXiv:2208.05358, 2022
2022 arXiv
-
[188]
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35: 0 2507...
2022
-
[190]
Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023 b
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023 b
2023 arXiv
-
[191]
Seed-bench: Benchmarking multimodal llms with generative comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023 b
2023 arXiv
-
[192]
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023 b
2023 arXiv
-
[193]
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A m...
2024
-
[194]
Mm-vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linfeng Ren, Linjie Li, Jianfeng Wang, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Lijuan Wang, and Xinchao Wang. Mm-vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities. arXiv preprint arXiv:2408.00765, 2024
2024 arXiv
-
[195]
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023
-
[196]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740--755. Springer, 2014
2014
-
[197]
Unanswerable questions about images and texts
Ernest Davis. Unanswerable questions about images and texts. Frontiers in Artificial Intelligence, 3: 0 51, 2020
2020
-
[198]
Question relevance in vqa: identifying non-visual and false-premise questions
Arijit Ray, Gordon Christie, Mohit Bansal, Dhruv Batra, and Devi Parikh. Question relevance in vqa: identifying non-visual and false-premise questions. arXiv preprint arXiv:1606.06622, 2016
2016 arXiv
-
[199]
Why did the chicken cross the road? rephrasing and analyzing ambiguous questions in vqa
Elias Stengel-Eskin, Jimena Guallar-Blasco, Yi Zhou, and Benjamin Van Durme. Why did the chicken cross the road? rephrasing and analyzing ambiguous questions in vqa. arXiv preprint arXiv:2211.07516, 2022
2022 arXiv
-
[200]
Why does a visual question have different answers? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4271--4280, 2019
Nilavra Bhattacharya, Qing Li, and Danna Gurari. Why does a visual question have different answers? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4271--4280, 2019
2019
-
[201]
Know what you don't know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don't know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822, 2018
2018 arXiv
-
[202]
Document understanding dataset and evaluation (dude)
Jordy Van Landeghem, Rub \`e n Tito, ukasz Borchmann, Micha Pietruszka, Pawel Joziak, Rafal Powalski, Dawid Jurkiewicz, Micka \"e l Coustaty, Bertrand Anckaert, Ernest Valveny, et al. Document understanding dataset and evaluation (dude). In Proceedings of the IEEE/CVF Internat...
2023
-
[203]
Question part relevance and editing for cooperative and context-aware vqa (c2vqa)
Andeep S Toor, Harry Wechsler, and Michele Nappi. Question part relevance and editing for cooperative and context-aware vqa (c2vqa). In Proceedings of the 15th International Workshop on Content-Based Multimedia Indexing, pages 1--6, 2017
2017
-
[204]
Do explanations make vqa models more predictable to a human? arXiv preprint arXiv:1810.12366, 2018
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. Do explanations make vqa models more predictable to a human? arXiv preprint arXiv:1810.12366, 2018
2018 arXiv
-
[205]
Robust visual question answering via semantic cross modal augmentation
Akib Mashrur, Wei Luo, Nayyar A Zaidi, and Antonio Robles-Kelly. Robust visual question answering via semantic cross modal augmentation. Computer Vision and Image Understanding, page 103862, 2023
2023
-
[206]
The promise of premise: Harnessing question premises in visual question answering
Aroma Mahendru, Viraj Prabhu, Akrit Mohapatra, Dhruv Batra, and Stefan Lee. The promise of premise: Harnessing question premises in visual question answering. arXiv preprint arXiv:1705.00601, 2017
2017 arXiv
-
[207]
Wordnet: a lexical database for english
George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38 0 (11): 0 39--41, 1995
1995
-
[208]
A dataset of clinically generated visual questions and answers about radiology images
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5 0 (1): 0 1--10, 2018 b
2018
-
[209]
Medical visual question answering via conditional reasoning
Li-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen, and Xiao-Ming Wu. Medical visual question answering via conditional reasoning. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2345--2354, 2020
2020
-
[210]
Multiple meta-model quantifying for medical visual question answering
Tuong Do, Binh X Nguyen, Erman Tjiputra, Minh Tran, Quang D Tran, and Anh Nguyen. Multiple meta-model quantifying for medical visual question answering. In Medical Image Computing and Computer Assisted Intervention--MICCAI 2021: 24th International Conference, Strasbourg, Franc...
2021
-
[211]
Overcoming data limitation in medical visual question answering
Binh D Nguyen, Thanh-Toan Do, Binh X Nguyen, Tuong Do, Erman Tjiputra, and Quang D Tran. Overcoming data limitation in medical visual question answering. In Medical Image Computing and Computer Assisted Intervention--MICCAI 2019: 22nd International Conference, Shenzhen, China,...
2019
-
[212]
Contrastive pre-training and representation distillation for medical visual question answering based on radiology images
Bo Liu, Li-Ming Zhan, and Xiao-Ming Wu. Contrastive pre-training and representation distillation for medical visual question answering based on radiology images. In Medical Image Computing and Computer Assisted Intervention--MICCAI 2021: 24th International Conference, Strasbou...
2021
-
[213]
Umass at imageclef medical visual question answering (med-vqa) 2018 task
Yalei Peng, Feifan Liu, and Max P Rosen. Umass at imageclef medical visual question answering (med-vqa) 2018 task. In CLEF (Working Notes), 2018
2018
-
[214]
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608--3...
2018
-
[215]
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121--137. Springer, 2020 a
2020
-
[216]
Visual question answering from remote sensing images
Sylvain Lobry, Jesse Murray, Diego Marcos, and Devis Tuia. Visual question answering from remote sensing images. In IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium, pages 4951--4954. IEEE, 2019
2019
-
[217]
Rsvqa: Visual question answering for remote sensing data
Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia. Rsvqa: Visual question answering for remote sensing data. IEEE Transactions on Geoscience and Remote Sensing, 58 0 (12): 0 8555--8566, 2020
2020
-
[218]
How to find a good image-text embedding for remote sensing visual question answering? arXiv preprint arXiv:2109.11848, 2021
Christel Chappuis, Sylvain Lobry, Benjamin Kellenberger, Bertrand Le Saux, and Devis Tuia. How to find a good image-text embedding for remote sensing visual question answering? arXiv preprint arXiv:2109.11848, 2021
2021 arXiv
-
[219]
Cross-modal visual question answering for remote sensing data: The international conference on digital image computing: Techniques and applications (dicta 2021)
Rafael Felix, Boris Repasky, Samuel Hodge, Reza Zolfaghari, Ehsan Abbasnejad, and Jamie Sherrah. Cross-modal visual question answering for remote sensing data: The international conference on digital image computing: Techniques and applications (dicta 2021). In 2021 Digital Im...
2021
-
[220]
Bi-modal transformer-based approach for visual question answering in remote sensing imagery
Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mohamed Lamine Mekhalfi, Mansour Abdulaziz Al Zuair, and Farid Melgani. Bi-modal transformer-based approach for visual question answering in remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 60: 0 1--11, 2022
2022
-
[221]
Mutual attention inception network for remote sensing visual question answering
Xiangtao Zheng, Binqiang Wang, Xingqian Du, and Xiaoqiang Lu. Mutual attention inception network for remote sensing visual question answering. IEEE Transactions on Geoscience and Remote Sensing, 60: 0 1--14, 2021
2021
-
[222]
A spatial hierarchical reasoning network for remote sensing visual question answering
Zixiao Zhang, Licheng Jiao, Lingling Li, Xu Liu, Puhua Chen, Fang Liu, Yuxuan Li, and Zhicheng Guo. A spatial hierarchical reasoning network for remote sensing visual question answering. IEEE Transactions on Geoscience and Remote Sensing, 2023 a
2023
-
[223]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020
1901
-
[224]
Recommending themes for ad creative design via visual-linguistic representations
Yichao Zhou, Shaunak Mishra, Manisha Verma, Narayan Bhamidipati, and Wei Wang. Recommending themes for ad creative design via visual-linguistic representations. In Proceedings of The Web Conference 2020, pages 2521--2527, 2020
2020
-
[225]
A diagram is worth a dozen images
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14, pages 235--25...
2016
-
[226]
Isaaq--mastering textbook questions with pre-trained transformers and bottom-up and top-down attention
Jose Manuel Gomez-Perez and Raul Ortega. Isaaq--mastering textbook questions with pre-trained transformers and bottom-up and top-down attention. arXiv preprint arXiv:2010.00562, 2020
2010 arXiv
-
[227]
Moqa-a multi-modal question answering architecture
Monica Haurilet, Ziad Al-Halah, and Rainer Stiefelhagen. Moqa-a multi-modal question answering architecture. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0--0, 2018
2018
-
[228]
Textbook question answering under instructor guidance with memory networks
Juzheng Li, Hang Su, Jun Zhu, Siyu Wang, and Bo Zhang. Textbook question answering under instructor guidance with memory networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3655--3663, 2018 a
2018
-
[229]
Spatial-semantic collaborative graph network for textbook question answering
Yaxian Wang, Bifan Wei, Jun Liu, Qika Lin, Lingling Zhang, and Yaqiang Wu. Spatial-semantic collaborative graph network for textbook question answering. IEEE Transactions on Circuits and Systems for Video Technology, 33 0 (7): 0 3214--3228, 2023. doi:10.1109/TCSVT.2022.3231463
2023
-
[230]
Weakly supervised learning for textbook question answering
Jie Ma, Qi Chai, Jingyue Huang, Jun Liu, Yang You, and Qinghua Zheng. Weakly supervised learning for textbook question answering. IEEE Transactions on Image Processing, 31: 0 7378--7388, 2022
2022
-
[231]
Essay-anchor attentive multi-modal bilinear pooling for textbook question answering
Juzheng Li, Hang Su, Jun Zhu, and Bo Zhang. Essay-anchor attentive multi-modal bilinear pooling for textbook question answering. In 2018 IEEE International Conference on Multimedia and Expo (ICME), pages 1--6. IEEE, 2018 b
2018
-
[232]
Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, and Xiaodan Liang. Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression. arXiv preprint arXiv:2212.02746, 2022 b
2022 arXiv
-
[233]
A multi-modal neural geometric solver with textual clauses parsed from diagram
Ming-Liang Zhang, Fei Yin, and Cheng-Lin Liu. A multi-modal neural geometric solver with textual clauses parsed from diagram. arXiv preprint arXiv:2302.11097, 2023 b
2023 arXiv
-
[234]
An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Jie Cao and Jing Xiao. An augmented benchmark dataset for geometric question answering through dual parallel text encoding. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1511--1520, Gyeongju, Republic of Korea, October 2022. Internatio...
2022
-
[235]
An educational robot system of visual question answering for preschoolers
Bin He, Meng Xia, Xinguo Yu, Pengpeng Jian, Hao Meng, and Zhanwen Chen. An educational robot system of visual question answering for preschoolers. In 2017 2nd international conference on robotics and automation engineering (ICRAE), pages 441--445. IEEE, 2017
2017
-
[236]
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al. An empirical study of training end-to-end vision-and-language transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...
2022
-
[237]
Abc-cnn: An attention based convolutional neural network for visual question answering
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia. Abc-cnn: An attention based convolutional neural network for visual question answering. arXiv preprint arXiv:1511.05960, 2015
2015 arXiv
-
[238]
Pythia v0
Yu Jiang, Vivek Natarajan, Xinlei Chen, Marcus Rohrbach, Dhruv Batra, and Devi Parikh. Pythia v0. 1: the winning entry to the vqa challenge 2018. arXiv preprint arXiv:1807.09956, 2018
2018 arXiv
-
[239]
Proto: Program-guided transformer for program-guided tasks
Zelin Zhao, Karan Samel, Binghong Chen, et al. Proto: Program-guided transformer for program-guided tasks. Advances in Neural Information Processing Systems, 34: 0 17021--17036, 2021
2021
-
[240]
Learning conditioned graph structures for interpretable visual question answering
Will Norcliffe-Brown, Stathis Vafeias, and Sarah Parisot. Learning conditioned graph structures for interpretable visual question answering. Advances in neural information processing systems, 31, 2018
2018
-
[241]
Multi-modality latent interaction network for visual question answering
Peng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang, and Hongsheng Li. Multi-modality latent interaction network for visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5825--5835, 2019
2019
-
[242]
Chain of reasoning for visual question answering
Chenfei Wu, Jinlai Liu, Xiaojie Wang, and Xuan Dong. Chain of reasoning for visual question answering. Advances in Neural Information Processing Systems, 31, 2018
2018
-
[243]
Relation-aware graph attention network for visual question answering
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. Relation-aware graph attention network for visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10313--10322, 2019 b
2019
-
[244]
Learning by abstraction: The neural state machine
Drew Hudson and Christopher D Manning. Learning by abstraction: The neural state machine. Advances in Neural Information Processing Systems, 32, 2019 b
2019
-
[245]
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849, 2020
2004 arXiv
-
[246]
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104--120. Springer, 2020
2020
-
[247]
How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021
2021 arXiv
-
[248]
Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. arXiv preprint arXiv:2012.15409, 2020 b
2012 arXiv
-
[249]
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5579--5588, 2021
2021
-
[250]
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning, pages 1931--1942. PMLR, 2021
1931
-
[251]
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1780--1790, 2021
2021
-
[252]
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3208--3216, 2021
2021
-
[253]
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021
2021 arXiv
-
[254]
Unimo-2: End-to-end unified vision-language grounded learning
Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. Unimo-2: End-to-end unified vision-language grounded learning. arXiv preprint arXiv:2203.09067, 2022 a
2022 arXiv
-
[255]
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022 c
2022 arXiv
-
[256]
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022 d
2022 arXiv
-
[257]
High-performance large-scale image recognition without normalization
Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. In International Conference on Machine Learning, pages 1059--1071. PMLR, 2021
2021
-
[258]
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022 c
2022 arXiv
-
[259]
mplug: Effective and efficient vision-language learning by cross-modal skip-connections
Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, Hehong Chen, Guohai Xu, Zheng Cao, et al. mplug: Effective and efficient vision-language learning by cross-modal skip-connections. arXiv preprint arXiv:2205.12005, 2022 b
2022 arXiv
-
[260]
C hart QA : A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. C hart QA : A benchmark for question answering about charts with visual and logical reasoning. In Smaranda Muresan and Aline Nakov, Preslav andbai2023qwen Villavicencio, editors, Findings of the Associatio...
2022 doi
-
[261]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296--26306, 2024 b
2024
-
[262]
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36: 0 46595--46623, 2023
2023
-
[263]
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020
2003 arXiv
-
[264]
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019
1909 arXiv
-
[265]
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654, 2020 b
2006 arXiv
-
[266]
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao. Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering. IEEE transactions on neural networks and learning systems, 29 0 (12): 0 5947--5959, 2018
2018
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.