Pith. sign in

REVIEW 5 major objections 6 minor 5 cited by

Visual question answering: from early developments to recent advances -- a survey

T0 review · 5 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper's central claim is that a single taxonomy of VQA architectures—organized by vision encoder, language encoder, fusion machine, and answer decoder—can structure the whole field and expose its shift from simple fusion to large…

desk verdict A useful but sloppy VQA survey: structure and LVLM coverage are fine, but citation and transcription errors in the dataset and model tables break its value as a reference until fixed. read the letter →

arxiv 2501.03939 v2 pith:QRQR2TJZ submitted 2025-01-07 cs.CV cs.MM

classification cs.CVcs.MM
keywords visualquestionansweringmultimodallearningvision-languagepretraininglargelanguagemodelsattentionmechanismsbenchmarkdatasetstaxonomysurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey sets out to give the visual question answering field a structured map: a taxonomy that sorts VQA architectures by their vision encoder, language encoder, fusion mechanism, and answer decoder. The authors argue that earlier surveys covered only pieces, such as fusion techniques or datasets, whereas an updated survey is needed because large visual language models (LVLMs) have reshaped the field since 2019. If the taxonomy is right, researchers can compare any VQA model along the same four axes and see how the field moved from simple fusion, through attention and bilinear pooling, to pretrained LVLMs. The paper also assembles the main datasets, evaluation metrics, and application domains as part of the same organizing scheme.

What carries the argument

The load-bearing object is the taxonomy shown in Figure 3, which decomposes every VQA system into the same four functional blocks and then subdivides each block into the technique families used in the literature. It does the organizing work of the survey: every model reviewed in the paper is placed in this grid, and Table 3 and Figure 4 translate the taxonomy into quantitative comparisons by reporting each model's accuracy on benchmarks such as VQA v1.0, VQA v2.0, DAQUAR, Visual7W, COCO-QA, CLEVR, and GQA. The taxonomy is what allows the paper to draw timeline conclusions, such as object-based vision encoders giving way to ViT patch encoders and transformer-based fusion displacing earlier attention designs.

What would settle it

Select a random sample of rows from Table 3, locate the cited original papers, and verify the reported accuracy values and dataset names; if a substantial share of entries differ from the originals, the paper's claim to be a reliable VQA resource fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the whole of VQA architecture design can be organized by four components: vision encoder (grid-based, object-based, or ViT patch-based), language encoder (bag-of-words, RNN, CNN, or transformer), fusion machine (simple fusion, attention, bilinear pooling, relation networks, neural module networks, or large visual-language models), and answer decoder (closed-vocabulary or open-vocabulary). Within this taxonomy, the survey charts a historical progression: early models used CNNs and LSTMs with simple fusion, attention-based and bilinear pooling methods then raised accuracy, and after 2019 LVLMs pretrained on image-text data became the dominant approach, with accuracy on VQA v2.0 climbing from about 52 percent in 2015 to above 84 percent by 2022. The authors' stated intent is that this structured framework makes comparative analysis and evaluation straightforward and gives researchers and practitioners a single resource covering methods, datasets, metrics, and applications.

Load-bearing premise

The survey's usefulness as a reference depends on the accuracy numbers in Table 3 and the dataset attributions in Table 2 being transcribed correctly from the cited papers; if those transcriptions are unreliable, the comparison tables cannot be trusted.

Editorial extensions

If this is right

  • Researchers can locate any VQA model's design choices on the taxonomy's four axes and compare models that differ in only one component, isolating which part drives accuracy.
  • The four-period timeline implies that future advances will come from LVLMs, with ViT-based patch encoders and pretrained large language model backbones dominating the field.
  • The accuracy tables imply that transformer-based fusion, especially self-attention and cross-attention, was the strongest fusion family before the LVLM era, with MCAN and its extensions setting the pre-LVLM ceiling on VQA v2.0.
  • Because the survey maps datasets by domain, it implies that progress in medical, remote sensing, education, cultural heritage, and advertising VQA tracks dataset creation as much as model design.
  • For closed-vocabulary VQA tasks, the decoder taxonomy suggests that open-ended and multiple-choice setups should be treated as distinct problems with distinct evaluation metrics.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy is taken as the field's organizing scheme, an implicit testable prediction is that future state-of-the-art VQA models will be describable as LVLMs with ViT encoders and LLM backbones, making the older fusion categories mostly historical.
  • The survey's own section on question relevance reports that a large share of human questions are unrelated to their images; a production VQA system could therefore benefit from an explicit relevance-detection gate before answering, a design the paper does not develop.
  • A consequence the authors leave implicit is that Table 3's benchmark rankings conflate architecture choice with model scale; separating those two effects would require controlled comparisons the table does not provide.
  • One testable extension would be to use the taxonomy's four axes as features for predicting which new VQA datasets a model will transfer to, treating the survey's classification as input to a prediction study.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This survey proposes a taxonomy of VQA architectures organized by vision encoder, language encoder, fusion method, and answer decoder, and uses it to structure a review of deep-learning VQA methods, LVLMs, datasets, metrics, applications, and future directions. The paper includes a large table of datasets (Table 2) and a large table of model accuracies (Table 3), and it describes applications in medicine, accessibility, remote sensing, cultural heritage, advertising, and education. The stated contribution is to serve as a comprehensive, up-to-date reference resource for VQA researchers and practitioners as of early 2025.

Significance. If the factual content were reliable, this would be a genuinely useful synthesis: the taxonomy in Figure 3 is reasonable, the coverage of historical periods from 2015 to 2024 is broad, and the treatment of LVLMs as the current dominant paradigm is appropriate. The paper also covers multiple applied domains and gives a clear account of the standard VQA pipeline. However, the survey's main value as a reference depends on the accuracy of its tables and citations, and the manuscript currently contains several concrete traceability failures. Those failures are not cosmetic; they directly undermine the paper's central claim to be a comprehensive resource for lookup and comparison. I therefore evaluate the manuscript as promising but not yet publishable in its present form.

major comments (5)
  1. [Table 2, rows 37–38] RSVQA-low and RSVQA-high are attributed to Gurari et al. [2018a], but Section 7.3 correctly attributes RSVQA to Lobry et al. [2019] and Lobry et al. [2020]. A reader using Table 2 as a lookup cannot trace the dataset to its primary source, so the table fails its stated reference function.
  2. [Table 3, Qwen-VL 7B row] The Qwen-VL 7B row cites Masry et al. [2022b] for both the image encoder and the language encoder, but Qwen-VL is introduced in Bai et al. [2023]; Masry et al. is the ChartQA paper. The 78.80 accuracy should be verified against the Qwen-VL report, and the row should cite the correct source.
  3. [Section 7.4] The sentence 'To overcome the need for expert-generated descriptions, Bongini Brown et al. [2020] used GPT-3 to automatically generate detailed descriptions of artworks' cites a work that does not appear in the bibliography. The closest entry, Brown et al. [2020] 'Language models are few-shot learners', describes no such artwork study, so this claim is currently unsupported and needs a correct reference or removal.
  4. [Table 2, row 10] AI2D is cited to Sheng et al. [2016], the same source used for Art-VQA in row 9, but the actual AI2D source is Kembhavi et al. [2016], which is cited in the text. This is another instance in which the dataset table provides incorrect provenance for an entry.
  5. [Section 8.3, Period III] The text states that VQA v2.0 accuracy during Period III ranged from 71% to 78%, but Table 3 lists Florence (80.16), SimVLM (80.34), and VLMo (82.78) with 2021 dates inside that period. The text and table contradict each other; the range should be corrected or the period boundaries clarified.
minor comments (6)
  1. [Section 3.3.3] The example question is 'What is the object on the refrigerator?', but the described reasoning step says 'identifying the object on the table'; this internal inconsistency should be fixed.
  2. [Section 3.3.1] The sentence 'simple fusion techniques need to provide insight into how the visual and textual inputs are combined' appears to mean the opposite; it should say 'fail to provide insight' or 'provide no insight'.
  3. [Figure 4] The axis labels are garbled, for example 'VI515/0(20 - 08)017/2 7/08(201 - 8)19/020 /089201( - )1/08202 1/08(202 - )222/120'; the figure should be regenerated with legible period labels.
  4. [Section 4.2] The text says questions 'must begin with one of six letters: what, where, how, why, who, and when', but these are interrogative words rather than letters; the wording should be corrected to 'six question categories' or similar.
  5. [References] The bibliography contains duplicate entries (e.g., Antol et al. 2015a/2015b, Devlin et al. 2018/2019, Liu et al. 2019a/2019b) and a malformed author string in the Masry et al. [2022b] entry ('andbai2023qwen Villavicencio'); these need cleanup.
  6. [Section 5] The final sentence about the SimpsonsVQA dataset appears without citation or connection to the surrounding discussion of question relevance; it should be integrated or removed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's contributions are descriptive attributions to external prior work, with no derivation or fitted prediction that reduces to its own inputs.

full rationale

This is a literature survey, not a derivation-driven paper. Its central claim is that it introduces a taxonomy of VQA architectures and organizes existing datasets, models, and metrics. A taxonomy is a categorization scheme, not a result derived from equations or fitted parameters; the survey does not claim to predict any quantity from its own assumptions. All technical claims are attributed to external references, and I found no self-citation chain that is load-bearing: the authors do not appear to rely on their own prior work to justify the taxonomy or any other central premise. The paper contains factual traceability problems that a reader should weigh under correctness risk, not circularity: Table 2 attributes RSVQA to Gurari et al. rather than Lobry et al., Section 7.4 cites a non-existent 'Bongini Brown et al.' for a GPT-3 artwork-description study, and Table 3's Qwen-VL row cites the ChartQA paper as the model source instead of the Qwen technical report. These are citation/attribution errors that reduce the survey's usefulness as a reference, but they do not make any claim definitionally equivalent to its input. Because the survey is self-contained as a review and makes no circular derivation, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey's usefulness rests on two unproven premises: that the cited papers are accurately represented, and that the selected literature is representative and complete. Both are strained by the citation errors documented in the red flags.

assumptions (2)
  • domain assumption The cited papers are accurately represented in the survey text and tables
    The entire value of the survey rests on this; it is contradicted by the misattributions in Table 2 (RSVQA to Gurari et al.) and Table 3 (Qwen-VL to Masry et al.).
  • domain assumption The selected papers form a representative and complete sample of the VQA literature as of early 2025
    The paper claims to be comprehensive but provides no search strategy or inclusion criteria; this assumption is not verified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Visual question answering: from early developments to recent advances -- a survey." pith.science (2026). https://pith.science/paper/QRQR2TJZ

@misc{pith2026250103939,
  author       = {Pith},
  title        = {Pith review of: Visual question answering: from early developments to recent advances -- a survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QRQR2TJZ}},
  note         = {Machine review of arXiv:2501.03939}
}
read the original abstract

Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text embedding, natural language understanding, and language generation. With the growth of multimodal data research, VQA has gained significant attention due to its broad applications, including interactive educational tools, medical image diagnosis, customer service, entertainment, and social media captioning. Additionally, VQA plays a vital role in assisting visually impaired individuals by generating descriptive content from images. This survey introduces a taxonomy of VQA architectures, categorizing them based on design choices and key components to facilitate comparative analysis and evaluation. We review major VQA approaches, focusing on deep learning-based methods, and explore the emerging field of Large Visual Language Models (LVLMs) that have demonstrated success in multimodal tasks like VQA. The paper further examines available datasets and evaluation metrics essential for measuring VQA system performance, followed by an exploration of real-world VQA applications. Finally, we highlight ongoing challenges and future directions in VQA research, presenting open questions and potential areas for further development. This survey serves as a comprehensive resource for researchers and practitioners interested in the latest advancements and future

Figures

Figures reproduced from arXiv: 2501.03939 by the authors.

Figure 1
Figure 1. Overview of a VQA system [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. General VQA system. extracts textual representations from the input question using Natural Language Processing (NLP) techniques like RNNs, Transformers, etc. Further details on both encoders can be found in Section 3.1 and Section 3.2, respectively. The output vectors from the Image and Language Encoders are then passed through a fusion component, which combines them using a technique such as element-wise product or… view at source ↗
Figure 3
Figure 3. VQA Taxonomy. 5 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The accuracy of the model on different datasets over the years (from May 2015 to December 2022). [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free SpecFlow condenses VLM visual tokens via kNN heat diffusion, adaptive quadtree budgets, and coreset sinks, retaining 95.6% LLaVA-1.5 performance after pruning 88.9% of tokens.

  2. Affordance Benchmark for MLLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 2,000-question benchmark finds multimodal AI models recognize object affordances far worse than humans, with top model Gemini-2.0-Pro at 18.05% versus 85.34% human best.

  3. Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniVQA is a first open-source dataset and benchmark for 360-degree visual question answering, and 360-R1 uses GRPO with three LLM-based rewards to improve an existing multimodal model on it.

  4. Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring

    cs.CV 2025-07 reject novelty 5.0 of 10

    Four open-source VQA models reach moderate accuracy on a new classroom video dataset, with yes/no questions easiest and counting/reasoning hardest.

  5. From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A staged knowledge-prompting method (KAAR) improves LLM test accuracy on ARC by about 5 absolute points over repeated-sampling plan-guided code generation, reaching 35% with GPT-o3-mini.

Reference graph

Works this paper leans on

264 extracted references · 11 canonical work pages · cited by 5 Pith papers

  1. [1]

    Multimodal machine learning: A survey and taxonomy

    Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41 0 (2): 0 423--443, 2019. doi:10.1109/TPAMI.2018.2798607

  2. [2]

    Hershey, Tim K

    Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi. Attention-based multimodal fusion for video description. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 4203--4212, 2017. doi:10.1109/ICCV.2017.450

  3. [3]

    Multimodal language analysis in the wild: CMU - MOSEI dataset and interpretable dynamic fusion graph

    AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. Multimodal language analysis in the wild: CMU - MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236--2246, Melbourne, Australia, July...

  4. [4]

    Merlot reserve: Neural script knowledge through vision and language and sound

    Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. Merlot reserve: Neural script knowledge through vision and language and sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16375--16387, June 2022

  5. [5]

    Multimodal learning with graphs

    Yasha Ektefaie, George Dasoulas, Ayush Noori, Maha Farhat, and Marinka Zitnik. Multimodal learning with graphs. Nature Machine Intelligence, pages 1--11, 2023

  6. [6]

    Simvlm: Simple visual language model pretraining with weak supervision

    Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904, 2021

  7. [7]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. arXiv preprint arXiv:2204.14198, 2022

  8. [8]

    Visual language integration: A survey and open challenges

    Sang-Min Park and Young-Gab Kim. Visual language integration: A survey and open challenges. Computer Science Review, 48: 0 100548, 2023. ISSN 1574-0137. doi:https://doi.org/10.1016/j.cosrev.2023.100548. URL https://www.sciencedirect.com/science/article/pii/S1574013723000151

Show all 264 references
  1. [9]

    Lawrence Zitnick, and Devi Parikh

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), December 2015 a

  2. [10]

    Exploring models and data for image question answering

    Mengye Ren, Ryan Kiros, and Richard Zemel. Exploring models and data for image question answering. Advances in neural information processing systems, 28, 2015 a

  3. [11]

    Yin and yang: Balancing and answering binary visual questions

    Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Yin and yang: Balancing and answering binary visual questions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016

  4. [12]

    Learning to answer questions from image using convolutional neural network

    Lin Ma, Zhengdong Lu, and Hang Li. Learning to answer questions from image using convolutional neural network. In Thirtieth AAAI Conference on Artificial Intelligence, 2016

  5. [13]

    Making the v in vqa matter: Elevating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017 a

  6. [14]

    Lawrence Zitnick, and Devi Parikh

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. VQA : V isual Q uestion A nswering. In International Conference on Computer Vision (ICCV), 2015 b

  7. [15]

    Shamma, Michael S

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li - Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei - Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations...

  8. [16]

    From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

    Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2: 0 67--78, 2014. doi:10.1162/tacl...

  9. [17]

    Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

    Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR...

  10. [18]

    Medical visual question answering: A survey

    Zhihong Lin, Donghao Zhang, Qingyi Tac, Danli Shi, Gholamreza Haffari, Qi Wu, Mingguang He, and Zongyuan Ge. Medical visual question answering: A survey. arXiv preprint arXiv:2111.10056, 2021

  11. [19]

    Cross-modal self-attention with multi-task pre-training for medical visual question answering

    Haifan Gong, Guanqi Chen, Sishuo Liu, Yizhou Yu, and Guanbin Li. Cross-modal self-attention with multi-task pre-training for medical visual question answering. In Proceedings of the 2021 International Conference on Multimedia Retrieval, ICMR '21, page 456–460, New York, NY, US...

  12. [20]

    A sequence-to-sequence model approach for imageclef 2018 medical domain visual question answering

    Rahul Ambati and Chakravardhan Reddy Dudyala. A sequence-to-sequence model approach for imageclef 2018 medical domain visual question answering. In 2018 15th IEEE India Council International Conference (INDICON), pages 1--6, 2018. doi:10.1109/INDICON45594.2018.8987108

  13. [21]

    An encoder-decoder model for visual question answering in the medical domain

    Imane Allaouzi, Mohamed Ben Ahmed, and Badr Benamrou. An encoder-decoder model for visual question answering in the medical domain. In Conference and Labs of the Evaluation Forum, 2019

  14. [22]

    Sysu-hcp at vqa-med 2021: A data-centric model with efficient training methodology for medical visual question answering

    Haifan Gong, Ricong Huang, Guanqi Chen, and Guanbin Li. Sysu-hcp at vqa-med 2021: A data-centric model with efficient training methodology for medical visual question answering. In CLEF, 2021 b

  15. [23]

    Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller

    Sadid A. Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Matthew Lungren, and Henning M\"uller. Overview of the ImageCLEF 2018 medical domain visual question answering task. In CLEF2018 Working Notes, CEUR Workshop Proceedings, Avignon, France, September 10-14 2018 a . CEUR-WS.or...

  16. [24]

    Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman

    Jason J. Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific Data, 5, 2018 a

  17. [25]

    Hasan, and Henning M\"uller

    Asma Ben Abacha , Mourad Sarrouti, Dina Demner-Fushman, Sadid A. Hasan, and Henning M\"uller. Overview of the vqa-med task at imageclef 2021: Visual question answering and generation in the medical domain. In CLEF 2021 Working Notes, CEUR Workshop Proceedings, Bucharest, Roman...

  18. [26]

    Iqa: Visual question answering in interactive environments

    Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. Iqa: Visual question answering in interactive environments. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018

  19. [27]

    Show and tell: A neural image caption generator

    Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015

  20. [28]

    Yuille, and Kevin Murphy

    Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L. Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016

  21. [29]

    Transvg: End-to-end visual grounding with transformers

    Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li. Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1769--1779, October 2021

  22. [30]

    Hierarchical lstms with adaptive attention for visual captioning

    Lianli Gao, Xiangpeng Li, Jingkuan Song, and Heng Tao Shen. Hierarchical lstms with adaptive attention for visual captioning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42 0 (5): 0 1112--1131, 2020. doi:10.1109/TPAMI.2019.2894139

  23. [31]

    Semantic compositional networks for visual captioning

    Zhe Gan, Chuang Gan, Xiaodong He, Yunchen Pu, Kenneth Tran, Jianfeng Gao, Lawrence Carin, and Li Deng. Semantic compositional networks for visual captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017

  24. [32]

    Information fusion in visual question answering: A survey

    Dongxiang Zhang, Rui Cao, and Sai Wu. Information fusion in visual question answering: A survey. Information Fusion, 52: 0 268--280, 2019

  25. [33]

    Visual question answering: methodologies and challenges

    Liyana Sahir Kallooriyakath, MV Jithin, PV Bindu, and PP Adith. Visual question answering: methodologies and challenges. In 2020 International Conference on Smart Technologies in Computing, Electrical and Electronics (ICSTCEE), pages 402--407. IEEE, 2020

  26. [34]

    Visual question answering: a state-of-the-art review

    Sruthy Manmadhan and Binsu C Kovoor. Visual question answering: a state-of-the-art review. Artificial Intelligence Review, 53: 0 5705--5745, 2020

  27. [35]

    A survey on vqa: Datasets and approaches

    Yeyun Zou and Qiyu Xie. A survey on vqa: Datasets and approaches. In 2020 2nd International Conference on Information Technology and Computer Application (ITCA), pages 289--297. IEEE, 2020

  28. [36]

    A survey of methods, datasets and evaluation metrics for visual question answering

    Himanshu Sharma and Anand Singh Jalal. A survey of methods, datasets and evaluation metrics for visual question answering. Image and Vision Computing, 116: 0 104327, 2021

  29. [37]

    Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32, 2019

  30. [38]

    Visualbert: A simple and performant baseline for vision and language

    Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019 a

  31. [39]

    Vl-bert: Pre-training of generic visual-linguistic representations

    Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530, 2019

  32. [40]

    Towards vqa models that can read

    Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317--8326, 2019

  33. [41]

    Tag: Boosting text-vqa via text-aware visual question-answer generation

    Jun Wang, Mingfei Gao, Yuqian Hu, Ramprasaath R Selvaraju, Chetan Ramaiah, Ran Xu, Joseph F JaJa, and Larry S Davis. Tag: Boosting text-vqa via text-aware visual question-answer generation. arXiv preprint arXiv:2208.01813, 2022 a

  34. [42]

    Ocr-vqa: Visual question answering by reading text in images

    Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 947--952, 2019. doi:10.1109/ICDAR.2019.00156

  35. [43]

    Visualmrc: Machine reading comprehension on document images

    Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. Visualmrc: Machine reading comprehension on document images. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13878--13888, 2021

  36. [44]

    Scene text visual question answering

    Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Mar c al Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas. Scene text visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4291--4301, 2019

  37. [45]

    Docvqa: A dataset for vqa on document images

    Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200--2209, 2021

  38. [46]

    Explicit knowledge-based reasoning for visual question answering

    Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick. Explicit knowledge-based reasoning for visual question answering. arXiv preprint arXiv:1511.02570, 2015

  39. [47]

    Ask me anything: Free-form visual question answering based on knowledge from external sources

    Qi Wu, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. Ask me anything: Free-form visual question answering based on knowledge from external sources. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4622--4630, 2016

  40. [48]

    Ok-vqa: A visual question answering benchmark requiring external knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pages 3195--3204, 2019

  41. [49]

    A-okvqa: A benchmark for visual question answering using world knowledge

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. A-okvqa: A benchmark for visual question answering using world knowledge. In Computer Vision--ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23--27, 2022, Proceedings, P...

  42. [50]

    Grounding answers for visual questions asked by visually impaired people

    Chongyan Chen, Samreen Anjum, and Danna Gurari. Grounding answers for visual questions asked by visually impaired people. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19098--19107, 2022 a

  43. [51]

    WSDM Cup 2023 Challenge on Visual Question Answering

    Dmitry Ustalov, Nikita Pavlichenko, Daniil Likhobaba, and Alisa Smirnova. WSDM Cup 2023 Challenge on Visual Question Answering . In Proceedings of the 4th Crowd Science Workshop on Collaboration of Humans and Learning Algorithms for Data Labeling, pages 1--7, Singapore, 2023. ...

  44. [52]

    Flava: A foundational language and vision alignment model

    Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. Flava: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15638--1...

  45. [53]

    Masked vision and language modeling for multi-modal representation learning

    Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, and Stefano Soatto. Masked vision and language modeling for multi-modal representation learning. arXiv preprint arXiv:2208.02131, 2022

  46. [54]

    Self-supervised learning from images with a joint-embedding predictive architecture

    Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...

  47. [55]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll \'a r, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000--16009, 2022

  48. [56]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...

  49. [57]

    Scaling up visual and vision-language representation learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904--4916...

  50. [58]

    Sigmoid loss for language image pre-training, 2023

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training, 2023

  51. [59]

    Chameleon: Mixed-modal early-fusion foundation models, 2024

    Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models, 2024. URL https://arxiv.org/abs/2405.09818

  52. [60]

    Gpt-4 technical report, 2023

    OpenAI. Gpt-4 technical report, 2023

  53. [61]

    Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

    Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning...

  54. [62]

    Visual instruction tuning, 2023 a

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023 a

  55. [63]

    Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models, 2021. URL https://arxiv.org/abs/2106.13884

  56. [64]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023 a

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023 a . URL https://arxiv.org/abs/2301.12597

  57. [65]

    Qwen technical report, 2023

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Si...

  58. [66]

    A simple neural network module for relational reasoning

    Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. A simple neural network module for relational reasoning. Advances in neural information processing systems, 30, 2017

  59. [67]

    Graph-structured representations for visual question answering

    Damien Teney, Lingqiao Liu, and Anton van Den Hengel. Graph-structured representations for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1--9, 2017

  60. [68]

    Multimodal compact bilinear pooling for visual question answering and visual grounding

    Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv preprint arXiv:1606.01847, 2016

  61. [69]

    Mutan: Multimodal tucker fusion for visual question answering

    Hedi Ben-Younes, R \'e mi Cadene, Matthieu Cord, and Nicolas Thome. Mutan: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2612--2620, 2017

  62. [70]

    Neural module networks

    Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. Neural module networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 39--48, 2016

  63. [71]

    Learning to reason: End-to-end module networks for visual question answering

    Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. Learning to reason: End-to-end module networks for visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 804--813, 2017

  64. [72]

    Deep modular co-attention networks for visual question answering

    Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6281--6290, 2019

  65. [73]

    An improved attention for visual question answering

    Tanzila Rahman, Shih-Han Chou, Leonid Sigal, and Giuseppe Carenini. An improved attention for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1653--1662, 2021

  66. [74]

    Hierarchical question-image co-attention for visual question answering

    Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. Hierarchical question-image co-attention for visual question answering. Advances in neural information processing systems, 29, 2016

  67. [75]

    Dual attention networks for multimodal reasoning and matching

    Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. Dual attention networks for multimodal reasoning and matching. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 299--307, 2017

  68. [76]

    Stacked attention networks for image question answering

    Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. Stacked attention networks for image question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016

  69. [77]

    Bottom-up and top-down attention for image captioning and visual question answering

    Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 60...

  70. [78]

    Ask your neurons: A neural-based approach to answering questions about images

    Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz. Ask your neurons: A neural-based approach to answering questions about images. In Proceedings of the IEEE international conference on computer vision, pages 1--9, 2015

  71. [79]

    BERT : Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Lan...

  72. [81]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023 a

  73. [82]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90\ quality, March 2023. URL https://lmsys.org/blog/2023...

  74. [83]

    Long short-term memory

    J \"u rgen Schmidhuber, Sepp Hochreiter, et al. Long short-term memory. Neural Comput, 9 0 (8): 0 1735--1780, 1997

  75. [84]

    Schuster and K.K

    M. Schuster and K.K. Paliwal. Bidirectional recurrent neural networks. IEEE Transactions on Signal Processing, 45 0 (11): 0 2673--2681, 1997. doi:10.1109/78.650093

  76. [85]

    Learning phrase representations using RNN encoder-decoder for statistical machine translation

    Kyunghyun Cho, Bart van Merrienboer, C aglar G \" u l c ehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations using RNN encoder-decoder for statistical machine translation. CoRR, abs/1406.1078, 2014. URL http://arxiv.org/abs/1406.1078

  77. [86]

    Efficient estimation of word representations in vector space

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781, 2013

  78. [87]

    G lo V e: Global vectors for word representation

    Jeffrey Pennington, Richard Socher, and Christopher Manning. G lo V e: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 1532--1543, Doha, Qatar, October 2014. Association for Com...

  79. [88]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv...

  80. [89]

    You only look once: Unified, real-time object detection

    Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016

  81. [90]

    Faster r-cnn: Towards real-time object detection with region proposal networks

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran ...

  82. [91]

    Spatial pyramid pooling in deep convolutional networks for visual recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37 0 (9): 0 1904--1916, 2015. doi:10.1109/TPAMI.2015.2389824

  83. [92]

    LeCun, B

    Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. Backpropagation applied to handwritten zip code recognition. Neural Computation, 1 0 (4): 0 541--551, 1989. doi:10.1162/neco.1989.1.4.541

  84. [93]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248--255, 2009. doi:10.1109/CVPR.2009.5206848

  85. [94]

    Very deep convolutional networks for large-scale image recognition

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedin...

  86. [95]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770--778, 2016. doi:10.1109/CVPR.2016.90

  87. [96]

    Going deeper with convolutions

    Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1--9, 2015. doi:1...

  88. [97]

    In defense of grid features for visual question answering

    Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen. In defense of grid features for visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020

  89. [98]

    W ea QA : Weak supervision via captions for visual question answering

    Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, and Chitta Baral. W ea QA : Weak supervision via captions for visual question answering. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3420--3435, Online, August 2021. Association for Computat...

  90. [99]

    Question-guided feature pyramid network for medical visual question answering

    Yonglin Yu, Haifeng Li, Hanrong Shi, Lin Li, and Jun Xiao. Question-guided feature pyramid network for medical visual question answering. Expert Systems with Applications, 214: 0 119148, 2023 a . ISSN 0957-4174. doi:https://doi.org/10.1016/j.eswa.2022.119148. URL https://www.s...

  91. [100]

    Spatial transformer networks

    Max Jaderberg, Karen Simonyan, Andrew Zisserman, and koray kavukcuoglu. Spatial transformer networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 28. Curran Associates, Inc., 2015. URL https...

  92. [101]

    Principe

    Ryan Burt, Mihael Cudic, and Jose C. Principe. Fusing attention with visual question answering. In 2017 International Joint Conference on Neural Networks (IJCNN), pages 949--953, 2017. doi:10.1109/IJCNN.2017.7965954

  93. [102]

    From images to textual prompts: Zero-shot VQA with frozen large language models, 2023

    Jiaxian Guo, Junnan Li, Dongxu Li, Anthony Tiong, Boyang Li, Dacheng Tao, and Steven HOI. From images to textual prompts: Zero-shot VQA with frozen large language models, 2023. URL https://openreview.net/forum?id=Ck1UtnVukP8

  94. [103]

    Focal loss for dense object detection

    Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct 2017

  95. [104]

    Vilt: Vision-and-language transformer without convolution or region supervision

    Wonjae Kim, Bokyung Son, and Ildoo Kim. Vilt: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning, pages 5583--5594. PMLR, 2021

  96. [105]

    Vlmo: Unified vision-language pre-training with mixture-of-modality-experts

    Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, and Furu Wei. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. arXiv preprint arXiv:2111.02358, 2021

  97. [106]

    Multi-grained vision language pre-training: Aligning texts with visual concepts

    Yan Zeng, Xinsong Zhang, and Hang Li. Multi-grained vision language pre-training: Aligning texts with visual concepts. Proceedings of The 33rd International Conference on Machine Learning, 2022 a

  98. [107]

    X2-vlm: All-in-one pre-trained model for vision-language tasks

    Yan Zeng, Xinsong Zhang, Hang Li, Jiawei Wang, Jipeng Zhang, and Wangchunshu Zhou. X2-vlm: All-in-one pre-trained model for vision-language tasks. arXiv preprint arXiv:2211.12402, 2022 b

  99. [108]

    Align before fuse: Vision and language representation learning with momentum distillation

    Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, 34: 0 9694--9705, 2021

  100. [109]

    Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision ...

  101. [110]

    Midas v3

    Reiner Birkl, D Wofk, and M M \"u ller. Midas v3. 1--a model zoo for robust monocular relative depth estimation. arxiv. arXiv preprint arXiv:2307.14460, 2023

  102. [111]

    Moai: Mixture of all intelligence for large language and vision models

    Byung-Kwan Lee, Beomchan Park, Chae Won Kim, and Yong Man Ro. Moai: Mixture of all intelligence for large language and vision models. arXiv preprint arXiv:2403.07508, 2024

  103. [112]

    Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024

  104. [113]

    Skip-thought vectors

    Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Skip-thought vectors. Advances in neural information processing systems, 28, 2015

  105. [114]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  106. [115]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018

  107. [116]

    Roberta: A robustly optimized bert pretraining approach

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019 b

  108. [117]

    mt5: A massively multilingual pre-trained text-to-text transformer

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020

  109. [118]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023 b

  110. [119]

    Revisiting visual question answering baselines

    Allan Jabri, Armand Joulin, and Laurens van der Maaten. Revisiting visual question answering baselines. In European conference on computer vision, pages 727--739. Springer, 2016

  111. [120]

    Dualnet: Domain-invariant network for visual question answering

    Kuniaki Saito, Andrew Shin, Yoshitaka Ushiku, and Tatsuya Harada. Dualnet: Domain-invariant network for visual question answering. In 2017 IEEE International Conference on Multimedia and Expo (ICME), pages 829--834. IEEE, 2017

  112. [121]

    Are you talking to a machine? dataset and methods for multilingual image question

    Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. Are you talking to a machine? dataset and methods for multilingual image question. Advances in neural information processing systems, 28, 2015

  113. [122]

    Answer-type prediction for visual question answering

    Kushal Kafle and Christopher Kanan. Answer-type prediction for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4976--4984, 2016

  114. [123]

    Neural machine translation by jointly learning to align and translate

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014

  115. [124]

    Multimodal residual learning for visual qa

    Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Multimodal residual learning for visual qa. Advances in neural information processing systems, 29, 2016 a

  116. [125]

    End-to-end memory networks

    Sainbayar Sukhbaatar, Jason Weston, Rob Fergus, et al. End-to-end memory networks. Advances in neural information processing systems, 28, 2015

  117. [126]

    Ask, attend and answer: Exploring question-guided spatial attention for visual question answering

    Huijuan Xu and Kate Saenko. Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. In European conference on computer vision, pages 451--466. Springer, 2016

  118. [127]

    Dynamic memory networks for visual and textual question answering

    Caiming Xiong, Stephen Merity, and Richard Socher. Dynamic memory networks for visual and textual question answering. In Maria Florina Balcan and Kilian Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Ma...

  119. [128]

    Ask me anything: Dynamic memory networks for natural language processing

    Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. Ask me anything: Dynamic memory networks for natural language processing. In International conference on machine learning, pages 1378--1387....

  120. [129]

    Where to look: Focus regions for visual question answering

    Kevin J Shih, Saurabh Singh, and Derek Hoiem. Where to look: Focus regions for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4613--4621, 2016

  121. [130]

    A focused dynamic attention model for visual question answering

    Ilija Ilievski, Shuicheng Yan, and Jiashi Feng. A focused dynamic attention model for visual question answering. arXiv preprint arXiv:1604.01485, 2016

  122. [131]

    Tips and tricks for visual question answering: Learnings from the 2017 challenge

    Damien Teney, Peter Anderson, Xiaodong He, and Anton Van Den Hengel. Tips and tricks for visual question answering: Learnings from the 2017 challenge. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4223--4232, 2018

  123. [132]

    From pixels to objects: Cubic visual attention for visual question answering

    Jingkuan Song, Pengpeng Zeng, Lianli Gao, and Heng Tao Shen. From pixels to objects: Cubic visual attention for visual question answering. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI-18 , pages 906--912. International J...

  124. [133]

    Bilinear attention networks

    Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. Bilinear attention networks. Advances in neural information processing systems, 31, 2018

  125. [134]

    Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering

    Duy-Kien Nguyen and Takayuki Okatani. Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6087--6096, 2018

  126. [135]

    Attention on attention for image captioning

    Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. Attention on attention for image captioning. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4634--4643, 2019

  127. [136]

    Trar: Routing the attention spans in transformer for visual question answering

    Yiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun, Jianzhuang Liu, Xinghao Ding, Mingliang Xu, and Rongrong Ji. Trar: Routing the attention spans in transformer for visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20...

  128. [137]

    Txt: Crossmodal end-to-end learning with transformers

    Jan-Martin O Steitz, Jonas Pfeiffer, Iryna Gurevych, and Stefan Roth. Txt: Crossmodal end-to-end learning with transformers. In Pattern Recognition: 43rd DAGM German Conference, DAGM GCPR 2021, Bonn, Germany, September 28--October 1, 2021, Proceedings, pages 405--420. Springer, 2022

  129. [138]

    Inferring and executing programs for visual reasoning

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Inferring and executing programs for visual reasoning. In Proceedings of the IEEE international conference on computer vision, pages 2989--2998, 2017 a

  130. [139]

    Explainable neural computation via stack neural module networks

    Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko. Explainable neural computation via stack neural module networks. In Proceedings of the European conference on computer vision (ECCV), pages 53--69, 2018

  131. [140]

    Neural-symbolic vqa: Disentangling reasoning from vision and language understanding

    Kexin Yi, Jiajun Wu, Chuang Gan, Antonio Torralba, Pushmeet Kohli, and Josh Tenenbaum. Neural-symbolic vqa: Disentangling reasoning from vision and language understanding. Advances in neural information processing systems, 31, 2018

  132. [141]

    Probabilistic neural symbolic models for interpretable visual question answering

    Ramakrishna Vedantam, Karan Desai, Stefan Lee, Marcus Rohrbach, Dhruv Batra, and Devi Parikh. Probabilistic neural symbolic models for interpretable visual question answering. In International Conference on Machine Learning, pages 6428--6437. PMLR, 2019

  133. [142]

    Finding frequent items in data streams

    Moses Charikar, Kevin Chen, and Martin Farach-Colton. Finding frequent items in data streams. In International Colloquium on Automata, Languages, and Programming, pages 693--703. Springer, 2002

  134. [143]

    Hadamard product for low-rank bilinear pooling

    Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. Hadamard product for low-rank bilinear pooling. arXiv preprint arXiv:1610.04325, 2016 b

  135. [144]

    Multi-modal factorized bilinear pooling with co-attention learning for visual question answering

    Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 1821--1830, 2017

  136. [145]

    Film: Visual reasoning with a general conditioning layer

    Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018

  137. [146]

    Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection

    Hedi Ben-Younes, Remi Cadene, Nicolas Thome, and Matthieu Cord. Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8102--8109, 2019

  138. [147]

    A tutorial on energy-based learning

    Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and Fujie Huang. A tutorial on energy-based learning. Predicting structured data, 1 0 (0), 2006

  139. [148]

    Slip: Self-supervision meets language-image pre-training

    Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie. Slip: Self-supervision meets language-image pre-training. In European conference on computer vision, pages 529--544. Springer, 2022

  140. [149]

    Modeling caption diversity in contrastive vision-language pretraining, 2024

    Samuel Lavoie, Polina Kirichenko, Mark Ibrahim, Mahmoud Assran, Andrew Gordon Wilson, Aaron Courville, and Nicolas Ballas. Modeling caption diversity in contrastive vision-language pretraining, 2024

  141. [150]

    Lxmert: Learning cross-modality encoder representations from transformers

    Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490, 2019

  142. [151]

    Coca: Contrastive captioners are image-text foundation models

    Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022

  143. [152]

    Fast high-resolution image synthesis with latent adversarial diffusion distillation

    Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast high-resolution image synthesis with latent adversarial diffusion distillation. arXiv preprint arXiv:2403.12015, 2024

  144. [153]

    Stable diffusion 3: Research paper

    Stability.ai. Stable diffusion 3: Research paper. https://stability.ai/news/stable-diffusion-3-research-paper, 2024

  145. [154]

    meta-llama-3, 2023

    Meta. meta-llama-3, 2023. URL https://ai.meta.com/blog/meta-llama-3

  146. [155]

    Falcon 2

    TII. Falcon 2. In European conference on computer vision, 2024. URL https://falconllm.tii.ae/

  147. [156]

    Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a . URL https://llava-vl.github.io/blog/2024-01-30-llava-next/

  148. [157]

    Minigpt-4: Enhancing vision-language understanding with advanced large language models

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023

  149. [158]

    Visual genome: Connecting language and vision using crowdsourced dense image annotations

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer ...

  150. [159]

    A dataset for multimodal question answering in the cultural heritage domain

    Shurong Sheng, Luc Van Gool, and Marie-Francine Moens. A dataset for multimodal question answering in the cultural heritage domain. In Proceedings of the COLING 2016 Workshop on Language Technology Resources and Tools for Digital Humanities (LT4DH), pages 10--17. ACL, 2016

  151. [160]

    Fvqa: Fact-based visual question answering

    Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. Fvqa: Fact-based visual question answering. IEEE transactions on pattern analysis and machine intelligence, 40 0 (10): 0 2413--2427, 2017

  152. [161]

    A multi-world approach to question answering about real-world scenes based on uncertain input

    Mateusz Malinowski and Mario Fritz. A multi-world approach to question answering about real-world scenes based on uncertain input. Advances in neural information processing systems, 27, 2014

  153. [162]

    Visual7w: Grounded question answering in images

    Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. Visual7w: Grounded question answering in images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4995--5004, 2016

  154. [163]

    Making the v in vqa matter: Elevating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904--6913, 2017 b

  155. [164]

    Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognitio...

  156. [165]

    Don't just assume; look and answer: Overcoming priors for visual question answering

    Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. Don't just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4971--4980, 2018

  157. [166]

    Automatic understanding of image and video advertisements

    Zaeem Hussain, Mingda Zhang, Xiaozhong Zhang, Keren Ye, Christopher Thomas, Zuha Agha, Nathan Ong, and Adriana Kovashka. Automatic understanding of image and video advertisements. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1705--1715, 2017

  158. [167]

    Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

    Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In Proceedings of the IEEE Conference on Computer Vision and Pattern rec...

  159. [168]

    Figureqa: An annotated figure dataset for visual reasoning

    Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, \'A kos K \'a d \'a r, Adam Trischler, and Yoshua Bengio. Figureqa: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300, 2017

  160. [169]

    Overview of imageclef 2018 medical domain visual question answering task

    Sadid A Hasan, Yuan Ling, Oladimeji Farri, Joey Liu, Henning M \"u ller, and Matthew P Lungren. Overview of imageclef 2018 medical domain visual question answering task. In CLEF (Working Notes), 2018 b

  161. [170]

    Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

    Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3608--3617, 201...

  162. [171]

    Tallyqa: Answering complex counting questions

    Manoj Acharya, Kushal Kafle, and Christopher Kanan. Tallyqa: Answering complex counting questions. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 8076--8084, 2019

  163. [172]

    Dvqa: Understanding data visualizations via question answering

    Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. Dvqa: Understanding data visualizations via question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5648--5656, 2018

  164. [173]

    Vqa-med: Overview of the medical visual question answering task at imageclef 2019

    Asma Ben Abacha, Sadid A Hasan, Vivek V Datla, Joey Liu, Dina Demner-Fushman, and Henning M \"u ller. Vqa-med: Overview of the medical visual question answering task at imageclef 2019. CLEF (working notes), 2 0 (6), 2019

  165. [174]

    On the general value of evidence, and bilingual scene-text visual question answering

    Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng, Canjie Luo, Lianwen Jin, Chee Seng Chan, Anton van den Hengel, and Liangwei Wang. On the general value of evidence, and bilingual scene-text visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vi...

  166. [175]

    Gqa: A new dataset for real-world visual reasoning and compositional question answering

    Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6700--6709, 2019 a

  167. [176]

    Leaf-qa: Locate, encode & attend for figure question answering

    Ritwick Chaudhry, Sumit Shekhar, Utkarsh Gupta, Pranav Maneriker, Prann Bansal, and Ajay Joshi. Leaf-qa: Locate, encode & attend for figure question answering. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3512--3521, 2020

  168. [177]

    From recognition to cognition: Visual commonsense reasoning

    Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019

  169. [178]

    A dataset and baselines for visual question answering on art

    Noa Garcia, Chentao Ye, Zihua Liu, Qingtao Hu, Mayu Otani, Chenhui Chu, Yuta Nakashima, and Teruko Mitamura. A dataset and baselines for visual question answering on art. In Computer Vision--ECCV 2020 Workshops: Glasgow, UK, August 23--28, 2020, Proceedings, Part II 16, pages ...

  170. [179]

    Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain

    Asma Ben Abacha, Vivek V Datla, Sadid A Hasan, Dina Demner-Fushman, and Henning M \"u ller. Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain. In CLEF (Working Notes), 2020

  171. [180]

    Towards visual dialog for radiology

    Olga Kovaleva, Chaitanya Shivade, Satyananda Kashyap, Karina Kanjaria, Joy Wu, Deddeh Ballah, Adam Coy, Alexandros Karargyris, Yufan Guo, David Beymer Beymer, et al. Towards visual dialog for radiology. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Proce...

  172. [181]

    Pathvqa: 30000+ questions for medical visual question answering

    Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathvqa: 30000+ questions for medical visual question answering. arXiv preprint arXiv:2003.10286, 2020 a

  173. [182]

    Khapra, and Pratyush Kumar

    Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. Plotqa: Reasoning over scientific plots. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), March 2020

  174. [183]

    The hateful memes challenge: Detecting hate speech in multimodal memes

    Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in neural information processing systems, 33: 0 2611--2624, 2020

  175. [184]

    Slake: a semantically-labeled knowledge-enhanced dataset for medical visual question answering

    Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: a semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1650--1654. IEEE, 2021 a

  176. [185]

    Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning

    Jiaqi Chen, Jianheng Tang, Jinghui Qin, Xiaodan Liang, Lingbo Liu, Eric P Xing, and Liang Lin. Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning. arXiv preprint arXiv:2105.14517, 2021

  177. [186]

    Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning

    Pan Lu, Liang Qiu, Jiaqi Chen, Tony Xia, Yizhou Zhao, Wei Zhang, Zhou Yu, Xiaodan Liang, and Song-Chun Zhu. Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning. In The 35th Conference on Neural Information Processing Systems (NeurIPS) Track...

  178. [187]

    Clevr-math: A dataset for compositional language, visual and mathematical reasoning

    Adam Dahlgren Lindstr \"o m and Savitha Sam Abraham. Clevr-math: A dataset for compositional language, visual and mathematical reasoning. arXiv preprint arXiv:2208.05358, 2022

  179. [188]

    Learn to explain: Multimodal reasoning via thought chains for science question answering

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35: 0 2507...

  180. [190]

    Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023 b

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023 b

  181. [191]

    Seed-bench: Benchmarking multimodal llms with generative comprehension

    Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023 b

  182. [192]

    Mm-vet: Evaluating large multimodal models for integrated capabilities

    Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023 b

  183. [193]

    Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A m...

  184. [194]

    Mm-vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities

    Weihao Yu, Zhengyuan Yang, Linfeng Ren, Linjie Li, Jianfeng Wang, Kevin Lin, Chung-Ching Lin, Zicheng Liu, Lijuan Wang, and Xinchao Wang. Mm-vet v2: A challenging benchmark to evaluate large multimodal models for integrated capabilities. arXiv preprint arXiv:2408.00765, 2024

  185. [195]

    Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023

  186. [196]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \'a r, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740--755. Springer, 2014

  187. [197]

    Unanswerable questions about images and texts

    Ernest Davis. Unanswerable questions about images and texts. Frontiers in Artificial Intelligence, 3: 0 51, 2020

  188. [198]

    Question relevance in vqa: identifying non-visual and false-premise questions

    Arijit Ray, Gordon Christie, Mohit Bansal, Dhruv Batra, and Devi Parikh. Question relevance in vqa: identifying non-visual and false-premise questions. arXiv preprint arXiv:1606.06622, 2016

  189. [199]

    Why did the chicken cross the road? rephrasing and analyzing ambiguous questions in vqa

    Elias Stengel-Eskin, Jimena Guallar-Blasco, Yi Zhou, and Benjamin Van Durme. Why did the chicken cross the road? rephrasing and analyzing ambiguous questions in vqa. arXiv preprint arXiv:2211.07516, 2022

  190. [200]

    Why does a visual question have different answers? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4271--4280, 2019

    Nilavra Bhattacharya, Qing Li, and Danna Gurari. Why does a visual question have different answers? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4271--4280, 2019

  191. [201]

    Know what you don't know: Unanswerable questions for squad

    Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don't know: Unanswerable questions for squad. arXiv preprint arXiv:1806.03822, 2018

  192. [202]

    Document understanding dataset and evaluation (dude)

    Jordy Van Landeghem, Rub \`e n Tito, ukasz Borchmann, Micha Pietruszka, Pawel Joziak, Rafal Powalski, Dawid Jurkiewicz, Micka \"e l Coustaty, Bertrand Anckaert, Ernest Valveny, et al. Document understanding dataset and evaluation (dude). In Proceedings of the IEEE/CVF Internat...

  193. [203]

    Question part relevance and editing for cooperative and context-aware vqa (c2vqa)

    Andeep S Toor, Harry Wechsler, and Michele Nappi. Question part relevance and editing for cooperative and context-aware vqa (c2vqa). In Proceedings of the 15th International Workshop on Content-Based Multimedia Indexing, pages 1--6, 2017

  194. [204]

    Do explanations make vqa models more predictable to a human? arXiv preprint arXiv:1810.12366, 2018

    Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. Do explanations make vqa models more predictable to a human? arXiv preprint arXiv:1810.12366, 2018

  195. [205]

    Robust visual question answering via semantic cross modal augmentation

    Akib Mashrur, Wei Luo, Nayyar A Zaidi, and Antonio Robles-Kelly. Robust visual question answering via semantic cross modal augmentation. Computer Vision and Image Understanding, page 103862, 2023

  196. [206]

    The promise of premise: Harnessing question premises in visual question answering

    Aroma Mahendru, Viraj Prabhu, Akrit Mohapatra, Dhruv Batra, and Stefan Lee. The promise of premise: Harnessing question premises in visual question answering. arXiv preprint arXiv:1705.00601, 2017

  197. [207]

    Wordnet: a lexical database for english

    George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 38 0 (11): 0 39--41, 1995

  198. [208]

    A dataset of clinically generated visual questions and answers about radiology images

    Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images. Scientific data, 5 0 (1): 0 1--10, 2018 b

  199. [209]

    Medical visual question answering via conditional reasoning

    Li-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen, and Xiao-Ming Wu. Medical visual question answering via conditional reasoning. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2345--2354, 2020

  200. [210]

    Multiple meta-model quantifying for medical visual question answering

    Tuong Do, Binh X Nguyen, Erman Tjiputra, Minh Tran, Quang D Tran, and Anh Nguyen. Multiple meta-model quantifying for medical visual question answering. In Medical Image Computing and Computer Assisted Intervention--MICCAI 2021: 24th International Conference, Strasbourg, Franc...

  201. [211]

    Overcoming data limitation in medical visual question answering

    Binh D Nguyen, Thanh-Toan Do, Binh X Nguyen, Tuong Do, Erman Tjiputra, and Quang D Tran. Overcoming data limitation in medical visual question answering. In Medical Image Computing and Computer Assisted Intervention--MICCAI 2019: 22nd International Conference, Shenzhen, China,...

  202. [212]

    Contrastive pre-training and representation distillation for medical visual question answering based on radiology images

    Bo Liu, Li-Ming Zhan, and Xiao-Ming Wu. Contrastive pre-training and representation distillation for medical visual question answering based on radiology images. In Medical Image Computing and Computer Assisted Intervention--MICCAI 2021: 24th International Conference, Strasbou...

  203. [213]

    Umass at imageclef medical visual question answering (med-vqa) 2018 task

    Yalei Peng, Feifan Liu, and Max P Rosen. Umass at imageclef medical visual question answering (med-vqa) 2018 task. In CLEF (Working Notes), 2018

  204. [214]

    Vizwiz grand challenge: Answering visual questions from blind people

    Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3608--3...

  205. [215]

    Oscar: Object-semantics aligned pre-training for vision-language tasks

    Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision, pages 121--137. Springer, 2020 a

  206. [216]

    Visual question answering from remote sensing images

    Sylvain Lobry, Jesse Murray, Diego Marcos, and Devis Tuia. Visual question answering from remote sensing images. In IGARSS 2019-2019 IEEE International Geoscience and Remote Sensing Symposium, pages 4951--4954. IEEE, 2019

  207. [217]

    Rsvqa: Visual question answering for remote sensing data

    Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia. Rsvqa: Visual question answering for remote sensing data. IEEE Transactions on Geoscience and Remote Sensing, 58 0 (12): 0 8555--8566, 2020

  208. [218]

    How to find a good image-text embedding for remote sensing visual question answering? arXiv preprint arXiv:2109.11848, 2021

    Christel Chappuis, Sylvain Lobry, Benjamin Kellenberger, Bertrand Le Saux, and Devis Tuia. How to find a good image-text embedding for remote sensing visual question answering? arXiv preprint arXiv:2109.11848, 2021

  209. [219]

    Cross-modal visual question answering for remote sensing data: The international conference on digital image computing: Techniques and applications (dicta 2021)

    Rafael Felix, Boris Repasky, Samuel Hodge, Reza Zolfaghari, Ehsan Abbasnejad, and Jamie Sherrah. Cross-modal visual question answering for remote sensing data: The international conference on digital image computing: Techniques and applications (dicta 2021). In 2021 Digital Im...

  210. [220]

    Bi-modal transformer-based approach for visual question answering in remote sensing imagery

    Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Mohamed Lamine Mekhalfi, Mansour Abdulaziz Al Zuair, and Farid Melgani. Bi-modal transformer-based approach for visual question answering in remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 60: 0 1--11, 2022

  211. [221]

    Mutual attention inception network for remote sensing visual question answering

    Xiangtao Zheng, Binqiang Wang, Xingqian Du, and Xiaoqiang Lu. Mutual attention inception network for remote sensing visual question answering. IEEE Transactions on Geoscience and Remote Sensing, 60: 0 1--14, 2021

  212. [222]

    A spatial hierarchical reasoning network for remote sensing visual question answering

    Zixiao Zhang, Licheng Jiao, Lingling Li, Xu Liu, Puhua Chen, Fang Liu, Yuxuan Li, and Zhicheng Guo. A spatial hierarchical reasoning network for remote sensing visual question answering. IEEE Transactions on Geoscience and Remote Sensing, 2023 a

  213. [223]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020

  214. [224]

    Recommending themes for ad creative design via visual-linguistic representations

    Yichao Zhou, Shaunak Mishra, Manisha Verma, Narayan Bhamidipati, and Wei Wang. Recommending themes for ad creative design via visual-linguistic representations. In Proceedings of The Web Conference 2020, pages 2521--2527, 2020

  215. [225]

    A diagram is worth a dozen images

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14, pages 235--25...

  216. [226]

    Isaaq--mastering textbook questions with pre-trained transformers and bottom-up and top-down attention

    Jose Manuel Gomez-Perez and Raul Ortega. Isaaq--mastering textbook questions with pre-trained transformers and bottom-up and top-down attention. arXiv preprint arXiv:2010.00562, 2020

  217. [227]

    Moqa-a multi-modal question answering architecture

    Monica Haurilet, Ziad Al-Halah, and Rainer Stiefelhagen. Moqa-a multi-modal question answering architecture. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0--0, 2018

  218. [228]

    Textbook question answering under instructor guidance with memory networks

    Juzheng Li, Hang Su, Jun Zhu, Siyu Wang, and Bo Zhang. Textbook question answering under instructor guidance with memory networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3655--3663, 2018 a

  219. [229]

    Spatial-semantic collaborative graph network for textbook question answering

    Yaxian Wang, Bifan Wei, Jun Liu, Qika Lin, Lingling Zhang, and Yaqiang Wu. Spatial-semantic collaborative graph network for textbook question answering. IEEE Transactions on Circuits and Systems for Video Technology, 33 0 (7): 0 3214--3228, 2023. doi:10.1109/TCSVT.2022.3231463

  220. [230]

    Weakly supervised learning for textbook question answering

    Jie Ma, Qi Chai, Jingyue Huang, Jun Liu, Yang You, and Qinghua Zheng. Weakly supervised learning for textbook question answering. IEEE Transactions on Image Processing, 31: 0 7378--7388, 2022

  221. [231]

    Essay-anchor attentive multi-modal bilinear pooling for textbook question answering

    Juzheng Li, Hang Su, Jun Zhu, and Bo Zhang. Essay-anchor attentive multi-modal bilinear pooling for textbook question answering. In 2018 IEEE International Conference on Multimedia and Expo (ICME), pages 1--6. IEEE, 2018 b

  222. [232]

    Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression

    Jiaqi Chen, Tong Li, Jinghui Qin, Pan Lu, Liang Lin, Chongyu Chen, and Xiaodan Liang. Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression. arXiv preprint arXiv:2212.02746, 2022 b

  223. [233]

    A multi-modal neural geometric solver with textual clauses parsed from diagram

    Ming-Liang Zhang, Fei Yin, and Cheng-Lin Liu. A multi-modal neural geometric solver with textual clauses parsed from diagram. arXiv preprint arXiv:2302.11097, 2023 b

  224. [234]

    An augmented benchmark dataset for geometric question answering through dual parallel text encoding

    Jie Cao and Jing Xiao. An augmented benchmark dataset for geometric question answering through dual parallel text encoding. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1511--1520, Gyeongju, Republic of Korea, October 2022. Internatio...

  225. [235]

    An educational robot system of visual question answering for preschoolers

    Bin He, Meng Xia, Xinguo Yu, Pengpeng Jian, Hao Meng, and Zhanwen Chen. An educational robot system of visual question answering for preschoolers. In 2017 2nd international conference on robotics and automation engineering (ICRAE), pages 441--445. IEEE, 2017

  226. [236]

    An empirical study of training end-to-end vision-and-language transformers

    Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al. An empirical study of training end-to-end vision-and-language transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...

  227. [237]

    Abc-cnn: An attention based convolutional neural network for visual question answering

    Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia. Abc-cnn: An attention based convolutional neural network for visual question answering. arXiv preprint arXiv:1511.05960, 2015

  228. [238]

    Pythia v0

    Yu Jiang, Vivek Natarajan, Xinlei Chen, Marcus Rohrbach, Dhruv Batra, and Devi Parikh. Pythia v0. 1: the winning entry to the vqa challenge 2018. arXiv preprint arXiv:1807.09956, 2018

  229. [239]

    Proto: Program-guided transformer for program-guided tasks

    Zelin Zhao, Karan Samel, Binghong Chen, et al. Proto: Program-guided transformer for program-guided tasks. Advances in Neural Information Processing Systems, 34: 0 17021--17036, 2021

  230. [240]

    Learning conditioned graph structures for interpretable visual question answering

    Will Norcliffe-Brown, Stathis Vafeias, and Sarah Parisot. Learning conditioned graph structures for interpretable visual question answering. Advances in neural information processing systems, 31, 2018

  231. [241]

    Multi-modality latent interaction network for visual question answering

    Peng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang, and Hongsheng Li. Multi-modality latent interaction network for visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5825--5835, 2019

  232. [242]

    Chain of reasoning for visual question answering

    Chenfei Wu, Jinlai Liu, Xiaojie Wang, and Xuan Dong. Chain of reasoning for visual question answering. Advances in Neural Information Processing Systems, 31, 2018

  233. [243]

    Relation-aware graph attention network for visual question answering

    Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. Relation-aware graph attention network for visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10313--10322, 2019 b

  234. [244]

    Learning by abstraction: The neural state machine

    Drew Hudson and Christopher D Manning. Learning by abstraction: The neural state machine. Advances in Neural Information Processing Systems, 32, 2019 b

  235. [245]

    Pixel-bert: Aligning image pixels with text by deep multi-modal transformers

    Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849, 2020

  236. [246]

    Uniter: Universal image-text representation learning

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104--120. Springer, 2020

  237. [247]

    How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021

    Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. How much can clip benefit vision-and-language tasks? arXiv preprint arXiv:2107.06383, 2021

  238. [248]

    Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning

    Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. Unimo: Towards unified-modal understanding and generation via cross-modal contrastive learning. arXiv preprint arXiv:2012.15409, 2020 b

  239. [249]

    Vinvl: Revisiting visual representations in vision-language models

    Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5579--5588, 2021

  240. [250]

    Unifying vision-and-language tasks via text generation

    Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning, pages 1931--1942. PMLR, 2021

  241. [251]

    Mdetr-modulated detection for end-to-end multi-modal understanding

    Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1780--1790, 2021

  242. [252]

    Ernie-vil: Knowledge enhanced vision-language representations through scene graphs

    Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 3208--3216, 2021

  243. [253]

    Florence: A new foundation model for computer vision

    Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432, 2021

  244. [254]

    Unimo-2: End-to-end unified vision-language grounded learning

    Wei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, and Haifeng Wang. Unimo-2: End-to-end unified vision-language grounded learning. arXiv preprint arXiv:2203.09067, 2022 a

  245. [255]

    Git: A generative image-to-text transformer for vision and language

    Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang. Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022 c

  246. [256]

    Image as a foreign language: Beit pretraining for all vision and vision-language tasks

    Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022 d

  247. [257]

    High-performance large-scale image recognition without normalization

    Andy Brock, Soham De, Samuel L Smith, and Karen Simonyan. High-performance large-scale image recognition without normalization. In International Conference on Machine Learning, pages 1059--1071. PMLR, 2021

  248. [258]

    Pali: A jointly-scaled multilingual language-image model

    Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022 c

  249. [259]

    mplug: Effective and efficient vision-language learning by cross-modal skip-connections

    Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, Hehong Chen, Guohai Xu, Zheng Cao, et al. mplug: Effective and efficient vision-language learning by cross-modal skip-connections. arXiv preprint arXiv:2205.12005, 2022 b

  250. [260]

    C hart QA : A benchmark for question answering about charts with visual and logical reasoning

    Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. C hart QA : A benchmark for question answering about charts with visual and logical reasoning. In Smaranda Muresan and Aline Nakov, Preslav andbai2023qwen Villavicencio, editors, Findings of the Associatio...

  251. [261]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296--26306, 2024 b

  252. [262]

    Judging llm-as-a-judge with mt-bench and chatbot arena

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36: 0 46595--46623, 2023

  253. [263]

    Electra: Pre-training text encoders as discriminators rather than generators

    Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020

  254. [264]

    Albert: A lite bert for self-supervised learning of language representations

    Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2019

  255. [265]

    Deberta: Decoding-enhanced bert with disentangled attention

    Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: Decoding-enhanced bert with disentangled attention. arXiv preprint arXiv:2006.03654, 2020 b

  256. [266]

    Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering

    Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao. Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering. IEEE transactions on neural networks and learning systems, 29 0 (12): 0 5947--5959, 2018

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.