REVIEW 3 major objections 5 minor 229 references
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This survey claims to be the first to organize vision-language methods that use pre-trained models according to the classic challenge each method tackles, and it backs the map with comparative tables across images and videos.
desk verdict A useful survey with a sensible challenge-based taxonomy, despite a few citation slip-ups and an asserted rather than argued choice of four challenges. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the survey's challenge-based taxonomy itself, laid out in Table II. Each of the four challenges anchors a family of paradigms: data scarcity is attacked by direct inference, uni-modal training, or pseudo-pair generation; escalating reasoning complexity is attacked by divide-and-conquer or chain-of-thought decomposition; generalization to novel samples is attacked by extracting semantic context from a language model or distilling teacher knowledge from a vision-language model; and task diversity is attacked by continual learning or by planning with natural language or code statements. Illustrated pipeline diagrams make each paradigm concrete, and the benchmark tables anchor each family to measured performance.
What would settle it
Find a vision-language method whose stated motivation is purely efficiency or safety—cutting inference cost or preventing harmful outputs—with no dependence on the four challenges; its absence from the survey's taxonomy would show that the challenge set is not exhaustive.
Extended reading notes
Core claim
The paper's central claim is that vision-language research today is best understood as a set of responses to four classic challenges, and that pre-trained models supply the capabilities—language priors, a shared image-text space, world knowledge, and in-context flexibility—that let those responses work where earlier methods failed. It argues that before pre-training, each challenge was only partially addressed: semi- and weakly supervised methods overfitted, fixed multi-step reasoners could not scale their reasoning hops, knowledge-base lookups were rigid and thin, and multi-task models suffered catastrophic forgetting. Pre-trained models change the options by enabling direct inference on test samples, training from unlabeled uni-modal data through the CLIP common space, generating pseudo-paired data, divide-and-conquer and chain-of-thought reasoning, extracting semantic context from language models, distilling teacher knowledge from vision-language models, continual learning, and language-model-planned tool use. The paper supports the map with comparative tables showing that methods in each paradigm beat their pre-training-era baselines.
Load-bearing premise
The taxonomy holds only if the four challenges—data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity—are the main bottlenecks and are distinct enough that each method can be filed under one of them.
Editorial extensions
If this is right
- If the survey's map is right, a practitioner short on annotated data has three proven routes: direct inference with an LLM plus CLIP, text-only training through the CLIP common space, or generating pseudo-paired data, with the pseudo-pair route approaching fully supervised captioning scores.
- Decomposing questions or reasoning paths—divide-and-conquer or chain-of-thought—consistently beats one-step reasoning on OK-VQA, A-OKVQA, VCR, SNLI-VE, and ScienceQA, so complexity should be attacked by decomposition rather than by bigger single-step models.
- For novel samples, querying an LLM for class descriptions improves CLIP's open-vocabulary classification, and distilling a VLM into a close-set detector raises novel-class average precision while keeping base-class performance.
- A single general system with an LLM planner calling tools can cover many tasks zero-shot, making task-specific training unnecessary for the covered tasks.
- The same pre-trained capabilities carry risks—hallucination, outdated knowledge, concept-association bias, and compositional confusion—so methods that use them should include verification, retrieval, or code execution to counter those risks.
Reading between the lines
- Editorial inference: the four-challenge frame doubles as a design checklist; a new vision-language method can be positioned by naming the bottleneck it attacks, which suggests the taxonomy will seed future method papers even if its boundaries blur.
- Editorial inference: the survey's risk section implies that the next performance gains will come from pairing methods—for example, LLM semantic context to compensate what VLM distillation loses, or retrieval to fix outdated knowledge—though the paper only hints at such combinations.
- Editorial inference: one could test the taxonomy's completeness by checking whether methods motivated purely by efficiency or safety (e.g., inference-cost reduction or harm avoidance) can be classified; the survey does not cover such motivations.
- Editorial inference: since the four challenges overlap in practice—novel samples often require reasoning, and task diversity is a generalization problem across tasks—a quantitative study measuring how single methods score on more than one challenge could refine or merge the categories.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey organizes recent methods that integrate large pre-trained models (both LLMs and VLMs) into vision-language tasks around four challenges: data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity. For each challenge, it reviews the corresponding paradigms (e.g., direct inference, chain-of-thought, knowledge distillation, LLM-as-planner), illustrates them with pipeline figures, and collects performance tables for image captioning, complex reasoning benchmarks, open-vocabulary classification and detection, and general modular systems. The paper closes with a discussion of four risks introduced by pre-trained models (hallucination, outdated knowledge, concept association bias, and compositional concept confusion) and suggests mitigation strategies. The central claim, stated in Section I, is that this is the first survey focused specifically on how vision-language tasks benefit from large pre-trained models and the first to categorize methods according to the challenges they tackle.
Significance. If the taxonomy is accepted, the survey provides a useful organizational map of a fast-moving area and a convenient entry point for researchers: the pipeline diagrams in Figures 2 through 5 are pedagogical, and the performance tables make cross-paradigm comparisons accessible. The paper covers both images and videos, both discriminative and generative models, and includes a thoughtful discussion of risks that is often absent from method-oriented surveys. The comparisons in Tables III-VI are concrete and falsifiable in that they name specific methods and benchmarks. The main value is synthetic: it brings together method families that are usually scattered across separate papers and frames them under a small set of recurring challenges. No machine-checked artifacts are shipped, but the claims are checkable against the cited literature.
major comments (3)
- [Section III and Table II] The paper's central contribution is the challenge-based taxonomy, but the taxonomy is asserted rather than validated, and the assignment of methods to exactly one challenge is often ambiguous. For instance, ZeroCap [61] is filed under 'Data Scarcity – Direct inference on test samples,' yet its goal is to caption images never seen during training, which overlaps with the definition of 'Generalization to Novel Samples' given in Section III-C; likewise, the pseudo-paired-data methods [84]-[86] create synthetic samples that address data scarcity and simultaneously improve robustness to novel samples. The survey should either provide a concrete decision rule for assigning methods, allow a method to appear under multiple challenges, or explicitly frame the four challenges as one useful partition rather than the 'main challenges' faced by all models. Without this, the claimed novelty of categorizing methods by challenge is not fully supported.
- [Section III and Section VI] The paper claims that the four challenges are the 'main challenges' and describes the survey as comprehensive, but the taxonomy omits challenges that are prominent in current vision-language research, such as alignment and safety, computational efficiency, and robustness beyond novel-sample generalization. The four risks discussed in Section VI (hallucination, outdated knowledge, concept association bias, and compositional concept confusion) are not mapped back to the challenge taxonomy, so it is unclear how these risks interact with the four categories. This omission is not fatal if the paper narrows its scope, but as written the 'comprehensive' claim in the abstract and Section I is stronger than the taxonomy supports.
- [Section V-D] The grouping of continual learning and LLM-as-planner under 'Task Diversity' is not self-evident. Continual learning addresses catastrophic forgetting and parameter efficiency across a sequence of tasks, not primarily the diversity of input-output workflows; LLM-as-planner systems address compositional reasoning, modularity, and tool use. The survey should explain why these paradigms are classified under task diversity rather than under reasoning complexity or generalization, and should state how 'task diversity' is distinguished from the other three challenges.
minor comments (5)
- [Section V-C and Table V] The sentence 'Table V reports the comparison results of methods [9], [10], [12], [203]' is inconsistent with the actual content of Table V, which lists only VCD [9], LCDAtt [10], and CPHC [12]; reference [203] (OVR-CNN) appears in Table VI as an object-detection baseline and is not an image-classification method. Please correct the text or add the missing row to Table V.
- [Section II] The sentence 'Table I shows the differences between our survey and the existing related surveys in terms of content and coverage' is immediately repeated with slightly different wording; the duplicate should be removed.
- [Table VII] In Table VII, the planning-format entry for MM-VID reads 'natrual lang', which appears to be a typo for 'natural lang'.
- [Section VI-A] The benchmark name 'HALLUSIONBENCH' should be written as 'HallusionBench' to match the cited paper and standard usage.
- [Header and metadata] The page header 'JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021' appears to be a leftover from a template, since the arXiv submission is dated December 2024 and the actual journal/volume information is not provided; this should be corrected or removed.
Circularity Check
No material circularity; the survey contains only peripheral self-citations used as method examples, and its central taxonomy is asserted rather than derived, which is a validity concern, not a circularity concern.
full rationale
This is a survey paper, not a derivation or prediction pipeline. There is no fitted parameter that is later relabeled as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in through citation. The paper's central claim is that it is pioneering in focusing on how vision-language tasks benefit from pre-trained models and in categorizing methods by the challenges they tackle (Section I: 'To the best of our knowledge, this survey is pioneering in its focus on this topic and in categorizing methods according to the challenges they tackle.'). That claim is a novelty assertion, not a derived result, so it cannot be circular in the sense of reducing to its own input. The four challenges in Section III (data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity) are organizing categories introduced by the survey; the paper does not pretend to derive them from first principles. Whether the taxonomy is exhaustive or whether individual methods are assigned to the right bucket is a correctness and completeness concern, not a circularity concern. The only self-citations by the survey authors are reference [195] (Y. Qi, W. Zhao, and X. Wu, 'Relational distant supervision for image captioning without image-text pairs') and reference [204] (Y. Shi, X. Wu, H. Lin, and J. Luo, 'Commonsense knowledge prompting for few-shot action recognition in videos'). Both are cited as examples of existing methods: [195] appears in the summary of pre-pretraining-era approaches as 'constructing pseudo-paired data based on object relationships [195]', and [204] appears in Section V-C as 'Shi et al. [204] enhance the semantics of action categories by using BERT [52] to collect text proposals containing language descriptions of actions.' Neither citation is used to justify the survey's novelty claim, its taxonomy, or any evaluative conclusion; both are ordinary literature pointers within a survey. Consequently, these self-citations are not load-bearing and do not raise the circularity score beyond 1. The paper is also self-contained against external benchmarks: Tables III-VI compare independently published methods on standard datasets, and the survey's discussion of risks is drawn from external studies. Therefore, the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Vision-language tasks face exactly four main challenges: data scarcity, escalating reasoning complexity, generalization to novel samples, and task diversity.
- domain assumption The benchmark numbers reported in the cited papers are accurate and comparable across the tables in this survey.
- domain assumption Large pre-trained models possess transferable capabilities, including zero-shot inference, stored world knowledge, and a shared vision-language embedding space.
Cite this review
Pith. "Pith review of How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey." pith.science (2026). https://pith.science/paper/GYLNSXOC
@misc{pith2026241208158,
author = {Pith},
title = {Pith review of: How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/GYLNSXOC}},
note = {Machine review of arXiv:2412.08158}
}
read the original abstract
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelligence and continuously attracts the research community's attention. Despite the improvements in overall performance, classic challenges still exist in vision-language tasks and hinder the development of this area. In recent years, the rise of pre-trained models is driving the research on vision-language tasks. Thanks to the massive scale of training data and model parameters, pre-trained models have exhibited excellent performance in numerous downstream tasks. Inspired by the powerful capabilities of pre-trained models, new paradigms have emerged to solve the classic challenges. Such methods have become mainstream in current research with increasing attention and rapid advances. In this paper, we present a comprehensive overview of how vision-language tasks benefit from pre-trained models. First, we review several main challenges in vision-language tasks and discuss the limitations of previous solutions before the era of pre-training. Next, we summarize the recent advances in incorporating pre-trained models to address the challenges in vision-language tasks. Finally, we analyze the potential risks associated with the inherent limitations of pre-trained models and discuss possible solutions, attempting to provide future research directions.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[9]
Visual classification via description from large language models,
S. Menon and C. V ondrick, “Visual classification via description from large language models,” in The Eleventh International Conference on Learning Representations, 2022
2022
-
[10]
Learning concise and descriptive attributes for visual recognition,
A. Yan, Y . Wang, Y . Zhong, C. Dong, Z. He, Y . Lu, W. Y . Wang, J. Shang, and J. McAuley, “Learning concise and descriptive attributes for visual recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3090–3100
2023
-
[12]
Chatgpt-powered hierarchical comparisons for image classification,
Z. Ren, Y . Su, and X. Liu, “Chatgpt-powered hierarchical comparisons for image classification,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[203]
Open-vocabulary object detection using captions,
A. Zareian, K. D. Rosa, D. H. Hu, and S.-F. Chang, “Open-vocabulary object detection using captions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 393–14 402
2021
-
[110]
Open-vocabulary object detec- tion via vision and language knowledge distillation,
X. Gu, T.-Y . Lin, W. Kuo, and Y . Cui, “Open-vocabulary object detec- tion via vision and language knowledge distillation,” in International Conference on Learning Representations , 2022
2022
-
[114]
Open- vocabulary object detection upon frozen vision and language models,
W. Kuo, Y . Cui, X. Gu, A. Piergiovanni, and A. Angelova, “Open- vocabulary object detection upon frozen vision and language models,” in The Eleventh International Conference on Learning Representations, 2023
2023
-
[115]
Cohoz: Contrastive multimodal prompt tuning for hierarchical open- set zero-shot recognition,
N. Liao, Y . Liu, L. Xiaobo, C. Lei, G. Wang, X.-S. Hua, and J. Yan, “Cohoz: Contrastive multimodal prompt tuning for hierarchical open- set zero-shot recognition,” in Proceedings of the 30th ACM Interna- tional Conference on Multimedia , 2022, pp. 3262–3271
2022
-
[117]
Lmc: Large model collaboration with cross-assessment for training-free open-set object recognition,
H. Qu, X. Hui, Y . Cai, and J. Liu, “Lmc: Large model collaboration with cross-assessment for training-free open-set object recognition,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[61]
Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,
Y . Tewel, Y . Shalev, I. Schwartz, and L. Wolf, “Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic,” in Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 918–17 928
2022
-
[84]
Image captioning with multi-context synthetic data,
F. Ma, Y . Zhou, F. Rao, Y . Zhang, and X. Sun, “Image captioning with multi-context synthetic data,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 5, 2024, pp. 4089–4097
2024
-
[86]
Freemask: Synthetic images with dense annotations make stronger segmentation models,
L. Yang, X. Xu, B. Kang, Y . Shi, and H. Zhao, “Freemask: Synthetic images with dense annotations make stronger segmentation models,” arXiv preprint arXiv:2310.15160 , 2023
arXiv 2023
Show all 229 references
-
[1]
Multimodal research in vision and language: A review of current and emerging trends,
S. Uppal, S. Bhagat, D. Hazarika, N. Majumder, S. Poria, R. Zimmer- mann, and A. Zadeh, “Multimodal research in vision and language: A review of current and emerging trends,” Information Fusion , vol. 77, pp. 149–171, 2022
2022
-
[2]
Vision+ language applications: A survey,
Y . Zhou and N. Shimada, “Vision+ language applications: A survey,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 826–842
2023
-
[3]
Bottom-up and top-down attention for image captioning and visual question answering,
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6077–6086
2018
-
[4]
Swinbert: End-to-end transformers with sparse attention for video captioning,
K. Lin, L. Li, C.-C. Lin, F. Ahmed, Z. Gan, Z. Liu, Y . Lu, and L. Wang, “Swinbert: End-to-end transformers with sparse attention for video captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 949–17 958
2022
-
[5]
Vqa: Visual question answering,
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “Vqa: Visual question answering,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. 2425– 2433
2015
-
[6]
Movieqa: Understanding stories in movies through question- answering,
M. Tapaswi, Y . Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler, “Movieqa: Understanding stories in movies through question- answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4631–4640
2016
-
[7]
Context-aware attention network for image-text retrieval,
Q. Zhang, Z. Lei, Z. Zhang, and S. Z. Li, “Context-aware attention network for image-text retrieval,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 3536–3545
2020
-
[8]
Fine-grained video-text retrieval with hierarchical graph reasoning,
S. Chen, Y . Zhao, Q. Jin, and Q. Wu, “Fine-grained video-text retrieval with hierarchical graph reasoning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 10 638–10 647
2020
-
[11]
Exploring large language models for multi-modal out-of-distribution detection,
Y . Dai, H. Lang, K. Zeng, F. Huang, and Y . Li, “Exploring large language models for multi-modal out-of-distribution detection,” in The 2023 Conference on Empirical Methods in Natural Language Processing, 2023
2023
-
[13]
Vision-language pre-training: Basics, recent advances, and future trends,
Z. Gan, L. Li, C. Li, L. Wang, Z. Liu, J. Gao et al., “Vision-language pre-training: Basics, recent advances, and future trends,” Foundations and Trends® in Computer Graphics and Vision , vol. 14, no. 3–4, pp. 163–352, 2022. JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 19
2022
-
[14]
Show and tell: A neural image caption generator,
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” inProceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164
2015
-
[15]
Deep correlation for matching images and text,
F. Yan and K. Mikolajczyk, “Deep correlation for matching images and text,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3441–3450
2015
-
[16]
Draw: A recurrent neural network for image generation,
K. Gregor, I. Danihelka, A. Graves, D. Rezende, and D. Wierstra, “Draw: A recurrent neural network for image generation,” in Interna- tional conference on machine learning. PMLR, 2015, pp. 1462–1471
2015
-
[17]
Memory- attended recurrent network for video captioning,
W. Pei, J. Zhang, X. Wang, L. Ke, X. Shen, and Y .-W. Tai, “Memory- attended recurrent network for video captioning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 8347–8356
2019
-
[18]
Heterogeneous memory enhanced multimodal attention model for video question answering,
C. Fan, X. Zhang, S. Zhang, W. Wang, C. Zhang, and H. Huang, “Heterogeneous memory enhanced multimodal attention model for video question answering,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 1999–2007
2019
-
[19]
Mocogan: Decompos- ing motion and content for video generation,
S. Tulyakov, M.-Y . Liu, X. Yang, and J. Kautz, “Mocogan: Decompos- ing motion and content for video generation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1526–1535
2018
-
[20]
Rethinking the bottom-up framework for query-based video localiza- tion,
L. Chen, C. Lu, S. Tang, J. Xiao, D. Zhang, C. Tan, and X. Li, “Rethinking the bottom-up framework for query-based video localiza- tion,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 10 551–10 558
2020
-
[21]
Object detection in 20 years: A survey,
Z. Zou, K. Chen, Z. Shi, Y . Guo, and J. Ye, “Object detection in 20 years: A survey,” Proceedings of the IEEE , vol. 111, no. 3, pp. 257– 276, 2023
2023
-
[22]
Neural motifs: Scene graph parsing with global context,
R. Zellers, M. Yatskar, S. Thomson, and Y . Choi, “Neural motifs: Scene graph parsing with global context,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 5831–5840
2018
-
[23]
From recognition to cognition: Visual commonsense reasoning,
R. Zellers, Y . Bisk, A. Farhadi, and Y . Choi, “From recognition to cognition: Visual commonsense reasoning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6720–6731
2019
-
[24]
Visual entailment: A novel task for fine-grained image understanding,
N. Xie, F. Lai, D. Doran, and A. Kadav, “Visual entailment: A novel task for fine-grained image understanding,” arXiv preprint arXiv:1901.06706, 2019
1901 arXiv
-
[25]
The abduction of sherlock holmes: A dataset for visual abductive reasoning,
J. Hessel, J. D. Hwang, J. S. Park, R. Zellers, C. Bhagavatula, A. Rohrbach, K. Saenko, and Y . Choi, “The abduction of sherlock holmes: A dataset for visual abductive reasoning,” in European Con- ference on Computer Vision . Springer, 2022, pp. 558–575
2022
-
[26]
Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,
N. Bitton-Guetta, Y . Bitton, J. Hessel, L. Schmidt, Y . Elovici, G. Stanovsky, and R. Schwartz, “Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional im- ages,” in Proceedings of the IEEE/CVF International Conference on Computer Vision...
2023
-
[27]
Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,
S. Zhong, Z. Huang, S. Gao, W. Wen, L. Lin, M. Zitnik, and P. Zhou, “Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2024
2024
-
[28]
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys, vol. 55, no. 9, pp. 1–35, 2023
2023
-
[29]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi `ere, N. Goyal, E. Hambro, F. Azhar et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[30]
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
-
[31]
Qwen technical report,
J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huang et al. , “Qwen technical report,” arXiv preprint arXiv:2309.16609, 2023
2023 arXiv
-
[32]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
-
[33]
Scaling up visual and vision-language representation learning with noisy text supervision,
C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in International conference on machine learning . PMLR, 2021, pp. 4904–4916
2021
-
[34]
Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre- training with mixture-of-modality-experts,” Advances in Neural Infor- mation Processing Systems , vol. 35, pp. 32 897–32 912, 2022
2022
-
[35]
Filip: Fine-grained interactive language-image pre-training,
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu, “Filip: Fine-grained interactive language-image pre-training,” arXiv preprint arXiv:2111.07783 , 2021
2021 arXiv
-
[36]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International conference on machine learning . PMLR, 2022, pp. 12 888–12 900
2022
-
[37]
Minigpt-4: En- hancing vision-language understanding with advanced large language models,
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, “Minigpt-4: En- hancing vision-language understanding with advanced large language models,” arXiv preprint arXiv:2304.10592 , 2023
2023 arXiv
-
[38]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, 2023
2023
-
[39]
Video-llama: An instruction-tuned audio-visual language model for video understanding,
H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” arXiv preprint arXiv:2306.02858, 2023
2023 arXiv
-
[40]
A survey of large language models,
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong et al. , “A survey of large language models,” arXiv preprint arXiv:2303.18223 , 2023
2023 arXiv
-
[41]
A comprehensive overview of large language models,
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Barnes, and A. Mian, “A comprehensive overview of large language models,” arXiv preprint arXiv:2307.06435 , 2023
2023 arXiv
-
[42]
A survey on evaluation of large language models,
Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang et al. , “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology , vol. 15, no. 3, pp. 1–45, 2024
2024
-
[43]
Large-scale multi-modal pre-trained models: A com- prehensive survey,
X. Wang, G. Chen, G. Qian, P. Gao, X.-Y . Wei, Y . Wang, Y . Tian, and W. Gao, “Large-scale multi-modal pre-trained models: A com- prehensive survey,” Machine Intelligence Research, vol. 20, no. 4, pp. 447–482, 2023
2023
-
[44]
Foundational models defining a new era in vision: A survey and outlook,
M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M.-H. Yang, and F. S. Khan, “Foundational models defining a new era in vision: A survey and outlook,” arXiv preprint arXiv:2307.13721 , 2023
2023 arXiv
-
[45]
A survey on multimodal large language models,
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen, “A survey on multimodal large language models,” arXiv preprint arXiv:2306.13549, 2023
2023 arXiv
-
[46]
Mm-llms: Recent advances in multimodal large language models,
D. Zhang, Y . Yu, C. Li, J. Dong, D. Su, C. Chu, and D. Yu, “Mm-llms: Recent advances in multimodal large language models,” arXiv preprint arXiv:2401.13601, 2024
2024 arXiv
-
[47]
Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,
Y . Wang, W. Chen, X. Han, X. Lin, H. Zhao, Y . Liu, B. Zhai, J. Yuan, Q. You, and H. Yang, “Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,” arXiv preprint arXiv:2401.06805 , 2024
2024 arXiv
-
[48]
Video understanding with large language models: A survey,
Y . Tang, J. Bi, S. Xu, L. Song, S. Liang, T. Wang, D. Zhang, J. An, J. Lin, R. Zhu et al., “Video understanding with large language models: A survey,” arXiv preprint arXiv:2312.17432 , 2023
2023
-
[49]
Vision-language models for vision tasks: A survey,
J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-language models for vision tasks: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024
2024
-
[50]
Llms meet multimodal generation and editing: A survey,
Y . He, Z. Liu, J. Chen, Z. Tian, H. Liu, X. Chi, R. Liu, R. Yuan, Y . Xing, W. Wang et al. , “Llms meet multimodal generation and editing: A survey,” arXiv preprint arXiv:2405.19334 , 2024
2024 arXiv
-
[51]
Generalized out-of-distribution detection and beyond in vision language model era: A survey,
A. Miyai, J. Yang, J. Zhang, Y . Ming, Y . Lin, Q. Yu, G. Irie, S. Joty, Y . Li, H. Li et al. , “Generalized out-of-distribution detection and beyond in vision language model era: A survey,” arXiv preprint arXiv:2407.21794, 2024
2024 arXiv
-
[52]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805, 2018
2018 arXiv
-
[53]
Ernie: Enhanced representation through knowledge integration,
Y . Sun, S. Wang, Y . Li, S. Feng, X. Chen, H. Zhang, X. Tian, D. Zhu, H. Tian, and H. Wu, “Ernie: Enhanced representation through knowledge integration,” arXiv preprint arXiv:1904.09223 , 2019
1904 arXiv
-
[54]
Language mod- els are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language mod- els are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[55]
Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach,
D.-J. Kim, J. Choi, T.-H. Oh, and I. S. Kweon, “Image captioning with very scarce supervised data: Adversarial semi-supervised learning approach,” arXiv preprint arXiv:1909.02201 , 2019
1909 arXiv
-
[56]
Semi-supervised cross-modal retrieval with label prediction,
D. Mandal, P. Rao, and S. Biswas, “Semi-supervised cross-modal retrieval with label prediction,” IEEE Transactions on Multimedia , vol. 22, no. 9, pp. 2345–2353, 2019. JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 20
2019
-
[57]
Weakly supervised dense event captioning in videos,
X. Duan, W. Huang, C. Gan, J. Wang, W. Zhu, and J. Huang, “Weakly supervised dense event captioning in videos,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
-
[58]
Weakly-supervised visual-retriever-reader for knowledge-based question answering,
M. Luo, Y . Zeng, P. Banerjee, and C. Baral, “Weakly-supervised visual-retriever-reader for knowledge-based question answering,” arXiv preprint arXiv:2109.04014, 2021
2021 arXiv
-
[59]
Unsupervised image captioning,
Y . Feng, L. Ma, W. Liu, and J. Luo, “Unsupervised image captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 4125–4134
2019
-
[60]
Towards unsupervised image captioning with shared multimodal embeddings,
I. Laina, C. Rupprecht, and N. Navab, “Towards unsupervised image captioning with shared multimodal embeddings,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 7414–7424
2019
-
[62]
Language models can see: Plugging visual controls in text generation,
Y . Su, T. Lan, Y . Liu, F. Liu, D. Yogatama, Y . Wang, L. Kong, and N. Collier, “Language models can see: Plugging visual controls in text generation,” arXiv preprint arXiv:2205.02655 , 2022
2022 arXiv
-
[63]
Conzic: Controllable zero-shot image captioning by sampling-based polishing,
Z. Zeng, H. Zhang, R. Lu, D. Wang, B. Chen, and Z. Wang, “Conzic: Controllable zero-shot image captioning by sampling-based polishing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 23 465–23 476
2023
-
[64]
Mea- cap: Memory-augmented zero-shot image captioning,
Z. Zeng, Y . Xie, H. Zhang, C. Chen, Z. Wang, and B. Chen, “Mea- cap: Memory-augmented zero-shot image captioning,” arXiv preprint arXiv:2403.03715, 2024
2024 arXiv
-
[65]
Text-only training for visual storytelling,
Y . Wang, W. Zhou, Z. Lu, and H. Li, “Text-only training for visual storytelling,” in Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 3686–3695
2023
-
[66]
Zero- shot video captioning with evolving pseudo-tokens,
Y . Tewel, Y . Shalev, R. Nadler, I. Schwartz, and L. Wolf, “Zero- shot video captioning with evolving pseudo-tokens,” arXiv preprint arXiv:2207.11100, 2022
2022 arXiv
-
[67]
Zero-shot dense video captioning by jointly optimizing text and moment,
Y . Jo, S. Lee, A. S. Lee, H. Lee, H. Oh, and M. Seo, “Zero-shot dense video captioning by jointly optimizing text and moment,” arXiv preprint arXiv:2307.02682, 2023
2023 arXiv
-
[68]
Clip models are few-shot learners: Empirical studies on vqa and visual entailment,
H. Song, L. Dong, W.-N. Zhang, T. Liu, and F. Wei, “Clip models are few-shot learners: Empirical studies on vqa and visual entailment,” arXiv preprint arXiv:2203.07190 , 2022
2022 arXiv
-
[69]
Towards counterfactual image manipulation via clip,
Y . Yu, F. Zhan, R. Wu, J. Zhang, S. Lu, M. Cui, X. Xie, X.-S. Hua, and C. Miao, “Towards counterfactual image manipulation via clip,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 3637–3645
2022
-
[70]
An empirical study of gpt-3 for few-shot knowledge-based vqa,
Z. Yang, Z. Gan, J. Wang, X. Hu, Y . Lu, Z. Liu, and L. Wang, “An empirical study of gpt-3 for few-shot knowledge-based vqa,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 3, 2022, pp. 3081–3089
2022
-
[71]
Plug-and- play vqa: Zero-shot vqa by conjoining large pretrained models with zero training,
A. M. H. Tiong, J. Li, B. Li, S. Savarese, and S. C. Hoi, “Plug-and- play vqa: Zero-shot vqa by conjoining large pretrained models with zero training,” arXiv preprint arXiv:2210.08773 , 2022
2022 arXiv
-
[72]
Language models with image descriptors are strong few-shot video-language learners,
Z. Wang, M. Li, R. Xu, L. Zhou, J. Lei, X. Lin, S. Wang, Z. Yang, C. Zhu, D. Hoiem et al. , “Language models with image descriptors are strong few-shot video-language learners,” Advances in Neural Information Processing Systems , vol. 35, pp. 8483–8497, 2022
2022
-
[73]
Language as the medium: Multimodal video classification through text only,
L. Hanu, A. L. Ver ˝o, and J. Thewlis, “Language as the medium: Multimodal video classification through text only,” arXiv preprint arXiv:2309.10783, 2023
2023 arXiv
-
[74]
A video is worth 4096 tokens: Verbalize story videos to understand them in zero shot,
A. Bhattacharya, Y . K. Singla, B. Krishnamurthy, R. R. Shah, and C. Chen, “A video is worth 4096 tokens: Verbalize story videos to understand them in zero shot,” arXiv preprint arXiv:2305.09758, 2023
2023 arXiv
-
[75]
Retrieving-to-answer: Zero-shot video question answering with frozen large language models,
J. Pan, Z. Lin, Y . Ge, X. Zhu, R. Zhang, Y . Wang, Y . Qiao, and H. Li, “Retrieving-to-answer: Zero-shot video question answering with frozen large language models,” arXiv preprint arXiv:2306.11732 , 2023
2023 arXiv
-
[76]
Text-only training for image captioning using noise-injected clip,
D. Nukrai, R. Mokady, and A. Globerson, “Text-only training for image captioning using noise-injected clip,” arXiv preprint arXiv:2211.00575, 2022
2022 arXiv
-
[77]
I can’t believe there’s no images! learning visual tasks using only language data,
S. Gu, C. Clark, and A. Kembhavi, “I can’t believe there’s no images! learning visual tasks using only language data,” arXiv preprint arXiv:2211.09778, 2022
2022 arXiv
-
[78]
Decap: Decoding clip la- tents for zero-shot captioning via text-only training,
W. Li, L. Zhu, L. Wen, and Y . Yang, “Decap: Decoding clip la- tents for zero-shot captioning via text-only training,” arXiv preprint arXiv:2303.03032, 2023
2023 arXiv
-
[79]
From association to generation: Text- only captioning by unsupervised cross-modal mapping,
J. Wang, M. Yan, and Y . Zhang, “From association to generation: Text- only captioning by unsupervised cross-modal mapping,” arXiv preprint arXiv:2304.13273, 2023
2023 arXiv
-
[80]
Zero-shot image captioning by anchor-augmented vision-language space alignment,
J. Wang, Y . Zhang, M. Yan, J. Zhang, and J. Sang, “Zero-shot image captioning by anchor-augmented vision-language space alignment,” arXiv preprint arXiv:2211.07275 , 2022
2022 arXiv
-
[81]
Transferable decoding with visual entities for zero-shot image captioning,
J. Fei, T. Wang, J. Zhang, Z. He, C. Wang, and F. Zheng, “Transferable decoding with visual entities for zero-shot image captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 3136–3146
2023
-
[82]
Language-free training for zero-shot video grounding,
D. Kim, J. Park, J. Lee, S. Park, and K. Sohn, “Language-free training for zero-shot video grounding,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 2539–2548
2023
-
[83]
Clip-gen: Language- free training of a text-to-image generator with clip,
Z. Wang, W. Liu, Q. He, X. Wu, and Z. Yi, “Clip-gen: Language- free training of a text-to-image generator with clip,” arXiv preprint arXiv:2203.00386, 2022
2022 arXiv
-
[85]
Improving cross-modal alignment with synthetic pairs for text-only image captioning,
Z. Liu, J. Liu, and F. Ma, “Improving cross-modal alignment with synthetic pairs for text-only image captioning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 4, 2024, pp. 3864–3872
2024
-
[87]
Towards language-free training for text-to-image generation,
Y . Zhou, R. Zhang, C. Chen, C. Li, C. Tensmeyer, T. Yu, J. Gu, J. Xu, and T. Sun, “Towards language-free training for text-to-image generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 907–17 917
2022
-
[88]
See, think, confirm: Interactive prompting between vision and language models for knowledge-based visual reasoning,
Z. Chen, Q. Zhou, Y . Shen, Y . Hong, H. Zhang, and C. Gan, “See, think, confirm: Interactive prompting between vision and language models for knowledge-based visual reasoning,” arXiv preprint arXiv:2301.05226 , 2023
2023 arXiv
-
[89]
Vicor: Bridging visual un- derstanding and commonsense reasoning with large language models,
K. Zhou, K. Lee, T. Misu, and X. E. Wang, “Vicor: Bridging visual un- derstanding and commonsense reasoning with large language models,” arXiv preprint arXiv:2310.05872 , 2023
2023 arXiv
-
[90]
Domino: A dual-system for multi-step visual language reasoning,
P. Wang, O. Golovneva, A. Aghajanyan, X. Ren, M. Chen, A. Celiky- ilmaz, and M. Fazel-Zarandi, “Domino: A dual-system for multi-step visual language reasoning,” arXiv preprint arXiv:2310.02804 , 2023
2023 arXiv
-
[91]
IdealGPT: Iteratively decomposing vision and language reasoning via large language models,
H. You, R. Sun, Z. Wang, L. Chen, G. Wang, H. Ayyubi, K.-W. Chang, and S.-F. Chang, “IdealGPT: Iteratively decomposing vision and language reasoning via large language models,” in The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [Online]. Availabl...
2023
-
[92]
Good questions help zero-shot image reasoning,
K. Yang, T. Shen, X. Tian, X. Geng, C. Tao, D. Tao, and T. Zhou, “Good questions help zero-shot image reasoning,” arXiv preprint arXiv:2312.01598, 2023
2023 arXiv
-
[93]
The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,
J. Qi, Z. Xu, Y . Shen, M. Liu, D. Jin, Q. Wang, and L. Huang, “The art of SOCRATIC QUESTIONING: Recursive thinking with large language models,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023. [Online]. Available: https://openreview.net/forum...
2023
-
[94]
Filling the image information gap for VQA: Prompting large language models to proactively ask questions,
Z. Wang, C. Chen, P. Li, and Y . Liu, “Filling the image information gap for VQA: Prompting large language models to proactively ask questions,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Associa...
2023
-
[95]
Multimodal multi-hop question answering through a conversation between tools and efficiently finetuned large language models,
H. Rajabzadeh, S. Wang, H. J. Kwon, and B. Liu, “Multimodal multi-hop question answering through a conversation between tools and efficiently finetuned large language models,” arXiv preprint arXiv:2309.08922, 2023
2023 arXiv
-
[96]
Learn to explain: Multimodal reasoning via thought chains for science question answering,
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” Advances in Neural Information Processing Systems , vol. 35, pp. 2507–2521, 2022
2022
-
[97]
Multimodal chain-of-thought reasoning in language models,
Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola, “Multimodal chain-of-thought reasoning in language models,” arXiv preprint arXiv:2302.00923, 2023
2023 arXiv
-
[98]
T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,
L. Wang, Y . Hu, J. He, X. Xu, N. Liu, H. Liu, and H. T. Shen, “T-sciq: Teaching multimodal chain-of-thought reasoning via large language model signals for science question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp...
2024
-
[99]
Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning,
D. Mondal, S. Modi, S. Panda, R. Singh, and G. S. Rao, “Kam-cot: Knowledge augmented multimodal chain-of-thoughts reasoning,” arXiv preprint arXiv:2401.12863, 2024. JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 21
2024 arXiv
-
[100]
Measuring and improving chain-of-thought reasoning in vision-language models,
Y . Chen, K. Sikka, M. Cogswell, H. Ji, and A. Divakaran, “Measuring and improving chain-of-thought reasoning in vision-language models,” arXiv preprint arXiv:2309.04461 , 2023
2023 arXiv
-
[101]
Efficient end-to-end visual document understanding with rationale distillation,
W. Zhu, A. Agarwal, M. Joshi, R. Jia, J. Thomason, and K. Toutanova, “Efficient end-to-end visual document understanding with rationale distillation,” arXiv preprint arXiv:2311.09612 , 2023
2023 arXiv
-
[102]
Compositional chain- of-thought prompting for large multimodal models,
C. Mitra, B. Huang, T. Darrell, and R. Herzig, “Compositional chain- of-thought prompting for large multimodal models,” arXiv preprint arXiv:2311.17076, 2023
2023 arXiv
-
[103]
Cocot: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs,
D. Zhang, J. Yang, H. Lyu, Z. Jin, Y . Yao, M. Chen, and J. Luo, “Cocot: Contrastive chain-of-thought prompting for large multimodal models with multiple image inputs,” arXiv preprint arXiv:2401.02582, 2024
2024 arXiv
-
[104]
Chain of images for intuitively reasoning,
F. Meng, H. Yang, Y . Wang, and M. Zhang, “Chain of images for intuitively reasoning,” arXiv preprint arXiv:2311.09241 , 2023
2023 arXiv
-
[105]
Visual chain of thought: Bridging logical gaps with multimodal infillings,
D. Rose, V . Himakunthala, A. Ouyang, R. He, A. Mei, Y . Lu, M. Saxon, C. Sonar, D. Mirza, and W. Y . Wang, “Visual chain of thought: Bridging logical gaps with multimodal infillings,” arXiv preprint arXiv:2305.02317, 2023
2023 arXiv
-
[106]
Let’s think frame by frame: Evaluating video chain of thought with video infilling and prediction,
V . Himakunthala, A. Ouyang, D. Rose, R. He, A. Mei, Y . Lu, C. Sonar, M. Saxon, and W. Y . Wang, “Let’s think frame by frame: Evaluating video chain of thought with video infilling and prediction,” arXiv preprint arXiv:2305.13903, 2023
2023 arXiv
-
[107]
Zero-shot visual relation detection via composite visual cues from large language models,
L. Li, J. Xiao, G. Chen, J. Shao, Y . Zhuang, and L. Chen, “Zero-shot visual relation detection via composite visual cues from large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[108]
Generating action-conditioned prompts for open- vocabulary video action recognition,
C. Jia, M. Luo, X. Chang, Z. Dang, M. Han, M. Wang, G. Dai, S. Dang, and J. Wang, “Generating action-conditioned prompts for open- vocabulary video action recognition,”arXiv preprint arXiv:2312.02226, 2023
2023 arXiv
-
[109]
Video- prompter: an ensemble of foundational models for zero-shot video understanding,
A. Yousaf, M. Naseer, S. Khan, F. S. Khan, and M. Shah, “Video- prompter: an ensemble of foundational models for zero-shot video understanding,” arXiv preprint arXiv:2310.15324 , 2023
2023 arXiv
-
[111]
Aligning bag of regions for open-vocabulary object detection,
S. Wu, W. Zhang, S. Jin, W. Liu, and C. C. Loy, “Aligning bag of regions for open-vocabulary object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 15 254–15 264
2023
-
[112]
Object-aware distillation pyramid for open-vocabulary object detection,
L. Wang, Y . Liu, P. Du, Z. Ding, Y . Liao, Q. Qi, B. Chen, and S. Liu, “Object-aware distillation pyramid for open-vocabulary object detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 11 186–11 196
2023
-
[113]
Open vocabulary object detection with pseudo bounding-box labels,
M. Gao, C. Xing, J. C. Niebles, J. Li, R. Xu, W. Liu, and C. Xiong, “Open vocabulary object detection with pseudo bounding-box labels,” in European Conference on Computer Vision . Springer, 2022, pp. 266–282
2022
-
[116]
M-tuning: Prompt tuning with mitigated label bias in open-set scenarios,
N. Liao, X. Zhang, M. Cao, J. Yan, and Q. Tian, “M-tuning: Prompt tuning with mitigated label bias in open-set scenarios,” arXiv preprint arXiv:2303.05122, 2023
2023 arXiv
-
[118]
Decouple before interact: Multi-modal prompt learning for continual visual question answering,
Z. Qian, X. Wang, X. Duan, P. Qin, Y . Li, and W. Zhu, “Decouple before interact: Multi-modal prompt learning for continual visual question answering,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2953–2962
2023
-
[119]
Coin: A benchmark of continual instruction tuning for multimodel large language model,
C. Chen, J. Zhu, X. Luo, H. Shen, L. Gao, and J. Song, “Coin: A benchmark of continual instruction tuning for multimodel large language model,” arXiv preprint arXiv:2403.08350 , 2024
2024 arXiv
-
[120]
Beyond anti-forgetting: Multimodal continual instruction tuning with positive forward transfer,
J. Zheng, Q. Ma, Z. Liu, B. Wu, and H. Feng, “Beyond anti-forgetting: Multimodal continual instruction tuning with positive forward transfer,” arXiv preprint arXiv:2401.09181 , 2024
2024 arXiv
-
[121]
Mm-react: Prompting chatgpt for multimodal reasoning and action,
Z. Yang, L. Li, J. Wang, K. Lin, E. Azarnasab, F. Ahmed, Z. Liu, C. Liu, M. Zeng, and L. Wang, “Mm-react: Prompting chatgpt for multimodal reasoning and action,” arXiv preprint arXiv:2303.11381 , 2023
2023 arXiv
-
[122]
Visual chatgpt: Talking, drawing and editing with visual foundation models,
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan, “Visual chatgpt: Talking, drawing and editing with visual foundation models,” arXiv preprint arXiv:2303.04671, 2023
2023 arXiv
-
[123]
Chameleon: Plug-and-play compositional reasoning with large language models,
P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y . N. Wu, S.-C. Zhu, and J. Gao, “Chameleon: Plug-and-play compositional reasoning with large language models,” ArXiv, vol. abs/2304.09842,
-
[124]
Mm-vid: Advancing video understanding with gpt-4v (ision),
K. Lin, F. Ahmed, L. Li, C.-C. Lin, E. Azarnasab, Z. Yang, J. Wang, L. Liang, Z. Liu, Y . Luet al., “Mm-vid: Advancing video understanding with gpt-4v (ision),” arXiv preprint arXiv:2310.19773 , 2023
2023 arXiv
-
[125]
Visual programming: Compositional visual reasoning without training,
T. Gupta and A. Kembhavi, “Visual programming: Compositional visual reasoning without training,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2022
2023
-
[126]
Vipergpt: Visual inference via python execution for reasoning,
D. Sur’is, S. Menon, and C. V ondrick, “Vipergpt: Visual inference via python execution for reasoning,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV) , 2023
2023
-
[127]
Zero-shot video question answering with procedural programs,
R. Choudhury, K. Niinuma, K. M. Kitani, and L. A. Jeni, “Zero-shot video question answering with procedural programs,” ArXiv, vol. abs/2312.00937, 2023. [Online]. Available: https://api.semanticscholar. org/CorpusID:265608945
2023 arXiv
-
[128]
Towards truly zero-shot compositional visual reasoning with llms as programmers,
A. Stani’c, S. Caelles, and M. Tschannen, “Towards truly zero-shot compositional visual reasoning with llms as programmers,” ArXiv, vol. abs/2401.01974, 2024. [Online]. Available: https://api.semanticscholar. org/CorpusID:266755924
2024 arXiv
-
[129]
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,
J. Johnson, B. Hariharan, L. Van Der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick, “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,” in Pro- ceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp...
2017
-
[130]
Microsoft coco: Common objects in context,
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll ´ar, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 . Springer,...
2014
-
[131]
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hocken- maier, and S. Lazebnik, “Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models,” in Proceedings of the IEEE international conference on computer vision , 2015, pp. ...
2015
-
[132]
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,
X. Wang, J. Wu, J. Chen, L. Li, Y .-F. Wang, and W. Y . Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 4581–4591
2019
-
[133]
Msr-vtt: A large video description dataset for bridging video and language,
J. Xu, T. Mei, T. Yao, and Y . Rui, “Msr-vtt: A large video description dataset for bridging video and language,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 5288–5296
2016
-
[134]
A dataset for movie description,
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele, “A dataset for movie description,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3202–3212
2015
-
[135]
Funqa: Towards surprising video comprehension,
B. Xie, S. Zhang, Z. Zhou, B. Li, Y . Zhang, J. Hessel, J. Yang, and Z. Liu, “Funqa: Towards surprising video comprehension,” arXiv preprint arXiv:2306.14899, 2023
2023 arXiv
-
[136]
Visual abductive reason- ing,
C. Liang, W. Wang, T. Zhou, and Y . Yang, “Visual abductive reason- ing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 15 565–15 575
2022
-
[137]
Ask, attend and answer: Exploring question- guided spatial attention for visual question answering,
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question- guided spatial attention for visual question answering,” in Com- puter Vision–ECCV 2016: 14th European Conference, Amsterdam, the Netherlands, October 11–14, 2016, Proceedings, Part VII 14. Springer, 2016, pp. 451–466
2016
-
[138]
Dynamic key-value memory enhanced multi- step graph reasoning for knowledge-based visual question answering,
M. Li and M.-F. Moens, “Dynamic key-value memory enhanced multi- step graph reasoning for knowledge-based visual question answering,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 10, 2022, pp. 10 983–10 992
2022
-
[139]
Explore multi-step reason- ing in video question answering,
X. Song, Y . Shi, X. Chen, and Y . Han, “Explore multi-step reason- ing in video question answering,” in Proceedings of the 26th ACM international conference on Multimedia , 2018, pp. 239–247
2018
-
[140]
Knowledge editing for large language models: A survey,
S. Wang, Y . Zhu, H. Liu, Z. Zheng, C. Chen et al., “Knowledge editing for large language models: A survey,”arXiv preprint arXiv:2310.16218, 2023
2023 arXiv
-
[141]
Noc-rek: novel object captioning with retrieved vocabulary from external knowledge,
D. M. V o, H. Chen, A. Sugimoto, and H. Nakayama, “Noc-rek: novel object captioning with retrieved vocabulary from external knowledge,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022, pp. 17 979–17 987. JOURNAL OF IEEE, VOL. 14, NO. 8...
2022
-
[142]
Straight to the facts: Learning knowledge base retrieval for factual visual question answering,
M. Narasimhan and A. G. Schwing, “Straight to the facts: Learning knowledge base retrieval for factual visual question answering,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 451–468
2018
-
[143]
Ask me anything: Free-form visual question answering based on knowledge from external sources,
Q. Wu, P. Wang, C. Shen, A. Dick, and A. Van Den Hengel, “Ask me anything: Free-form visual question answering based on knowledge from external sources,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4622–4630
2016
-
[144]
Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks,
G. Li, H. Su, and W. Zhu, “Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks,” arXiv preprint arXiv:1712.00733, 2017
2017 arXiv
-
[145]
Fvqa: Fact-based visual question answering,
P. Wang, Q. Wu, C. Shen, A. Dick, and A. Van Den Hengel, “Fvqa: Fact-based visual question answering,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 10, pp. 2413–2427, 2017
2017
-
[146]
Zero-shot visual question answering using knowledge graph,
Z. Chen, J. Chen, Y . Geng, J. Z. Pan, Z. Yuan, and H. Chen, “Zero-shot visual question answering using knowledge graph,” in The Semantic Web–ISWC 2021: 20th International Semantic Web Conference, ISWC 2021, Virtual Event, October 24–28, 2021, Proceedings 20 . Springer, 2021, ...
2021
-
[147]
Wordnet: a lexical database for english,
G. A. Miller, “Wordnet: a lexical database for english,” Communica- tions of the ACM , vol. 38, no. 11, pp. 39–41, 1995
1995
-
[148]
Wikidata: a free collaborative knowl- edgebase,
D. Vrande ˇci´c and M. Kr ¨otzsch, “Wikidata: a free collaborative knowl- edgebase,” Communications of the ACM , vol. 57, no. 10, pp. 78–85, 2014
2014
-
[149]
Unit: Multimodal multitask learning with a unified transformer,
R. Hu and A. Singh, “Unit: Multimodal multitask learning with a unified transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 1439–1449
2021
-
[150]
Multi-task learning of hierarchical vision-language representation,
D.-K. Nguyen and T. Okatani, “Multi-task learning of hierarchical vision-language representation,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2019, pp. 10 492– 10 501
2019
-
[151]
12-in-1: Multi-task vision and language representation learning,
J. Lu, V . Goswami, M. Rohrbach, D. Parikh, and S. Lee, “12-in-1: Multi-task vision and language representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, 2020, pp. 10 437–10 446
2020
-
[152]
Finetuned language models are zero-shot learners,
J. Wei, M. Bosma, V . Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le, “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations , 2021
2021
-
[153]
Instruction tuning for large language models: A survey,
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu et al., “Instruction tuning for large language models: A survey,” arXiv preprint arXiv:2308.10792 , 2023
2023
-
[154]
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
-
[155]
Deep reinforcement learning from human preferences,
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[156]
Align before fuse: Vision and language representation learning with momentum distillation,
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in neural information processing systems, vol. 34, pp. 9694–9705, 2021
2021
-
[157]
Modeling caption diversity in contrastive vision-language pretraining,
S. Lavoie, P. Kirichenko, M. Ibrahim, M. Assran, A. G. Wildon, A. Courville, and N. Ballas, “Modeling caption diversity in contrastive vision-language pretraining,” arXiv preprint arXiv:2405.00740 , 2024
2024 arXiv
-
[158]
Videoclip: Contrastive pre-training for zero-shot video-text understanding,
H. Xu, G. Ghosh, P.-Y . Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer, “Videoclip: Contrastive pre-training for zero-shot video-text understanding,” arXiv preprint arXiv:2109.14084, 2021
2021 arXiv
-
[159]
Actionclip: A new paradigm for video action recognition,
M. Wang, J. Xing, and Y . Liu, “Actionclip: A new paradigm for video action recognition,” arXiv preprint arXiv:2109.08472 , 2021
2021 arXiv
-
[160]
Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,” arXiv preprint arXiv:2301.12597 , 2023
2023 arXiv
-
[161]
High-resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684–10 695
2022
-
[162]
Aligning large multimodal models with factually augmented rlhf,
Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y . Shen, C. Gan, L.-Y . Gui, Y .-X. Wang, Y . Yanget al. , “Aligning large multimodal models with factually augmented rlhf,” arXiv preprint arXiv:2309.14525 , 2023
2023 arXiv
-
[163]
Using human feedback to fine-tune diffusion models without any reward model,
K. Yang, J. Tao, J. Lyu, C. Ge, J. Chen, W. Shen, X. Zhu, and X. Li, “Using human feedback to fine-tune diffusion models without any reward model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 8941–8951
2024
-
[164]
Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,
R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shammaet al., “Visual genome: Connect- ing language and vision using crowdsourced dense image annotations,” International journal of computer vision , vol. 123, pp. 32–73, 2017
2017
-
[165]
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in Proceedings of the 56th Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers) , 2018, pp....
2018
-
[166]
Frozen in time: A joint video and image encoder for end-to-end retrieval,
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 1728–1738
2021
-
[167]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
-
[168]
Videochat: Chat-centric video understanding,
K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P. Luo, Y . Wang, L. Wang, and Y . Qiao, “Videochat: Chat-centric video understanding,”arXiv preprint arXiv:2305.06355, 2023
2023 arXiv
-
[169]
Mimic-it: Multi-modal in-context instruction tuning,
B. Li, Y . Zhang, L. Chen, J. Wang, F. Pu, J. Yang, C. Li, and Z. Liu, “Mimic-it: Multi-modal in-context instruction tuning,” arXiv preprint arXiv:2306.05425, 2023
2023 arXiv
-
[170]
Otter: A multi-modal model with in-context instruction tuning,
B. Li, Y . Zhang, L. Chen, J. Wang, J. Yang, and Z. Liu, “Otter: A multi-modal model with in-context instruction tuning,” arXiv preprint arXiv:2305.03726, 2023
2023 arXiv
-
[171]
Nocaps: Novel object captioning at scale,
H. Agrawal, K. Desai, Y . Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson, “Nocaps: Novel object captioning at scale,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8948–8957
2019
-
[172]
Ok-vqa: A visual question answering benchmark requiring external knowledge,
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, 2019, pp. 3195–3204
2019
-
[173]
Vizwiz grand challenge: Answering visual questions from blind people,
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham, “Vizwiz grand challenge: Answering visual questions from blind people,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 3608–3617
2018
-
[174]
Modeling context in referring expressions,
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11- 14, 2016, Proceedings, Part II 14 . Springer, 2016, pp. 69–85
2016
-
[175]
Seed-bench: Benchmarking multimodal llms with generative comprehension,
B. Li, R. Wang, G. Wang, Y . Ge, Y . Ge, and Y . Shan, “Seed-bench: Benchmarking multimodal llms with generative comprehension,” arXiv preprint arXiv:2307.16125, 2023
2023 arXiv
-
[176]
Mmbench: Is your multi-modal model an all-around player?
Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu et al., “Mmbench: Is your multi-modal model an all-around player?” in European Conference on Computer Vision . Springer, 2025, pp. 216–233
2025
-
[177]
Teaching clip to count to ten,
R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, and T. Dekel, “Teaching clip to count to ten,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3170–3180
2023
-
[178]
Eyes wide shut? exploring the visual shortcomings of multimodal llms,
S. Tong, Z. Liu, Y . Zhai, Y . Ma, Y . LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” arXiv preprint arXiv:2401.06209, 2024
2024 arXiv
-
[179]
Do vision-language models understand compound nouns?
S. Kumar, S. Ghosh, S. Sakshi, U. Tyagi, and D. Manocha, “Do vision-language models understand compound nouns?” arXiv preprint arXiv:2404.00419, 2024
2024 arXiv
-
[180]
cola: A benchmark for compositional text-to-image re- trieval,
A. Ray, F. Radenovic, A. Dubey, B. Plummer, R. Krishna, and K. Saenko, “cola: A benchmark for compositional text-to-image re- trieval,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[181]
Crepe: Can vision-language foundation models reason compositionally?
Z. Ma, J. Hong, M. O. Gul, M. Gandhi, I. Gao, and R. Krishna, “Crepe: Can vision-language foundation models reason compositionally?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 10 910–10 921
2023
-
[182]
When and why vision-language models behave like bags-of-words, and what to do about it?
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou, “When and why vision-language models behave like bags-of-words, and what to do about it?” in The Eleventh International Conference on Learning Representations, 2022
2022
-
[183]
Sugar- crepe: Fixing hackable benchmarks for vision-language compositional- ity,
C.-Y . Hsieh, J. Zhang, Z. Ma, A. Kembhavi, and R. Krishna, “Sugar- crepe: Fixing hackable benchmarks for vision-language compositional- ity,” Advances in neural information processing systems, vol. 36, 2024
2024
-
[184]
Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models,
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y . Yacoob et al. , “Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models,” in Proceedings of the IEEE/CVF Conference on...
2024
-
[185]
Muffin or chihuahua? challenging multimodal large language models with multipanel vqa,
Y . Fan, J. Gu, K. Zhou, Q. Yan, S. Jiang, C.-C. Kuo, Y . Zhao, X. Guan, and X. Wang, “Muffin or chihuahua? challenging multimodal large language models with multipanel vqa,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: ...
2024
-
[186]
Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge,
A. Wang, B. Wu, S. Chen, Z. Chen, H. Guan, W.-N. Lee, L. E. Li, and C. Gan, “Sok-bench: A situated video reasoning benchmark with aligned open-world knowledge,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 384–13 394
2024
-
[187]
What if the tv was off? examining counterfactual reasoning abilities of multi- modal language models,
L. Zhang, X. Zhai, Z. Zhao, Y . Zong, X. Wen, and B. Zhao, “What if the tv was off? examining counterfactual reasoning abilities of multi- modal language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 21 853–21 862
2024
-
[188]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019
2019
-
[189]
Parallel refinements for lexically constrained text generation with bart,
X. He, “Parallel refinements for lexically constrained text generation with bart,” arXiv preprint arXiv:2109.12487 , 2021
2021 arXiv
-
[190]
Exploring the limits of transfer learn- ing with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learn- ing with a unified text-to-text transformer,” The Journal of Machine Learning Research, vol. 21, no. 1, pp. 5485–5551, 2020
2020
-
[191]
A style-based generator architecture for generative adversarial networks,
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4401–4410
2019
-
[192]
Clipcap: Clip prefix for image captioning,
R. Mokady, A. Hertz, and A. H. Bermano, “Clipcap: Clip prefix for image captioning,” arXiv preprint arXiv:2111.09734 , 2021
2021 arXiv
-
[193]
Taming transformers for high-resolution image synthesis,
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 12 873–12 883
2021
-
[194]
Recurrent relational memory network for unsupervised image captioning,
D. Guo, Y . Wang, P. Song, and M. Wang, “Recurrent relational memory network for unsupervised image captioning,” arXiv preprint arXiv:2006.13611, 2020
2006 arXiv
-
[195]
Relational distant supervision for image captioning without image-text pairs,
Y . Qi, W. Zhao, and X. Wu, “Relational distant supervision for image captioning without image-text pairs,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 5, 2024, pp. 4524– 4532
2024
-
[196]
Bleu: a method for automatic evaluation of machine translation,
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
-
[197]
Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72
2005
-
[198]
Cider: Consensus- based image description evaluation,
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus- based image description evaluation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 4566–4575
2015
-
[199]
Spice: Semantic propositional image caption evaluation,
P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, Octo- ber 11-14, 2016, Proceedings, Part V 14. Springer, 2016, pp. 382–398
2016
-
[200]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
-
[201]
A-okvqa: A benchmark for visual question answering using world knowledge,
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in European Conference on Computer Vision . Springer, 2022, pp. 146–162
2022
-
[202]
Lvis: A dataset for large vocabulary instance segmentation,
A. Gupta, P. Dollar, and R. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 5356–5364
2019
-
[204]
Commonsense knowledge prompt- ing for few-shot action recognition in videos,
Y . Shi, X. Wu, H. Lin, and J. Luo, “Commonsense knowledge prompt- ing for few-shot action recognition in videos,” IEEE Transactions on Multimedia, 2024
2024
-
[205]
Imagenet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
-
[206]
The caltech-ucsd birds-200-2011 dataset,
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie, “The caltech-ucsd birds-200-2011 dataset,” 2011
2011
-
[207]
Food-101–mining discriminative components with random forests,
L. Bossard, M. Guillaumin, and L. Van Gool, “Food-101–mining discriminative components with random forests,” in Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, Septem- ber 6-12, 2014, Proceedings, Part VI 13. Springer, 2014, pp. 446–461
2014
-
[208]
Places: A 10 million image database for scene recognition,
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba, “Places: A 10 million image database for scene recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 6, pp. 1452– 1464, 2017
2017
-
[209]
Cats and dogs,
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in 2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012, pp. 3498–3505
2012
-
[210]
De- scribing textures in the wild,
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “De- scribing textures in the wild,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 3606–3613
2014
-
[211]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020
2010 arXiv
-
[212]
Align before fuse: Vision and language representation learning with momentum distillation,
J. Li, R. R. Selvaraju, A. D. Gotmare, S. Joty, C. Xiong, and S. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in NeurIPS, 2021
2021
-
[213]
Zero-shot object detection,
A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran, “Zero-shot object detection,” in Proceedings of the European confer- ence on computer vision (ECCV) , 2018, pp. 384–400
2018
-
[214]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
-
[215]
Evaluating large language models trained on code,
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockmanet al., “Evaluating large language models trained on code,” arXiv preprint arXiv:2107.03374 , 2021
2021 arXiv
-
[216]
Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models,
F. Liu, T. Guan, Z. Li, L. Chen, Y . Yacoob, D. Manocha, and T. Zhou, “Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models,” arXiv preprint arXiv:2310...
-
[217]
Plausible may not be faithful: Probing object hallucination in vision-language pre-training,
W. Dai, Z. Liu, Z. Ji, D. Su, and P. Fung, “Plausible may not be faithful: Probing object hallucination in vision-language pre-training,” arXiv preprint arXiv:2210.07688 , 2022
2022 arXiv
-
[218]
Siren’s song in the ai ocean: a survey on hal- lucination in large language models,
Y . Zhang, Y . Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y . Zhang, Y . Chenet al., “Siren’s song in the ai ocean: a survey on hal- lucination in large language models,” arXiv preprint arXiv:2309.01219, 2023
2023 arXiv
-
[219]
How language model hallucinations can snowball,
M. Zhang, O. Press, W. Merrill, A. Liu, and N. A. Smith, “How language model hallucinations can snowball,” arXiv preprint arXiv:2305.13534, 2023
2023 arXiv
-
[220]
Con- tinual learning for large language models: A survey,
T. Wu, L. Luo, Y .-F. Li, S. Pan, T.-T. Vu, and G. Haffari, “Con- tinual learning for large language models: A survey,” arXiv preprint arXiv:2402.01364, 2024
2024 arXiv
-
[221]
Visrag: Vision-based retrieval-augmented generation on multi-modality documents,
S. Yu, C. Tang, B. Xu, J. Cui, J. Ran, Y . Yan, Z. Liu, S. Wang, X. Han, Z. Liu et al., “Visrag: Vision-based retrieval-augmented generation on multi-modality documents,” arXiv preprint arXiv:2410.10594 , 2024
2024 arXiv
-
[222]
Unirag: Universal retrieval augmentation for multi-modal large language mod- els,
S. Sharifymoghaddam, S. Upadhyay, W. Chen, and J. Lin, “Unirag: Universal retrieval augmentation for multi-modal large language mod- els,” arXiv preprint arXiv:2405.10311 , 2024
2024 arXiv
-
[223]
Wiki-llava: Hierarchical retrieval-augmented gen- eration for multimodal llms,
D. Caffagni, F. Cocchi, N. Moratelli, S. Sarto, M. Cornia, L. Baraldi, and R. Cucchiara, “Wiki-llava: Hierarchical retrieval-augmented gen- eration for multimodal llms,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , 2024, pp. 1818– 1826
2024
-
[224]
Snapntell: Enhanc- ing entity-centric visual question answering with retrieval augmented multimodal llm,
J. Qiu, A. Madotto, Z. Lin, P. A. Crook, Y . E. Xu, X. L. Dong, C. Faloutsos, L. Li, B. Damavandi, and S. Moon, “Snapntell: Enhanc- ing entity-centric visual question answering with retrieval augmented multimodal llm,” arXiv preprint arXiv:2403.04735 , 2024
2024 arXiv
-
[225]
Winoground: Probing vision and language models for visio-linguistic compositionality,
T. Thrush, R. Jiang, M. Bartolo, A. Singh, A. Williams, D. Kiela, and C. Ross, “Winoground: Probing vision and language models for visio-linguistic compositionality,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 5238–5248
2022
-
[226]
When are lemons purple? the concept association bias of vision-language models,
Y . Tang, Y . Yamada, Y . Zhang, and I. Yildirim, “When are lemons purple? the concept association bias of vision-language models,” in JOURNAL OF IEEE, VOL. 14, NO. 8, AUGUST 2021 24 Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, ...
2021
-
[227]
No token left behind: Explainability- aided image classification and generation,
R. Paiss, H. Chefer, and L. Wolf, “No token left behind: Explainability- aided image classification and generation,” in European Conference on Computer Vision. Springer, 2022, pp. 334–350
2022
-
[228]
Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena,
L. Parcalabescu, M. Cafagna, L. Muradjan, A. Frank, I. Calixto, and A. Gatt, “Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volum...
2022
-
[2023]
Available: https://api.semanticscholar.org/CorpusID: 258212542
[Online]. Available: https://api.semanticscholar.org/CorpusID: 258212542
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.