REVIEW 3 major objections 5 minor 14 cited by
Multimodal Autoregressive Pre-training of Large Vision Encoders
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Autoregressive prediction of both image patches and text produces generalist vision encoders that outperform contrastive models such as CLIP and SigLIP.
desk verdict AIMv2 is a clean, well-ablated empirical paper: the method—adding captioning to autoregressive image modeling—clearly works, but the headline claims overstate the evidence because the strongest comparisons are confounded with data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the unified autoregressive pre-training objective over a concatenated sequence of image patches and text tokens. A vision transformer encodes patches under a randomly sampled prefix attention mask, and a causal multimodal decoder predicts the shifted sequence: image patches are regressed with a normalized $\ell^2$ pixel loss (following He et al. [48]) and text tokens with cross-entropy, combined as $L = L_{\text{text}} + \alpha\,L_{\text{img}}$ with $\alpha \approx 0.4$. The prefix attention lets the encoder later switch to bidirectional attention without additional tuning, while the decoder provides dense supervision from every patch and token.
What would settle it
Retrain a CLIP or SigLIP model on exactly the AIMV2 12B mixture, including HQITP and the synthetic captions, with matched compute, and compare frozen-trunk and instruction-tuned benchmarks; if the margins vanish or reverse, the claim that multimodal autoregressive pre-training causes the gains is falsified. A second, cleaner test is to remove the proprietary HQITP and synthetic caption subsets from AIMV2's training mix and check whether its advantage persists.
Extended reading notes
Core claim
The paper's central claim is that multimodal autoregressive modeling—factorizing the joint sequence of image patches and caption text as $P(S) = \prod_j P(S_j \mid S_{<j})$ and training with a pixel MSE loss plus a text cross-entropy loss—is an effective objective for pre-training large vision encoders. With this objective, AIMV2-3B reaches 89.5% ImageNet-1k top-1 accuracy under attentive probing with a frozen trunk, and AIMV2 encoders outperform CLIP, SigLIP, and DINOv2 on most multimodal instruction-tuning benchmarks while remaining competitive on recognition, detection, and referring-expression comprehension. The paper further claims that the image-level objective adds signal beyond captioning alone, that the method scales consistently with data and parameters, and that it achieves these results while seeing fewer training samples than the contrastive baselines.
Load-bearing premise
The load-bearing premise is that AIMV2's advantage over CLIP and SigLIP comes from its training objective rather than from its particular 12B image-text mixture, which includes a proprietary high-quality set and synthetic captions, because the headline cross-model comparisons are not matched on data.
Editorial extensions
If this is right
- If correct, generative multimodal autoregression is a viable drop-in pre-training objective for generalist vision encoders, reducing the need for the large batch sizes and careful data filtering that contrastive methods require.
- AIMV2-3B's 89.5% ImageNet-1k accuracy with a frozen trunk implies that representation quality comparable to the best discriminative models can come from image-plus-text next-token prediction.
- The reported scaling behavior (performance improves with model size and sample count, while the optimal size grows with compute) suggests the recipe will keep improving as models and data grow, in line with LLM-style scaling.
- Consistent gains on text-rich benchmarks such as TextVQA, DocVQA, and ChartQA indicate that the multimodal objective is especially useful when downstream tasks require fine-grained reading and localization.
Reading between the lines
- What the paper leaves open is whether the margin over CLIP and SigLIP is an objective effect or a data effect; the inference that the objective alone drives the gains would be confirmed by swapping in matched data for the baselines.
- The dense patch-level supervision suggests an untested prediction: AIMV2 should degrade less than captioning-only models on tasks needing fine spatial detail, such as small-object detection, which the paper's detection results roughly support but do not isolate.
- One could extend the recipe to video or audio by treating frame or spectrogram patches as additional sequence tokens in the same factorization; the paper does not report such experiments.
- The prefix attention trick implies that the encoder is trained to produce useful representations from partial images, which may explain the robustness to cropping and tiling seen in the high-resolution evaluations.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AIMV2, a family of ViT-based vision encoders pre-trained with a multimodal autoregressive objective. The vision encoder uses prefix attention, and a causal multimodal decoder predicts both raw image patches (ℓ2 loss) and caption tokens (cross-entropy) on a 12B image-text mixture (DFN, COYO, proprietary HQITP, and synthetic captions). The authors evaluate the resulting encoders on frozen-trunk recognition, open-vocabulary detection and grounding, multimodal instruction tuning, in-context learning, zero-shot LiT, and native-resolution adaptation, reporting strong results including 89.5% ImageNet-1k top-1 with a frozen trunk. The central methodological claim is that joint autoregressive prediction of patches and text yields a simple, scalable, and generalist vision encoder that matches or outperforms contrastive pre-training.
Significance. If the central claim holds, this is a significant result: it demonstrates that a straightforward autoregressive multimodal objective can compete with contrastive objectives for vision-encoder pre-training, with the added benefits of dense supervision, modest batch sizes, and natural compatibility with LLM-based multimodal pipelines. The paper's strengths include a carefully controlled ablation in Table 9b comparing AIMV2 against CLIP and CapPa under identical architecture and data, a scaling analysis in Figure 2 that mirrors Hoffmann-style compute-optimal behavior, and a public code release. These elements make the core method claim credible. However, the broader advertised claim that AIMV2 'consistently outperforms state-of-the-art contrastive models' is only partially supported: the headline comparisons in Tables 3 and 7 mix method and data differences, and the paper's own results in Table 5 and Appendix D.3 show cases where baselines outperform AIMV2. The contribution is valuable but the presentation needs to separate method advantage from data advantage.
major comments (3)
- [§2.3, Table 2 vs. Tables 3 and 7; §5, Table 9b] The headline gains over SigLIP and CLIP are confounded with pre-training data. AIMV2 is trained on a 12B mixture that includes 3.8B synthetic DFN captions, 431.5M synthetic HQITP captions, and 564.6M proprietary HQITP alt-text pairs (Table 2), while the SigLIP and CLIP checkpoints used in Tables 3 and 7 were trained on different private corpora. The controlled same-data comparison in Table 9b covers only CLIP and CapPa at 2B pairs, not the deployed SigLIP models. A contrastive model trained on the same caption quality could plausibly close a substantial portion of the reported margins, since data filtering and synthetic captions are known to improve contrastive models as well. This does not falsify the method claim, but it means the abstract's 'consistently outperforms state-of-the-art contrastive models' is not fully supported by the evidence as presented. I recommend either adding a same-data SigLIP-style baseline or explicitly qualifying the claim as holding for AIMV2's data mixture.
- [Abstract and Conclusion vs. Table 5 and Table D3] The claim of 'consistently outperforms' is contradicted by results within the paper itself. Table 5 shows SigLIP ViT-So400m at 80.4 zero-shot ImageNet top-1 versus 77.0 for AIMV2-3B, and Table D3 shows DINOv2 outperforming AIMV2 on COCO detection/segmentation (55.5 vs. 54.0 AP). The text in Section 4.1 and Appendix D.2 acknowledges these cases, but the abstract and conclusion do not carry the same qualification. Please temper the claims to 'outperforms or matches' with explicit exceptions, or restrict the claim to the specific settings where the controlled evidence supports it.
- [§5, Table 9c–9f] Several design choices are recommended on the basis of differences that are likely within training noise. For example, Table 9c reports TextVQA 37.5 for α=0.4 versus 37.4 for α=0.2 and 0.6; Table 9e shows decoder width 512 at 35.9 versus 1536 at 36.9; and Table 9f shows depth 12 versus 16 at 37.5 versus 36.6. None of these experiments report multiple seeds or error bars. I am not asking for a full seed study, but the text should avoid presenting these differences as conclusive evidence for a particular hyperparameter or architecture choice, or the authors should add at least a few repeated runs for the key ablations.
minor comments (5)
- [§2.3 and Table 2] The term 'synthetic' captions is used without specifying the captioner model or its filtering procedure, despite citing Lai et al. [63]. A sentence describing the captioning pipeline and any quality filtering would help reproducibility.
- [§4.3.2, Table 8] The in-context learning comparison reports only results for OAI CLIP and DFN-CLIP as quoted from McKinzie et al. [85], without the MM1 ViT-L baseline under identical pre-training data. At minimum, clarify whether the ICL comparison holds the instruction-tuning data fixed.
- [Throughout] There are numerous typos and grammatical slips: 'factorizatized' (§2.1), 'task' for 'tasks' (§2.4), 'hyperaparmeters' in Tables A1, A2, C1, and 'the model’s predicted patch ˆxi(θ)' with mismatched parentheses in §2.1. A careful proofreading pass is needed.
- [§5, Table 9b caption] The caption reads 'AIMV2 vs. CLIP..' but the table also includes CapPa; the caption should list all three methods. Also, the CapPa row is trained at batch size 8k only, while CLIP is given at 8k and 16k; note in the text why CapPa was not run at 16k.
- [§4.1, Table 3] The comparison of AIMV2-3B at 448px against baselines at 224px is apples-to-oranges. The table caption notes the resolution, but the text should explicitly state that the 89.5% result uses a higher-resolution fine-tuned model, not the base pre-training resolution.
Circularity Check
No significant circularity: the central claims rest on external downstream benchmarks and a controlled same-data ablation, not on a derivation that reduces to its own inputs.
full rationale
AIMV2 is an empirical pre-training paper; it does not claim a mathematical derivation of benchmark performance from first principles. The main claims are that a vision encoder trained with a multimodal autoregressive objective (image-patch regression plus caption cross-entropy) transfers well to recognition, grounding, and multimodal instruction tuning. These claims are supported by evaluations on external benchmarks (ImageNet-1k, VQAv2, TextVQA, COCO, LVIS, etc.) that are not part of the training objective. The load-bearing controlled comparison is Table 9b, where AIMV2, CLIP, and CapPa are trained with identical architectures, data, and hyperparameters ("All models are trained using identical architectures, incorporating SwiGLU and RMSNorm, and are pre-trained using the same dataset of image-text pairs"), and AIMV2 wins by 11-13 points on TextVQA. That ablation is the evidence that the objective itself, not the data, drives the gains over contrastive and captioning baselines. Self-citations do appear: [33] (AIM, the authors' own prior work) is cited for prefix attention and as a baseline, [35] (DFN) for data filtering with overlapping authors, and [63] (Lai et al.) for synthetic captions. None of these is used as an unverified premise that forces the paper's conclusion; prefix attention is an architectural choice that is ablated (Table 9a), DFN is a public dataset, and the synthetic-caption pipeline is a data ingredient, not a proof step. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via citation that is then presented as externally validated. The headline comparisons against SigLIP and OAI CLIP are not fully controlled because AIMV2 uses 12B pairs including proprietary HQITP and roughly 4.2B synthetic captions while the baselines used different private corpora, and the abstract's "consistently outperforms" claim is qualified by the zero-shot LiT result in Table 5 where SigLIP leads by 3.4 points. These are legitimate correctness and external-validity concerns, but they are not circularity: the paper's central method claim survives the controlled ablation and is evaluated against external benchmarks. Accordingly, no circular step can be exhibited, and the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- alpha (pixel loss weight) =
0.4
- multimodal decoder width =
1024
- multimodal decoder depth =
12
assumptions (3)
- domain assumption Web-scale image-text pairs, including synthetic captions, provide a sufficient signal for visual representation learning.
- domain assumption Pixel-level MSE regression on normalized patches transfers to semantic downstream tasks.
- domain assumption Downstream benchmark results generalize to real-world performance and are not inflated by test leakage from pre-training data.
Cite this review
Pith. "Pith review of Multimodal Autoregressive Pre-training of Large Vision Encoders." pith.science (2026). https://pith.science/paper/SPRQZIL2
@misc{pith2026241114402,
author = {Pith},
title = {Pith review of: Multimodal Autoregressive Pre-training of Large Vision Encoders},
year = {2026},
howpublished = {\url{https://pith.science/paper/SPRQZIL2}},
note = {Machine review of arXiv:2411.14402}
}
read the original abstract
We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 14 Pith papers
-
CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.
-
Vision-Language Models Do Not Understand Negation
Current vision-language models largely ignore negation, and a new 79k-example benchmark plus synthetic fine-tuning data yields measurable but incomplete improvements.
-
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
A gradient alignment score against a reference model's weights selects and reweights training samples, improving data efficiency in image classification and CLIP pretraining.
-
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
Concept shift vectors, computed as mean activation differences, can partially recover fine-tuned multimodal LLM concepts and steer model outputs without additional training.
-
SigLIP-HD by Fine-to-Coarse Supervision
Fine-to-coarse L1 supervision lets a standard-resolution SigLIP 2 encoder produce better visual tokens for MLLMs without higher-resolution inference.
-
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Codec-guided sparse patch selection plus a lightweight speak/silent gate yields a 4B streaming VLM that is competitive on static tasks, stronger on video/spatial benchmarks, and much cheaper at inference.
-
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Inserting a pixel-shuffle plus residual patch-merge layer inside the vision encoder compresses visual tokens more efficiently than post-encoder compression, at modest accuracy cost.
-
Hierarchical Pre-Training of Vision Encoders with Large Language Model
A three-stage pre-training scheme that feeds multi-layer vision features into an LLM reports marginal benchmark gains, but lacks data, code, and ablations needed to support the claim.
-
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.
-
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.
-
Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models
Fine-tuning recent natural-domain foundation models, especially AIMv2, improves medical image classification accuracy across mammography, skin lesion, retinopathy, and chest X-ray benchmarks.
-
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Llama-SMoP-DEDR, a sparse mixture of projectors with modality-specific experts and routers, lowers word error rate for LLM-based AVSR on LRS3, mainly with smaller LLMs.
-
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
Adding an L2 loss that pushes the language model's image hidden states back toward the input image embeddings improves LLaVA-style models on several VQA benchmarks, with some benchmarks unaffected or slightly worse.
-
Visual RAG: Expanding MLLM visual knowledge without fine-tuning
Retrieval-selected demonstration examples let a multimodal LLM classify images as accurately as random many-shot prompting with far fewer examples.
Reference graph
Works this paper leans on
-
[1]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1
arXiv 2023
-
[2]
Nocaps: Novel object cap- tioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. Nocaps: Novel object cap- tioning at scale. In ICCV, 2019. 7, 16
2019
-
[3]
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, An- toine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Se- bastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Bink...
2022
-
[4]
Self-supervised learning from im- ages with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bo- janowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from im- ages with a joint-embedding predictive architecture. arXiv preprint arXiv:2301.08243, 2023. 9
arXiv 2023
-
[5]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 1
arXiv 2023
-
[6]
Qwen-vl: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966 ,
-
[7]
From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge
Peter Bandi, Oscar Geessink, Quirine Manson, Mar- cory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, et al. From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge. IEEE Transactions on Medical Imaging, 2018. 15
2018
-
[8]
BEiT: Bert pre- training of image transformers
Hangbo Bao, Li Dong, and Furu Wei. BEiT: Bert pre- training of image transformers. In ICLR, 2022. 1, 9
2022
Show all 137 references
-
[9]
Flexivit: One model for all patch sizes
Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, and Filip Pavetic. Flexivit: One model for all patch sizes. In CVPR, 2023. 4
2023
-
[10]
Food-101 – mining discriminative components with ran- dom forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with ran- dom forests. In ECCV, 2014. 15
2014
-
[11]
Time series analysis: forecasting and control
George EP Box, Gwilym M Jenkins, Gregory C Reinsel, and Greta M Ljung. Time series analysis: forecasting and control. John Wiley & Sons, 2015. 8
2015
-
[12]
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Nee- lakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. preprint arXiv:2005.14165, 2020. 8
2005 arXiv
-
[13]
Coyo-700m: Image-text pair dataset, 2022
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset, 2022. 3, 9
2022
-
[14]
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nico- las Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV,
-
[15]
Unsupervised learn- ing of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learn- ing of visual features by contrasting cluster assignments. In NeurIPS, 2020. 9
2020
-
[16]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021. 9
2021
-
[17]
A generative approach for wikipedia-scale visual entity recognition
Mathilde Caron, Ahmet Iscen, Alireza Fathi, and Cordelia Schmid. A generative approach for wikipedia-scale visual entity recognition. In CVPR, 2024. 9
2024
-
[18]
Mmdetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang,...
2019
-
[19]
Generative pre- training from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Hee- woo Jun, David Luan, and Ilya Sutskever. Generative pre- training from pixels. In ICML, 2020. 9
2020
-
[20]
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 9
2020
-
[21]
Microsoft coco captions: Data collection and eval- uation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and eval- uation server. arXiv preprint arXiv:1504.00325, 2015. 7
2015 arXiv
-
[22]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhang- hao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 7
2023
-
[23]
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arxiv 2022. arXiv preprint arXiv:2204.02311 ,
2022 arXiv
-
[24]
Functional map of the world
Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. 15
2018
-
[25]
Cimpoi, S
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi. Describing textures in the wild. In CVPR, 2014. 15
2014
-
[26]
Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution
Mostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek, Matthias Minderer, Mathilde Caron, An- dreas Steiner, Joan Puigcerver, Robert Geirhos, Ibrahim M 10 Alabdulmohsin, et al. Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution. Advances i...
2024
-
[27]
Imagenet: A large-scale hierarchical im- age database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical im- age database. In 2009 IEEE conference on computer vision and pattern recognition, 2009. 15
2009
-
[28]
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021. 9
2021
-
[29]
Unsu- pervised visual representation learning by context predic- tion
Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsu- pervised visual representation learning by context predic- tion. In ICCV, 2015. 1
2015
-
[30]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. In ICLR, 2021. 2
2021
-
[32]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783,
-
[33]
Scalable pre-training of large autoregressive image models
Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M Susskind, and Armand Joulin. Scalable pre-training of large autoregressive image models. arXiv preprint arXiv:2401.08541, 2024. 1, 2, 3, 5, 8, 9
2024 arXiv
-
[34]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. InCVPR,
-
[35]
Data filtering networks
Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. Data filtering networks. arXiv preprint arXiv:2309.17425, 2023. 1, 3, 5, 9
2023 arXiv
-
[36]
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. 2019 ieee. In ICCV, 2018. 1
2019
-
[37]
Improved base- lines for vision-language pre-training
Enrico Fini, Pietro Astolfi, Adriana Romero-Soriano, Jakob Verbeek, and Michal Drozdzal. Improved base- lines for vision-language pre-training. arXiv preprint arXiv:2305.08675, 2023. 9
2023 arXiv
-
[38]
Mme: A comprehensive eval- uation benchmark for multimodal large language models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. Mme: A comprehensive eval- uation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023. 16
2023 arXiv
-
[39]
Un- supervised representation learning by predicting image ro- tations
Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Un- supervised representation learning by predicting image ro- tations. arXiv preprint arXiv:1803.07728, 2018. 9
2018 arXiv
-
[40]
Making the v in vqa matter: El- evating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: El- evating the role of image understanding in visual question answering. In CVPR, 2017. 7, 8
2017
-
[41]
Making the v in vqa matter: El- evating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: El- evating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, 2017. 16
2017
-
[42]
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. NeurIPS, 2020. 9
2020
-
[43]
Lvis: A dataset for large vocabulary instance segmentation, 2019
Agrim Gupta, Piotr Doll ´ar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation, 2019. 6
2019
-
[44]
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In CVPR, 2018. 7
2018
-
[45]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR,
-
[46]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. In ICCV, 2017. 1
2017
-
[47]
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. 2018. 18
2018
-
[48]
Masked autoencoders are scal- able vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scal- able vision learners. In CVPR, 2022. 1, 2, 5, 9
2022
-
[49]
Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,
Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,
-
[50]
Training compute-optimal large language mod- els
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language mod- els. arXiv preprint arXiv:2203.15556, 2022. 1, 3, 4
2022 arXiv
-
[51]
Scaling up vision-language pre-training for image captioning
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. Scaling up vision-language pre-training for image captioning. In CVPR, 2022. 9
2022
-
[52]
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , 2019. 6, 16
2019
-
[53]
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019. 7
2019
-
[54]
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 2021. 1, 9 11
2021
-
[55]
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361,
2001 arXiv
-
[56]
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR,
-
[57]
Referitgame: Referring to objects in pho- tographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in pho- tographs of natural scenes. In Proceedings of the 2014 con- ference on empirical methods in natural language process- ing (EMNLP), 2014. 6
2014
-
[58]
Big transfer (bit): General visual representation learning
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In ECCV, 2020. 9
2020
-
[59]
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), 2013. 15
2013
-
[60]
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 15
2009
-
[61]
Imagenet classification with deep convolutional neural net- works
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural net- works. NeurIPS, 2012. 9
2012
-
[62]
Mammut: A simple architec- ture for joint learning for multimodal tasks
Weicheng Kuo, AJ Piergiovanni, Dahun Kim, Xiyang Luo, Ben Caine, Wei Li, Abhijit Ogale, Luowei Zhou, Andrew Dai, Zhifeng Chen, et al. Mammut: A simple architec- ture for joint learning for multimodal tasks. arXiv preprint arXiv:2303.16839, 2023. 9
2023 arXiv
-
[63]
Revisit large-scale image-caption data in pre-training multimodal foundation models
Zhengfeng Lai, Vasileios Saveris, Chen Chen, Hong-You Chen, Haotian Zhang, Bowen Zhang, Juan Lao Tebar, Wenze Hu, Zhe Gan, Peter Grasch, et al. Revisit large-scale image-caption data in pre-training multimodal foundation models. arXiv preprint arXiv:2410.02740, 2024. 3
-
[64]
Seed-bench: Benchmarking multi- modal llms with generative comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yix- iao Ge, and Ying Shan. Seed-bench: Benchmarking multi- modal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. 16
2023 arXiv
-
[65]
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. NeurIPS, 2021. 9
2021
-
[66]
Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation. In ICML, 2022. 9
2022
-
[67]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023. 9
2023
-
[68]
Exploring plain vision transformer backbones for object de- tection
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object de- tection. In European conference on computer vision, 2022. 6, 17, 18
2022
-
[69]
Exploring plain vision transformer backbones for object de- tection, 2022
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object de- tection, 2022. 6
2022
-
[70]
Mini-gemini: Mining the potential of multi-modality vision language models
Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814,
-
[71]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th Eu- ropean Conference, Zurich, Switzerland, September 6-12, 2014, Proceed...
2014
-
[72]
Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models
Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023. 6, 7
2023 arXiv
-
[73]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In CVPR, 2024. 1, 6, 7, 9, 16, 17
2024
-
[74]
Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection. 2024. 6
2024
-
[75]
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICLR, 2017. 15
2017
-
[76]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 15, 16
2017 arXiv
-
[77]
Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action
Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Anirud- dha Kembhavi. Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action. In CVPR, 2024. 9
2024
-
[78]
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Ad- vances in Neural Information Processing Systems, 2022. 16
2022
-
[79]
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016. 6
2016
-
[80]
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In CVPR, 2019. 7
2019
-
[81]
Ok-vqa: A visual question answering benchmark requiring external knowledge
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. InConference on Computer Vision and Pattern Recognition (CVPR) , 2019. 16
2019
-
[82]
Chartqa: A benchmark for question 12 answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question 12 answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244, 2022. 16
2022 arXiv
-
[83]
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Pro- ceedings of the IEEE/CVF winter conference on applica- tions of computer vision, 2021. 16
2021
-
[84]
Infographicvqa
Minesh Mathew, Viraj Bagal, Rub `en Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision, 2022. 16
2022
-
[85]
Mm1: Meth- ods, analysis & insights from multimodal llm pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al. Mm1: Meth- ods, analysis & insights from multimodal llm pre-training. arXiv preprint arXiv:2403.09611, 2024. 1, 6, 7, 9
2024 arXiv
-
[86]
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV,
-
[87]
Maxime Oquab, Timoth ´ee Darcet, Theo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernan- dez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Ass- ran, N...
2023
-
[88]
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V . Jawahar. Cats and dogs. In CVPR, 2012. 15
2012
-
[89]
Moment matching for multi-source domain adaptation
Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In ICCV, 2019. 15
2019
-
[90]
Plummer, Liwei Wang, Christopher M
Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazeb- nik. Flickr30k entities: Collecting region-to-phrase cor- respondences for richer image-to-sentence models. IJCV,
-
[91]
Dataset decomposition: Faster llm training with variable sequence length curriculum
Hadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pa- van Kumar Anasosalu Vasu, Cem Koc, Vaishaal Shankar, and Oncel Tuzel. Dataset decomposition: Faster llm training with variable sequence length curriculum. arXiv preprint arXiv:2405.13226, 2024. 4
2024 arXiv
-
[92]
Improving language understanding by genera- tive pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by genera- tive pre-training. 2018. 1, 8
2018
-
[93]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 2019. 1, 8
2019
-
[94]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 1, 5, 6, 8, 9, 16
2021
-
[95]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Ma- chine Learning Research, 21(1), 2020. 2, 3
2020
-
[96]
Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrit- twieser, et al. Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context. arXiv...
2024 arXiv
-
[97]
Imagenet-21k pretraining for the masses
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. arXiv preprint arXiv:2104.10972, 2021. 9
2021 arXiv
-
[98]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 9
2022
-
[99]
Learning visual representations with caption annotations
Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus. Learning visual representations with caption annotations. In Computer Vision–ECCV 2020: 16th European Confer- ence, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16. Springer, 2020. 9
2020
-
[100]
Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs. In NeurIPS Workshop, 2021. 9
2021
-
[101]
Objects365: A large-scale, high-quality dataset for object detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019. 6
2019
-
[102]
Glu variants improve transformer
Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020. 3
2002 arXiv
-
[103]
When do we not need larger vision mod- els? arXiv preprint arXiv:2403.13043, 2024
Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, and Trevor Darrell. When do we not need larger vision mod- els? arXiv preprint arXiv:2403.13043, 2024. 7
2024 arXiv
-
[104]
Unival: Unified model for image, video, audio and language tasks
Mustafa Shukor, Corentin Dancette, Alexandre Rame, and Matthieu Cord. Unival: Unified model for image, video, audio and language tasks. Transactions on Machine Learn- ing Research Journal, 2023. 9
2023
-
[105]
Textcaps: a dataset for image caption- ing with reading comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. Textcaps: a dataset for image caption- ing with reading comprehension. In ECCV, 2020. 7, 16
2020
-
[106]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, 2019. 8, 16
2019
-
[107]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In CVPR, 2019. 7
2019
-
[108]
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhi- nav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 9
2017
-
[109]
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 9
2023 arXiv
-
[110]
Generative pretraining in mul- timodality
Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in mul- timodality. arXiv preprint arXiv:2307.05222, 2023. 9 13
2023 arXiv
-
[111]
Generative multimodal models are in-context learners
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiy- ing Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative multimodal models are in-context learners. In CVPR, 2024. 9
2024
-
[112]
Taylor, B
J. Taylor, B. Earnshaw, B. Mabey, M. Victors, and J. Yosin- ski. Rxrx1: An image set for cellular morphological vari- ation across many experimental batches. In ICLR, 2019. 15
2019
-
[113]
Chameleon: Mixed-modal early-fusion foundation models
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818 ,
-
[114]
Particu- lar object retrieval with integral max-pooling of cnn activa- tions
Giorgos Tolias, Ronan Sicre, and Herv ´e J ´egou. Particu- lar object retrieval with integral max-pooling of cnn activa- tions. arXiv preprint arXiv:1511.05879, 2015. 1
2015 arXiv
-
[115]
Cambrian- 1: A fully open, vision-centric exploration of multimodal llms
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian- 1: A fully open, vision-centric exploration of multimodal llms. arXiv:2406.16860, 2024. 1, 6, 7, 16, 17
2024 arXiv
-
[116]
Llama: Open and efficient foundation language mod- els
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Bap- tiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language mod- els. arXiv preprint arXiv:2302.13971, 2023. 1, 3, 8
2023 arXiv
-
[117]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 3, 8
2023 arXiv
-
[118]
Image captioners are scalable vision learners too
Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiao- hua Zhai, Neil Houlsby, and Lucas Beyer. Image captioners are scalable vision learners too. NeurIPS, 2024. 1, 5, 6, 8, 9
2024
-
[119]
The inaturalist species classification and detection dataset
Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In CVPR, 2018. 15
2018
-
[120]
Rotation equivariant cnns for digital pathology
Bastiaan S Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling. Rotation equivariant cnns for digital pathology. In Medical Image Computing and Com- puter Assisted Intervention, 2018. 15
2018
-
[121]
Show and tell: A neural image caption gen- erator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Du- mitru Erhan. Show and tell: A neural image caption gen- erator. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015. 9
2015
-
[122]
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International conference on machine learn- i...
2022
-
[123]
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904, 2021. 9
2021 arXiv
-
[124]
Vila-u: a unified foundation model inte- grating visual understanding and generation
Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, et al. Vila-u: a unified foundation model inte- grating visual understanding and generation. arXiv preprint arXiv:2409.04429, 2024. 9
2024 arXiv
-
[125]
Show-o: One single transformer to unify multimodal under- standing and generation
Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal under- standing and generation. arXiv preprint arXiv:2408.12528,
-
[126]
Show, attend and tell: Neural image cap- tion generation with visual attention
Kelvin Xu. Show, attend and tell: Neural image cap- tion generation with visual attention. arXiv preprint arXiv:1502.03044, 2015. 9
2015 arXiv
-
[127]
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hock- enmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2, 2014. 6
2014
-
[128]
Vector-quantized image modeling with improved vqgan
Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. ArXiv, 2021. 9
2021
-
[129]
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mo- jtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. TMLR, 2022. 5, 9
2022
-
[130]
Berg, and Tamara L
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. Modeling context in referring expressions, 2016. 6
2016
-
[131]
Scaling autore- gressive multi-modal models: Pretraining and instruction tuning
Lili Yu, Bowen Shi, Ramakanth Pasunuru, Benjamin Muller, Olga Golovneva, Tianlu Wang, Arun Babu, Binh Tang, Brian Karrer, Shelly Sheynin, et al. Scaling autore- gressive multi-modal models: Pretraining and instruction tuning. arXiv preprint arXiv:2309.02591, 2023. 9
2023 arXiv
-
[132]
Lit: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning. In CVPR, 2022. 2, 6
2022
-
[133]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, 2023. 1, 3, 5, 9, 15, 16
2023
-
[134]
Root mean square layer normalization
Biao Zhang and Rico Sennrich. Root mean square layer normalization. NeurIPS, 2019. 3
2019
-
[135]
Colorful image colorization
Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In ECCV, 2016. 9
2016
-
[136]
An open and comprehensive pipeline for unified object grounding and detection, 2024
Xiangyu Zhao, Yicheng Chen, Shilin Xu, Xiangtai Li, Xin- jiang Wang, Yining Li, and Haian Huang. An open and comprehensive pipeline for unified object grounding and detection, 2024. 6
2024
-
[137]
ibot: Image bert pre- training with online tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre- training with online tokenizer. In ICLR, 2022. 1, 9
2022
-
[138]
supreme gasoline
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. In ICLR, 2024. 6 14 A. Hyperparamters Pre-training. We outline the optimization hyperaparmeters and data augmentations...
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.