Pith. sign in

REVIEW 3 major objections 5 minor 14 cited by

Multimodal Autoregressive Pre-training of Large Vision Encoders

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Autoregressive prediction of both image patches and text produces generalist vision encoders that outperform contrastive models such as CLIP and SigLIP.

desk verdict AIMv2 is a clean, well-ablated empirical paper: the method—adding captioning to autoregressive image modeling—clearly works, but the headline claims overstate the evidence because the strongest comparisons are confounded with data. read the letter →

arxiv 2411.14402 v1 pith:SPRQZIL2 submitted 2024-11-21 cs.CV cs.LG

classification cs.CVcs.LG
keywords visionencoderpre-trainingautoregressivemodelingmultimodallearningimage-texttransformergenerativerepresentationscalinglaws
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that generative autoregressive pre-training, applied jointly to image patches and text tokens, can produce general-purpose vision encoders that rival and often beat contrastive models such as CLIP and SigLIP. The proposed AIMV2 family couples a prefix-attention vision transformer with a causal multimodal decoder that predicts the next image patch and the next text token, drawing a training signal from every input. The authors argue that this simple recipe scales like large language models, is easy to implement without huge batch sizes, and yields strong frozen-trunk performance across recognition, detection, grounding, and multimodal question answering. If true, it would offer a simpler alternative to contrastive pre-training for building generalist vision backbones.

What carries the argument

The load-bearing mechanism is the unified autoregressive pre-training objective over a concatenated sequence of image patches and text tokens. A vision transformer encodes patches under a randomly sampled prefix attention mask, and a causal multimodal decoder predicts the shifted sequence: image patches are regressed with a normalized $\ell^2$ pixel loss (following He et al. [48]) and text tokens with cross-entropy, combined as $L = L_{\text{text}} + \alpha\,L_{\text{img}}$ with $\alpha \approx 0.4$. The prefix attention lets the encoder later switch to bidirectional attention without additional tuning, while the decoder provides dense supervision from every patch and token.

What would settle it

Retrain a CLIP or SigLIP model on exactly the AIMV2 12B mixture, including HQITP and the synthetic captions, with matched compute, and compare frozen-trunk and instruction-tuned benchmarks; if the margins vanish or reverse, the claim that multimodal autoregressive pre-training causes the gains is falsified. A second, cleaner test is to remove the proprietary HQITP and synthetic caption subsets from AIMV2's training mix and check whether its advantage persists.

Watch

Extended reading notes

Core claim

The paper's central claim is that multimodal autoregressive modeling—factorizing the joint sequence of image patches and caption text as $P(S) = \prod_j P(S_j \mid S_{<j})$ and training with a pixel MSE loss plus a text cross-entropy loss—is an effective objective for pre-training large vision encoders. With this objective, AIMV2-3B reaches 89.5% ImageNet-1k top-1 accuracy under attentive probing with a frozen trunk, and AIMV2 encoders outperform CLIP, SigLIP, and DINOv2 on most multimodal instruction-tuning benchmarks while remaining competitive on recognition, detection, and referring-expression comprehension. The paper further claims that the image-level objective adds signal beyond captioning alone, that the method scales consistently with data and parameters, and that it achieves these results while seeing fewer training samples than the contrastive baselines.

Load-bearing premise

The load-bearing premise is that AIMV2's advantage over CLIP and SigLIP comes from its training objective rather than from its particular 12B image-text mixture, which includes a proprietary high-quality set and synthetic captions, because the headline cross-model comparisons are not matched on data.

Editorial extensions

If this is right

  • If correct, generative multimodal autoregression is a viable drop-in pre-training objective for generalist vision encoders, reducing the need for the large batch sizes and careful data filtering that contrastive methods require.
  • AIMV2-3B's 89.5% ImageNet-1k accuracy with a frozen trunk implies that representation quality comparable to the best discriminative models can come from image-plus-text next-token prediction.
  • The reported scaling behavior (performance improves with model size and sample count, while the optimal size grows with compute) suggests the recipe will keep improving as models and data grow, in line with LLM-style scaling.
  • Consistent gains on text-rich benchmarks such as TextVQA, DocVQA, and ChartQA indicate that the multimodal objective is especially useful when downstream tasks require fine-grained reading and localization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • What the paper leaves open is whether the margin over CLIP and SigLIP is an objective effect or a data effect; the inference that the objective alone drives the gains would be confirmed by swapping in matched data for the baselines.
  • The dense patch-level supervision suggests an untested prediction: AIMV2 should degrade less than captioning-only models on tasks needing fine spatial detail, such as small-object detection, which the paper's detection results roughly support but do not isolate.
  • One could extend the recipe to video or audio by treating frame or spectrogram patches as additional sequence tokens in the same factorization; the paper does not report such experiments.
  • The prefix attention trick implies that the encoder is trained to produce useful representations from partial images, which may explain the robustness to cropping and tiling seen in the high-resolution evaluations.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces AIMV2, a family of ViT-based vision encoders pre-trained with a multimodal autoregressive objective. The vision encoder uses prefix attention, and a causal multimodal decoder predicts both raw image patches (ℓ2 loss) and caption tokens (cross-entropy) on a 12B image-text mixture (DFN, COYO, proprietary HQITP, and synthetic captions). The authors evaluate the resulting encoders on frozen-trunk recognition, open-vocabulary detection and grounding, multimodal instruction tuning, in-context learning, zero-shot LiT, and native-resolution adaptation, reporting strong results including 89.5% ImageNet-1k top-1 with a frozen trunk. The central methodological claim is that joint autoregressive prediction of patches and text yields a simple, scalable, and generalist vision encoder that matches or outperforms contrastive pre-training.

Significance. If the central claim holds, this is a significant result: it demonstrates that a straightforward autoregressive multimodal objective can compete with contrastive objectives for vision-encoder pre-training, with the added benefits of dense supervision, modest batch sizes, and natural compatibility with LLM-based multimodal pipelines. The paper's strengths include a carefully controlled ablation in Table 9b comparing AIMV2 against CLIP and CapPa under identical architecture and data, a scaling analysis in Figure 2 that mirrors Hoffmann-style compute-optimal behavior, and a public code release. These elements make the core method claim credible. However, the broader advertised claim that AIMV2 'consistently outperforms state-of-the-art contrastive models' is only partially supported: the headline comparisons in Tables 3 and 7 mix method and data differences, and the paper's own results in Table 5 and Appendix D.3 show cases where baselines outperform AIMV2. The contribution is valuable but the presentation needs to separate method advantage from data advantage.

major comments (3)
  1. [§2.3, Table 2 vs. Tables 3 and 7; §5, Table 9b] The headline gains over SigLIP and CLIP are confounded with pre-training data. AIMV2 is trained on a 12B mixture that includes 3.8B synthetic DFN captions, 431.5M synthetic HQITP captions, and 564.6M proprietary HQITP alt-text pairs (Table 2), while the SigLIP and CLIP checkpoints used in Tables 3 and 7 were trained on different private corpora. The controlled same-data comparison in Table 9b covers only CLIP and CapPa at 2B pairs, not the deployed SigLIP models. A contrastive model trained on the same caption quality could plausibly close a substantial portion of the reported margins, since data filtering and synthetic captions are known to improve contrastive models as well. This does not falsify the method claim, but it means the abstract's 'consistently outperforms state-of-the-art contrastive models' is not fully supported by the evidence as presented. I recommend either adding a same-data SigLIP-style baseline or explicitly qualifying the claim as holding for AIMV2's data mixture.
  2. [Abstract and Conclusion vs. Table 5 and Table D3] The claim of 'consistently outperforms' is contradicted by results within the paper itself. Table 5 shows SigLIP ViT-So400m at 80.4 zero-shot ImageNet top-1 versus 77.0 for AIMV2-3B, and Table D3 shows DINOv2 outperforming AIMV2 on COCO detection/segmentation (55.5 vs. 54.0 AP). The text in Section 4.1 and Appendix D.2 acknowledges these cases, but the abstract and conclusion do not carry the same qualification. Please temper the claims to 'outperforms or matches' with explicit exceptions, or restrict the claim to the specific settings where the controlled evidence supports it.
  3. [§5, Table 9c–9f] Several design choices are recommended on the basis of differences that are likely within training noise. For example, Table 9c reports TextVQA 37.5 for α=0.4 versus 37.4 for α=0.2 and 0.6; Table 9e shows decoder width 512 at 35.9 versus 1536 at 36.9; and Table 9f shows depth 12 versus 16 at 37.5 versus 36.6. None of these experiments report multiple seeds or error bars. I am not asking for a full seed study, but the text should avoid presenting these differences as conclusive evidence for a particular hyperparameter or architecture choice, or the authors should add at least a few repeated runs for the key ablations.
minor comments (5)
  1. [§2.3 and Table 2] The term 'synthetic' captions is used without specifying the captioner model or its filtering procedure, despite citing Lai et al. [63]. A sentence describing the captioning pipeline and any quality filtering would help reproducibility.
  2. [§4.3.2, Table 8] The in-context learning comparison reports only results for OAI CLIP and DFN-CLIP as quoted from McKinzie et al. [85], without the MM1 ViT-L baseline under identical pre-training data. At minimum, clarify whether the ICL comparison holds the instruction-tuning data fixed.
  3. [Throughout] There are numerous typos and grammatical slips: 'factorizatized' (§2.1), 'task' for 'tasks' (§2.4), 'hyperaparmeters' in Tables A1, A2, C1, and 'the model’s predicted patch ˆxi(θ)' with mismatched parentheses in §2.1. A careful proofreading pass is needed.
  4. [§5, Table 9b caption] The caption reads 'AIMV2 vs. CLIP..' but the table also includes CapPa; the caption should list all three methods. Also, the CapPa row is trained at batch size 8k only, while CLIP is given at 8k and 16k; note in the text why CapPa was not run at 16k.
  5. [§4.1, Table 3] The comparison of AIMV2-3B at 448px against baselines at 224px is apples-to-oranges. The table caption notes the resolution, but the text should explicitly state that the 89.5% result uses a higher-resolution fine-tuned model, not the base pre-training resolution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claims rest on external downstream benchmarks and a controlled same-data ablation, not on a derivation that reduces to its own inputs.

full rationale

AIMV2 is an empirical pre-training paper; it does not claim a mathematical derivation of benchmark performance from first principles. The main claims are that a vision encoder trained with a multimodal autoregressive objective (image-patch regression plus caption cross-entropy) transfers well to recognition, grounding, and multimodal instruction tuning. These claims are supported by evaluations on external benchmarks (ImageNet-1k, VQAv2, TextVQA, COCO, LVIS, etc.) that are not part of the training objective. The load-bearing controlled comparison is Table 9b, where AIMV2, CLIP, and CapPa are trained with identical architectures, data, and hyperparameters ("All models are trained using identical architectures, incorporating SwiGLU and RMSNorm, and are pre-trained using the same dataset of image-text pairs"), and AIMV2 wins by 11-13 points on TextVQA. That ablation is the evidence that the objective itself, not the data, drives the gains over contrastive and captioning baselines. Self-citations do appear: [33] (AIM, the authors' own prior work) is cited for prefix attention and as a baseline, [35] (DFN) for data filtering with overlapping authors, and [63] (Lai et al.) for synthetic captions. None of these is used as an unverified premise that forces the paper's conclusion; prefix attention is an architectural choice that is ablated (Table 9a), DFN is a public dataset, and the synthetic-caption pipeline is a data ingredient, not a proof step. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via citation that is then presented as externally validated. The headline comparisons against SigLIP and OAI CLIP are not fully controlled because AIMV2 uses 12B pairs including proprietary HQITP and roughly 4.2B synthetic captions while the baselines used different private corpora, and the abstract's "consistently outperforms" claim is qualified by the zero-shot LiT result in Table 5 where SigLIP leads by 3.4 points. These are legitimate correctness and external-validity concerns, but they are not circularity: the paper's central method claim survives the controlled ablation and is evaluated against external benchmarks. Accordingly, no circular step can be exhibited, and the circularity score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method is empirical; the central claim rests on the pre-training objective, the data mixture, and the downstream evaluation protocol. The main free parameter is the loss weight alpha; decoder width and depth are also tuned. No new entities are introduced, but several domain assumptions about data quality and generalization are unproved.

free parameters (3)
  • alpha (pixel loss weight) = 0.4
    Balancing weight for L_img in L_text + alpha * L_img; ablated in Table 9c, robust between 0.2 and 0.6.
  • multimodal decoder width = 1024
    Ablated in Table 9e; width 512 is lower, width 1536 slightly lower on IN-1k.
  • multimodal decoder depth = 12
    Ablated in Table 9f; depth 8 lower, depth 16 lower on TextVQA.
assumptions (3)
  • domain assumption Web-scale image-text pairs, including synthetic captions, provide a sufficient signal for visual representation learning.
    Section 2.3 and Table 2; no guarantee that caption quality or coverage does not confound comparison with baselines.
  • domain assumption Pixel-level MSE regression on normalized patches transfers to semantic downstream tasks.
    Section 2.1 and Section 5 ablations; supported empirically but not theoretically.
  • domain assumption Downstream benchmark results generalize to real-world performance and are not inflated by test leakage from pre-training data.
    Standard assumption for web-scale pre-training; the paper does not report contamination checks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Autoregressive Pre-training of Large Vision Encoders." pith.science (2026). https://pith.science/paper/SPRQZIL2

@misc{pith2026241114402,
  author       = {Pith},
  title        = {Pith review of: Multimodal Autoregressive Pre-training of Large Vision Encoders},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SPRQZIL2}},
  note         = {Machine review of arXiv:2411.14402}
}
read the original abstract

We introduce a novel method for pre-training of large-scale vision encoders. Building on recent advancements in autoregressive pre-training of vision models, we extend this framework to a multimodal setting, i.e., images and text. In this paper, we present AIMV2, a family of generalist vision encoders characterized by a straightforward pre-training process, scalability, and remarkable performance across a range of downstream tasks. This is achieved by pairing the vision encoder with a multimodal decoder that autoregressively generates raw image patches and text tokens. Our encoders excel not only in multimodal evaluations but also in vision benchmarks such as localization, grounding, and classification. Notably, our AIMV2-3B encoder achieves 89.5% accuracy on ImageNet-1k with a frozen trunk. Furthermore, AIMV2 consistently outperforms state-of-the-art contrastive models (e.g., CLIP, SigLIP) in multimodal image understanding across diverse settings.

Figures

Figures reproduced from arXiv: 2411.14402 by the authors.

Figure 1
Figure 1. AIMV2 pre-training Overview. (Left) Image patches are processed by a vision encoder trained with prefix attention [33, 95]. The resulting visual representations are concatenated with the text embeddings of their corresponding captions. This combined multimodal sequence is then processed by a joint decoder. The model is pre-trained to autoregressively reconstruct the shifted input. (Right) Pseudocode for the forward … view at source ↗
Figure 2
Figure 2. Scaling properties of AIMV2. (Left) Given a fixed pre-training data size, increasing the number of parameters always leads to an improvement in the validation loss. (Right) The optimal model size varies based on the pre-training compute budget. Larger models perform worse than smaller ones when severely undertrained but improves consistently as the compute budget increases. This behavior is consistent with that repo… view at source ↗
Figure 3
Figure 3. Scaling capacity and resolution. AIMV2 shows strong scalability with respect to model parameters, measured in frozen￾trunk top-1 accuracy for IN-1k. This behavior is consistent when scaling image resolution. 500M 1B 2B 4B 84 85 86 87 Image-text Pairs IN-1k Accuracy (%) AIMV2 Cap L H 1B 85 85.5 86 86.5 87 Model Size IN-1k Accuracy (%) AIMV2 Cap [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Impact of Scaling Resolution. The performance boost achieved byAIMV2 persists after scaling input resolution via tiling Lin et al. [72], Liu et al. [73] compared to popular vision backbones for VLMs such as OAI CLIP and SigLIP. MM-Grounding-DINO [74, 136] but adapt ViT…
Figure 6
Figure 6. Figure 6: Instruction tuning under different settings. [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  2. Vision-Language Models Do Not Understand Negation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Current vision-language models largely ignore negation, and a new 79k-example benchmark plus synthetic fine-tuning data yields measurable but incomplete improvements.

  3. Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A gradient alignment score against a reference model's weights selects and reweights training samples, improving data efficiency in image classification and CLIP pretraining.

  4. Analyzing Finetuning Representation Shift for Multimodal LLMs Steering

    cs.AI 2025-01 conditional novelty 6.0 of 10

    Concept shift vectors, computed as mean activation differences, can partially recover fine-tuned multimodal LLM concepts and steer model outputs without additional training.

  5. SigLIP-HD by Fine-to-Coarse Supervision

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Fine-to-coarse L1 supervision lets a standard-resolution SigLIP 2 encoder produce better visual tokens for MLLMs without higher-resolution inference.

  6. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Codec-guided sparse patch selection plus a lightweight speak/silent gate yields a 4B streaming VLM that is competitive on static tasks, stronger on video/spatial benchmarks, and much cheaper at inference.

  7. LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Inserting a pixel-shuffle plus residual patch-merge layer inside the vision encoder compresses visual tokens more efficiently than post-encoder compression, at modest accuracy cost.

  8. Hierarchical Pre-Training of Vision Encoders with Large Language Model

    cs.CV 2026-03 reject novelty 4.0 of 10

    A three-stage pre-training scheme that feeds multi-layer vision features into an LLM reports marginal benchmark gains, but lacks data, code, and ablations needed to support the claim.

  9. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.

  10. SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.

  11. Advancements in Medical Image Classification through Fine-Tuning Natural Domain Foundation Models

    eess.IV 2025-05 conditional novelty 4.0 of 10

    Fine-tuning recent natural-domain foundation models, especially AIMv2, improves medical image classification accuracy across mammography, skin lesion, retinopathy, and chest X-ray benchmarks.

  12. Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach

    eess.AS 2025-05 conditional novelty 4.0 of 10

    Llama-SMoP-DEDR, a sparse mixture of projectors with modality-specific experts and routers, lowers word error rate for LLM-based AVSR on LRS3, mainly with smaller LLMs.

  13. From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs

    cs.CV 2025-02 conditional novelty 4.0 of 10

    Adding an L2 loss that pushes the language model's image hidden states back toward the input image embeddings improves LLaVA-style models on several VQA benchmarks, with some benchmarks unaffected or slightly worse.

  14. Visual RAG: Expanding MLLM visual knowledge without fine-tuning

    cs.CV 2025-01 conditional novelty 4.0 of 10

    Retrieval-selected demonstration examples let a multimodal LLM classify images as accurately as random many-shot prompting with far fewer examples.

Reference graph

Works this paper leans on

137 extracted references · 37 canonical work pages · cited by 14 Pith papers

  1. [1]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1

  2. [2]

    Nocaps: Novel object cap- tioning at scale

    Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. Nocaps: Novel object cap- tioning at scale. In ICCV, 2019. 7, 16

  3. [3]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, An- toine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Se- bastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Bink...

  4. [4]

    Self-supervised learning from im- ages with a joint-embedding predictive architecture

    Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bo- janowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from im- ages with a joint-embedding predictive architecture. arXiv preprint arXiv:2301.08243, 2023. 9

  5. [5]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. 1

  6. [6]

    Qwen-vl: A frontier large vision-language model with versatile abilities

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966 ,

  7. [7]

    From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge

    Peter Bandi, Oscar Geessink, Quirine Manson, Mar- cory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, et al. From detection of individual metastases to classification of lymph node status at the pa- tient level: the camelyon17 challenge. IEEE Transactions on Medical Imaging, 2018. 15

  8. [8]

    BEiT: Bert pre- training of image transformers

    Hangbo Bao, Li Dong, and Furu Wei. BEiT: Bert pre- training of image transformers. In ICLR, 2022. 1, 9

Show all 137 references
  1. [9]

    Flexivit: One model for all patch sizes

    Lucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron, Simon Kornblith, Xiaohua Zhai, Matthias Minderer, Michael Tschannen, Ibrahim Alabdulmohsin, and Filip Pavetic. Flexivit: One model for all patch sizes. In CVPR, 2023. 4

  2. [10]

    Food-101 – mining discriminative components with ran- dom forests

    Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with ran- dom forests. In ECCV, 2014. 15

  3. [11]

    Time series analysis: forecasting and control

    George EP Box, Gwilym M Jenkins, Gregory C Reinsel, and Greta M Ljung. Time series analysis: forecasting and control. John Wiley & Sons, 2015. 8

  4. [12]

    Language models are few-shot learners

    Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Nee- lakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. preprint arXiv:2005.14165, 2020. 8

  5. [13]

    Coyo-700m: Image-text pair dataset, 2022

    Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset, 2022. 3, 9

  6. [14]

    End-to-end object detection with transformers

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nico- las Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV,

  7. [15]

    Unsupervised learn- ing of visual features by contrasting cluster assignments

    Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learn- ing of visual features by contrasting cluster assignments. In NeurIPS, 2020. 9

  8. [16]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021. 9

  9. [17]

    A generative approach for wikipedia-scale visual entity recognition

    Mathilde Caron, Ahmet Iscen, Alireza Fathi, and Cordelia Schmid. A generative approach for wikipedia-scale visual entity recognition. In CVPR, 2024. 9

  10. [18]

    Mmdetection: Open mmlab detection toolbox and benchmark

    Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, Zheng Zhang, Dazhi Cheng, Chenchen Zhu, Tianheng Cheng, Qijie Zhao, Buyu Li, Xin Lu, Rui Zhu, Yue Wu, Jifeng Dai, Jingdong Wang, Jianping Shi, Wanli Ouyang,...

  11. [19]

    Generative pre- training from pixels

    Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Hee- woo Jun, David Luan, and Ilya Sutskever. Generative pre- training from pixels. In ICML, 2020. 9

  12. [20]

    A simple framework for contrastive learning of visual representations

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 9

  13. [21]

    Microsoft coco captions: Data collection and eval- uation server

    Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and eval- uation server. arXiv preprint arXiv:1504.00325, 2015. 7

  14. [22]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhang- hao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 7

  15. [23]

    Palm: Scaling language modeling with pathways

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arxiv 2022. arXiv preprint arXiv:2204.02311 ,

  16. [24]

    Functional map of the world

    Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018. 15

  17. [25]

    Cimpoi, S

    M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi. Describing textures in the wild. In CVPR, 2014. 15

  18. [26]

    Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution

    Mostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek, Matthias Minderer, Mathilde Caron, An- dreas Steiner, Joan Puigcerver, Robert Geirhos, Ibrahim M 10 Alabdulmohsin, et al. Patch n’pack: Navit, a vision trans- former for any aspect ratio and resolution. Advances i...

  19. [27]

    Imagenet: A large-scale hierarchical im- age database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical im- age database. In 2009 IEEE conference on computer vision and pattern recognition, 2009. 15

  20. [28]

    Virtex: Learning visual representations from textual annotations

    Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021. 9

  21. [29]

    Unsu- pervised visual representation learning by context predic- tion

    Carl Doersch, Abhinav Gupta, and Alexei A Efros. Unsu- pervised visual representation learning by context predic- tion. In ICCV, 2015. 1

  22. [30]

    An image is worth 16x16 words: Trans- formers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. In ICLR, 2021. 2

  23. [32]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783,

  24. [33]

    Scalable pre-training of large autoregressive image models

    Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander Toshev, Vaishaal Shankar, Joshua M Susskind, and Armand Joulin. Scalable pre-training of large autoregressive image models. arXiv preprint arXiv:2401.08541, 2024. 1, 2, 3, 5, 8, 9

  25. [34]

    Taming transformers for high-resolution image synthesis

    Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. InCVPR,

  26. [35]

    Data filtering networks

    Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. Data filtering networks. arXiv preprint arXiv:2309.17425, 2023. 1, 3, 5, 9

  27. [36]

    Slowfast networks for video recognition

    Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. 2019 ieee. In ICCV, 2018. 1

  28. [37]

    Improved base- lines for vision-language pre-training

    Enrico Fini, Pietro Astolfi, Adriana Romero-Soriano, Jakob Verbeek, and Michal Drozdzal. Improved base- lines for vision-language pre-training. arXiv preprint arXiv:2305.08675, 2023. 9

  29. [38]

    Mme: A comprehensive eval- uation benchmark for multimodal large language models

    Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, et al. Mme: A comprehensive eval- uation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023. 16

  30. [39]

    Un- supervised representation learning by predicting image ro- tations

    Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Un- supervised representation learning by predicting image ro- tations. arXiv preprint arXiv:1803.07728, 2018. 9

  31. [40]

    Making the v in vqa matter: El- evating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: El- evating the role of image understanding in visual question answering. In CVPR, 2017. 7, 8

  32. [41]

    Making the v in vqa matter: El- evating the role of image understanding in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: El- evating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, 2017. 16

  33. [42]

    Bootstrap your own latent-a new approach to self-supervised learning

    Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. NeurIPS, 2020. 9

  34. [43]

    Lvis: A dataset for large vocabulary instance segmentation, 2019

    Agrim Gupta, Piotr Doll ´ar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation, 2019. 6

  35. [44]

    Vizwiz grand challenge: Answering visual questions from blind people

    Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham. Vizwiz grand challenge: Answering visual questions from blind people. In CVPR, 2018. 7

  36. [45]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR,

  37. [46]

    Mask r-cnn

    Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. In ICCV, 2017. 1

  38. [47]

    Mask r-cnn

    Kaiming He, Georgia Gkioxari, Piotr Doll ´ar, and Ross Gir- shick. Mask r-cnn. 2018. 18

  39. [48]

    Masked autoencoders are scal- able vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scal- able vision learners. In CVPR, 2022. 1, 2, 5, 9

  40. [49]

    Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,

    Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learn- ing benchmark for land use and land cover classification,

  41. [50]

    Training compute-optimal large language mod- els

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language mod- els. arXiv preprint arXiv:2203.15556, 2022. 1, 3, 4

  42. [51]

    Scaling up vision-language pre-training for image captioning

    Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. Scaling up vision-language pre-training for image captioning. In CVPR, 2022. 9

  43. [52]

    Gqa: A new dataset for real-world visual reasoning and compositional question answering

    Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , 2019. 6, 16

  44. [53]

    Gqa: A new dataset for real-world visual reasoning and compositional question answering

    Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019. 7

  45. [54]

    Scaling up visual and vision-language representation learning with noisy text supervision

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, 2021. 1, 9 11

  46. [55]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361,

  47. [56]

    Deep visual-semantic alignments for generating image descriptions

    Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In CVPR,

  48. [57]

    Referitgame: Referring to objects in pho- tographs of natural scenes

    Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in pho- tographs of natural scenes. In Proceedings of the 2014 con- ference on empirical methods in natural language process- ing (EMNLP), 2014. 6

  49. [58]

    Big transfer (bit): General visual representation learning

    Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In ECCV, 2020. 9

  50. [59]

    3d object representations for fine-grained categorization

    Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), 2013. 15

  51. [60]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 15

  52. [61]

    Imagenet classification with deep convolutional neural net- works

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural net- works. NeurIPS, 2012. 9

  53. [62]

    Mammut: A simple architec- ture for joint learning for multimodal tasks

    Weicheng Kuo, AJ Piergiovanni, Dahun Kim, Xiyang Luo, Ben Caine, Wei Li, Abhijit Ogale, Luowei Zhou, Andrew Dai, Zhifeng Chen, et al. Mammut: A simple architec- ture for joint learning for multimodal tasks. arXiv preprint arXiv:2303.16839, 2023. 9

  54. [63]

    Revisit large-scale image-caption data in pre-training multimodal foundation models

    Zhengfeng Lai, Vasileios Saveris, Chen Chen, Hong-You Chen, Haotian Zhang, Bowen Zhang, Juan Lao Tebar, Wenze Hu, Zhe Gan, Peter Grasch, et al. Revisit large-scale image-caption data in pre-training multimodal foundation models. arXiv preprint arXiv:2410.02740, 2024. 3

  55. [64]

    Seed-bench: Benchmarking multi- modal llms with generative comprehension

    Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yix- iao Ge, and Ying Shan. Seed-bench: Benchmarking multi- modal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. 16

  56. [65]

    Align before fuse: Vision and language representation learning with momentum distillation

    Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. NeurIPS, 2021. 9

  57. [66]

    Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation. In ICML, 2022. 9

  58. [67]

    Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023. 9

  59. [68]

    Exploring plain vision transformer backbones for object de- tection

    Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object de- tection. In European conference on computer vision, 2022. 6, 17, 18

  60. [69]

    Exploring plain vision transformer backbones for object de- tection, 2022

    Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He. Exploring plain vision transformer backbones for object de- tection, 2022. 6

  61. [70]

    Mini-gemini: Mining the potential of multi-modality vision language models

    Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814,

  62. [71]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th Eu- ropean Conference, Zurich, Switzerland, September 6-12, 2014, Proceed...

  63. [72]

    Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models

    Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023. 6, 7

  64. [73]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In CVPR, 2024. 1, 6, 7, 9, 16, 17

  65. [74]

    Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection. 2024. 6

  66. [75]

    Sgdr: Stochastic gradient descent with warm restarts

    Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In ICLR, 2017. 15

  67. [76]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 15, 16

  68. [77]

    Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action

    Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Anirud- dha Kembhavi. Unified-io 2: Scaling autoregressive mul- timodal models with vision language audio and action. In CVPR, 2024. 9

  69. [78]

    Learn to explain: Multimodal reasoning via thought chains for science question answering

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Ad- vances in Neural Information Processing Systems, 2022. 16

  70. [79]

    Generation and comprehension of unambiguous object descriptions

    Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016. 6

  71. [80]

    Ok-vqa: A visual question answering benchmark requiring external knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. In CVPR, 2019. 7

  72. [81]

    Ok-vqa: A visual question answering benchmark requiring external knowledge

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. Ok-vqa: A visual question answering benchmark requiring external knowledge. InConference on Computer Vision and Pattern Recognition (CVPR) , 2019. 16

  73. [82]

    Chartqa: A benchmark for question 12 answering about charts with visual and logical reasoning

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question 12 answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244, 2022. 16

  74. [83]

    Docvqa: A dataset for vqa on document images

    Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Pro- ceedings of the IEEE/CVF winter conference on applica- tions of computer vision, 2021. 16

  75. [84]

    Infographicvqa

    Minesh Mathew, Viraj Bagal, Rub `en Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Ap- plications of Computer Vision, 2022. 16

  76. [85]

    Mm1: Meth- ods, analysis & insights from multimodal llm pre-training

    Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al. Mm1: Meth- ods, analysis & insights from multimodal llm pre-training. arXiv preprint arXiv:2403.09611, 2024. 1, 6, 7, 9

  77. [86]

    Unsupervised learning of visual representations by solving jigsaw puzzles

    Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In ECCV,

  78. [87]

    Maxime Oquab, Timoth ´ee Darcet, Theo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernan- dez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Ass- ran, N...

  79. [88]

    O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V . Jawahar. Cats and dogs. In CVPR, 2012. 15

  80. [89]

    Moment matching for multi-source domain adaptation

    Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In ICCV, 2019. 15

  81. [90]

    Plummer, Liwei Wang, Christopher M

    Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazeb- nik. Flickr30k entities: Collecting region-to-phrase cor- respondences for richer image-to-sentence models. IJCV,

  82. [91]

    Dataset decomposition: Faster llm training with variable sequence length curriculum

    Hadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pa- van Kumar Anasosalu Vasu, Cem Koc, Vaishaal Shankar, and Oncel Tuzel. Dataset decomposition: Faster llm training with variable sequence length curriculum. arXiv preprint arXiv:2405.13226, 2024. 4

  83. [92]

    Improving language understanding by genera- tive pre-training

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by genera- tive pre-training. 2018. 1, 8

  84. [93]

    Language models are unsupervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 2019. 1, 8

  85. [94]

    Learn- ing transferable visual models from natural language super- vision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 1, 5, 6, 8, 9, 16

  86. [95]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Ma- chine Learning Research, 21(1), 2020. 2, 3

  87. [96]

    Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context

    Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrit- twieser, et al. Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context. arXiv...

  88. [97]

    Imagenet-21k pretraining for the masses

    Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. arXiv preprint arXiv:2104.10972, 2021. 9

  89. [98]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 9

  90. [99]

    Learning visual representations with caption annotations

    Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus. Learning visual representations with caption annotations. In Computer Vision–ECCV 2020: 16th European Confer- ence, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16. Springer, 2020. 9

  91. [100]

    Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs. In NeurIPS Workshop, 2021. 9

  92. [101]

    Objects365: A large-scale, high-quality dataset for object detection

    Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In ICCV, 2019. 6

  93. [102]

    Glu variants improve transformer

    Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020. 3

  94. [103]

    When do we not need larger vision mod- els? arXiv preprint arXiv:2403.13043, 2024

    Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, and Trevor Darrell. When do we not need larger vision mod- els? arXiv preprint arXiv:2403.13043, 2024. 7

  95. [104]

    Unival: Unified model for image, video, audio and language tasks

    Mustafa Shukor, Corentin Dancette, Alexandre Rame, and Matthieu Cord. Unival: Unified model for image, video, audio and language tasks. Transactions on Machine Learn- ing Research Journal, 2023. 9

  96. [105]

    Textcaps: a dataset for image caption- ing with reading comprehension

    Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh. Textcaps: a dataset for image caption- ing with reading comprehension. In ECCV, 2020. 7, 16

  97. [106]

    Towards vqa models that can read

    Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, 2019. 8, 16

  98. [107]

    Towards vqa models that can read

    Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In CVPR, 2019. 7

  99. [108]

    Revisiting unreasonable effectiveness of data in deep learning era

    Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhi- nav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, 2017. 9

  100. [109]

    Eva-clip: Improved training techniques for clip at scale

    Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 9

  101. [110]

    Generative pretraining in mul- timodality

    Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative pretraining in mul- timodality. arXiv preprint arXiv:2307.05222, 2023. 9 13

  102. [111]

    Generative multimodal models are in-context learners

    Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiy- ing Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative multimodal models are in-context learners. In CVPR, 2024. 9

  103. [112]

    Taylor, B

    J. Taylor, B. Earnshaw, B. Mabey, M. Victors, and J. Yosin- ski. Rxrx1: An image set for cellular morphological vari- ation across many experimental batches. In ICLR, 2019. 15

  104. [113]

    Chameleon: Mixed-modal early-fusion foundation models

    Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818 ,

  105. [114]

    Particu- lar object retrieval with integral max-pooling of cnn activa- tions

    Giorgos Tolias, Ronan Sicre, and Herv ´e J ´egou. Particu- lar object retrieval with integral max-pooling of cnn activa- tions. arXiv preprint arXiv:1511.05879, 2015. 1

  106. [115]

    Cambrian- 1: A fully open, vision-centric exploration of multimodal llms

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian- 1: A fully open, vision-centric exploration of multimodal llms. arXiv:2406.16860, 2024. 1, 6, 7, 16, 17

  107. [116]

    Llama: Open and efficient foundation language mod- els

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Bap- tiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language mod- els. arXiv preprint arXiv:2302.13971, 2023. 1, 3, 8

  108. [117]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 3, 8

  109. [118]

    Image captioners are scalable vision learners too

    Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiao- hua Zhai, Neil Houlsby, and Lucas Beyer. Image captioners are scalable vision learners too. NeurIPS, 2024. 1, 5, 6, 8, 9

  110. [119]

    The inaturalist species classification and detection dataset

    Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In CVPR, 2018. 15

  111. [120]

    Rotation equivariant cnns for digital pathology

    Bastiaan S Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling. Rotation equivariant cnns for digital pathology. In Medical Image Computing and Com- puter Assisted Intervention, 2018. 15

  112. [121]

    Show and tell: A neural image caption gen- erator

    Oriol Vinyals, Alexander Toshev, Samy Bengio, and Du- mitru Erhan. Show and tell: A neural image caption gen- erator. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2015. 9

  113. [122]

    Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

    Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International conference on machine learn- i...

  114. [123]

    Simvlm: Simple visual language model pretraining with weak supervision

    Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. arXiv preprint arXiv:2108.10904, 2021. 9

  115. [124]

    Vila-u: a unified foundation model inte- grating visual understanding and generation

    Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, et al. Vila-u: a unified foundation model inte- grating visual understanding and generation. arXiv preprint arXiv:2409.04429, 2024. 9

  116. [125]

    Show-o: One single transformer to unify multimodal under- standing and generation

    Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, Zhenheng Yang, and Mike Zheng Shou. Show-o: One single transformer to unify multimodal under- standing and generation. arXiv preprint arXiv:2408.12528,

  117. [126]

    Show, attend and tell: Neural image cap- tion generation with visual attention

    Kelvin Xu. Show, attend and tell: Neural image cap- tion generation with visual attention. arXiv preprint arXiv:1502.03044, 2015. 9

  118. [127]

    From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

    Peter Young, Alice Lai, Micah Hodosh, and Julia Hock- enmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2, 2014. 6

  119. [128]

    Vector-quantized image modeling with improved vqgan

    Jiahui Yu, Xin Li, Jing Yu Koh, Han Zhang, Ruoming Pang, James Qin, Alexander Ku, Yuanzhong Xu, Jason Baldridge, and Yonghui Wu. Vector-quantized image modeling with improved vqgan. ArXiv, 2021. 9

  120. [129]

    Coca: Contrastive captioners are image-text foundation models

    Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mo- jtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. TMLR, 2022. 5, 9

  121. [130]

    Berg, and Tamara L

    Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. Modeling context in referring expressions, 2016. 6

  122. [131]

    Scaling autore- gressive multi-modal models: Pretraining and instruction tuning

    Lili Yu, Bowen Shi, Ramakanth Pasunuru, Benjamin Muller, Olga Golovneva, Tianlu Wang, Arun Babu, Binh Tang, Brian Karrer, Shelly Sheynin, et al. Scaling autore- gressive multi-modal models: Pretraining and instruction tuning. arXiv preprint arXiv:2309.02591, 2023. 9

  123. [132]

    Lit: Zero-shot transfer with locked-image text tuning

    Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning. In CVPR, 2022. 2, 6

  124. [133]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, 2023. 1, 3, 5, 9, 15, 16

  125. [134]

    Root mean square layer normalization

    Biao Zhang and Rico Sennrich. Root mean square layer normalization. NeurIPS, 2019. 3

  126. [135]

    Colorful image colorization

    Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In ECCV, 2016. 9

  127. [136]

    An open and comprehensive pipeline for unified object grounding and detection, 2024

    Xiangyu Zhao, Yicheng Chen, Shilin Xu, Xiangtai Li, Xin- jiang Wang, Yining Li, and Haian Huang. An open and comprehensive pipeline for unified object grounding and detection, 2024. 6

  128. [137]

    ibot: Image bert pre- training with online tokenizer

    Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. ibot: Image bert pre- training with online tokenizer. In ICLR, 2022. 1, 9

  129. [138]

    supreme gasoline

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. In ICLR, 2024. 6 14 A. Hyperparamters Pre-training. We outline the optimization hyperaparmeters and data augmentations...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.