Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

FastVLM: Efficient Vision Encoding for Vision Language Models

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A single hybrid vision encoder, FastViTHD, that downsamples images 64x before self-attention produces far fewer visual tokens and encodes high-resolution images several times faster than ViT, SigLIP, and ConvNeXt encoders, while matching…

desk verdict FastViTHD is a genuinely useful token-efficient encoder with a real Pareto story, but the headline latency ratios rest on Apple-only, selectively reported benchmarks. read the letter →

arxiv 2412.13303 v2 pith:Q4FTJ4M2 submitted 2024-12-17 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords visionlanguagemodelshybridencodertime-to-first-tokenhigh-resolutionimageunderstandingvisualtokenefficiencyFastViTHDLLMprefillingon-deviceinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

FastVLM claims that the main bottleneck for high-resolution vision-language models is not just the vision encoder's own latency but the flood of visual tokens it sends to the language model. The paper introduces FastViTHD, a hybrid convolutional-transformer encoder with an extra downsampling stage that yields 4x fewer tokens than FastViT and 16x fewer than ViT-L/14 at the same input resolution. In controlled comparisons, FastViTHD achieves 3.2x faster time-to-first-token and 3.6x smaller size than SigLIP-SO400M, and 2.3x faster and 1.7x smaller than ConvNeXt, with equal or better benchmark accuracy. If correct, this means high-resolution, text-rich VLM inference could run on-device at a fraction of the current latency without sacrificing accuracy, and without needing token pruning heuristics.

What carries the argument

FastViTHD is a five-stage hybrid encoder: the first three stages use RepMixer convolutional blocks, and the last two stages use self-attention, with an added patch-embedding layer that downsamples the input by a total factor of 64. This makes self-attention operate on a small tensor (16x16 for a 1024x1024 image), cutting both encoder latency and the number of tokens passed to the LLM, which is what reduces LLM prefilling time and thus time-to-first-token. Multi-scale features pooled from earlier stages with depthwise convolutions add a small accuracy boost. The architecture is pretrained with MobileCLIP's reinforced image-text pipeline and then fine-tuned end-to-end in the LLaVA-1.5 two-stage recipe.

What would settle it

Run the exact same controlled comparison (same LLaVA-1.5 training, same LLM) on a different hardware platform, e.g., an NVIDIA GPU with standard PyTorch or TensorRT, benchmarking all encoders including those excluded from the MLX comparison, and compare TTFT at matched token counts; if FastViTHD's time-to-first-token advantage falls below roughly 2x, the central claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that a hybrid vision backbone with a 64x downsampled final stage dominates isotropic ViTs and pure convolutional encoders on the accuracy-latency frontier for vision-language models. FastViTHD runs self-attention only on a heavily downsampled feature map, so it encodes a 1024x1024 image in 235 ms on an M1 MacBook Pro while producing only 256 visual tokens, versus ViT-L/14's 576 tokens from a 336x336 image. Across LLMs of 0.5B, 1.5B, and 7B parameters, the Pareto-optimal curve of FastViTHD is over 2.5 points better on the Average-5 metric than the best FastViT curve, and it reaches a target VLM performance up to 3x faster. The paper further shows that scaling input resolution directly beats tiling (AnyRes) except at extreme resolutions, and that a hierarchical encoder with few tokens beats token-pruning methods applied to ViTs.

Load-bearing premise

The claimed speedups are measured only on Apple-silicon devices using CoreML for encoders and MLX for the LLM, and only for models that convert cleanly to those formats, so the latency comparisons could be artifacts of the conversion stack rather than of the architectures themselves.

Editorial extensions

If this is right

  • High-resolution VLMs for text-rich images could run at roughly one-third the time-to-first-token of prior ViT-based models, making on-device document and chart understanding practical.
  • The need for token-pruning and resampling modules disappears for hierarchical encoders: simply training at lower input resolution yields token counts as low as 16 while outperforming prune-then-feed ViT methods.
  • Because the encoder produces far fewer tokens, smaller LLMs (e.g., 0.5B) can handle high-resolution inputs better, and the Pareto analysis shows which (resolution, LLM size) pair is optimal for a given latency budget.
  • Scaling visual instruction-tuning data further improves FastVLM, suggesting that the efficient token representation transfers to larger datasets and stronger benchmarks.
  • Dynamic-resolution tiling is largely unnecessary with FastViTHD: static resolution scaling is the better accuracy-latency trade-off except at extreme resolutions like 1536x1536.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The headline 85x TTFT improvement versus LLaVA-OneVision mixes several factors at once: it compares a 256-token static-resolution model to a 7,290-token dynamic-resolution configuration with a different LLM, so the speedup is not solely attributable to the vision encoder.
  • The benchmark stack (CoreML for encoders, MLX for LLMs, on Apple silicon) may favor hybrid convolutional encoders that convert cleanly; on other hardware with different conversion overheads, the relative gap to ViT and SigLIP could shrink.
  • The finding that a 64x-downsampled self-attention stage is enough for competitive VLM accuracy suggests a design rule that could extend to video or multi-image inputs, where token budgets explode even faster.
  • A direct test of the mechanism would be to pretrain FastViTHD at even higher downsampling (e.g., 128x) or to scale resolution beyond 1024 in the same setup; if accuracy plateaus while latency keeps falling, the current 64x choice is close to optimal.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces FastViTHD, a five-stage hybrid vision encoder with an overall 64x downsampling, and FastVLM, a vision-language model that uses it. The central claim is that FastViTHD provides a better accuracy-latency frontier than ViT, SigLIP, and ConvNeXt encoders in the LLaVA-1.5 setup, with reported speedups of 3.2x vs SigLIP-SO400M, 2.3x vs ConvNeXt, and 85x vs LLaVA-OneVision, while being smaller or comparable in accuracy. The paper also presents an analysis of the interplay between resolution, visual token count, LLM size, and TTFT, and argues that static high-resolution scaling outperforms tile-based dynamic resolution except at extreme resolutions. The authors release code and multiple checkpoints, and report variance over three training runs for key ablations.

Significance. If the efficiency claims hold, this is a useful contribution: it demonstrates that a hierarchical hybrid encoder with aggressive downsampling can substantially reduce visual tokens and encoding latency without sacrificing VLM accuracy, and it provides a systematic empirical mapping of the resolution-LLM-token trade-off. The released checkpoints and code make the accuracy claims checkable, and the variance reporting (Sec. D.3) is a strength. However, the efficiency claims rest entirely on measurements on one Apple M1 Max machine using CoreML and MLX conversions, with latency reported only for models that convert favorably to those frameworks. This limits the generality of the headline speedups and of the Pareto-frontier comparisons, and the 85x figure conflates token-count reduction with encoder efficiency.

major comments (3)
  1. [Sec. 4 (Benchmarking) and Table 10] The latency comparison is restricted to models that are 'publicly available and in a format favorable to MLX' (Sec. 4), and Table 10 marks several baselines (e.g., MM1, ViT-H) as '-' due to export difficulty. This excludes models that convert poorly and biases the Pareto frontier in Fig. 4 and the headline speedups in favor of FastViTHD. Without latency measurements on a neutral platform (e.g., A100 with TensorRT or a standard PyTorch benchmark), or at least a report of failed conversions and their latencies, the central accuracy-latency claim is not established beyond Apple-specific conversion artifacts.
  2. [Sec. 4.1, Table 6 (R2 vs R4)] The 85x TTFT comparison between FastVLM (R4) and LLaVA-OneVision (R2) is dominated by LLM prefilling, not the vision encoder: Table 10 shows prefill of 11,402.4 ms vs 50.5 ms and encoder latency of 2,721.4 ms vs 116.3 ms. The prefill difference is due to 7,290 vs 256 visual tokens, and the claimed 85x ratio therefore conflates token-count reduction with encoder efficiency. The paper should report encoder-only latency separately, or compare at matched token budgets, to support the claim that the architecture, rather than token count, drives the speedup.
  3. [Sec. 3.2.1, Fig. 4] The claim that the Pareto-optimal curve for FastViTHD is 'significantly better' with 'an improvement of over 2.5 points on the Average-5 metric' is presented without the underlying per-configuration table. Since the Avg-5 metric is defined in this paper and the figure uses log-scale axes with overlapping points, the reader cannot verify the 2.5-point improvement quantitatively. Please provide the (resolution, LLM, Avg-5, TTFT) values for all points in Fig. 4, or a table, so the Pareto claim is checkable.
minor comments (6)
  1. [Abstract] The headline '85x faster TTFT' should be qualified as 'on Apple silicon' or 'in our benchmarking setup' to avoid over-generalization to other hardware.
  2. [Sec. 4 (Benchmarking)] The definition of TTFT as vision encoder latency plus LLM prefill excludes the first decode step; clarify that this is a prefill-only metric rather than the standard time to first generated token.
  3. [Table 6 caption] The footnote about 'format favorable to MLX' is easy to miss; consider moving the selection-bias caveat to the main text of Sec. 4.
  4. [Fig. 4] The marker for FastViT at resolution 2048^2 appears to be missing the '2048' label in the figure; please check.
  5. [Table 10] Rows R3 and R3* are not distinguished by an asterisk in the table; fix the formatting.
  6. [Sec. 4.1] The sentence 'FastVLM (R40) outperforms Cambrian-1 (R44) ... while being 7.9x faster' relies on TTFT values from Table 10 that are only available for some models; ensure the 7.9x ratio is computed consistently and note any missing entries.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: FastVLM's claims are empirical measurements on external VLM benchmarks; self-cited FastViT/MobileCLIP are released prior work, and the MLX-only benchmark selection is a validity risk, not a circular derivation.

full rationale

The paper's derivation chain is architectural design (FastViTHD), CLIP-style pretraining, LLaVA-1.5-style visual instruction tuning, and evaluation on external benchmarks (GQA, TextVQA, POPE, DocVQA, SeedBench, MMMU, etc.). No equation in the paper defines an output in terms of a fitted constant or a target metric, and no 'prediction' is constructed from the benchmark scores it claims to predict. The marginal analysis in Sec. 3.2.1, including the claim that 'the Pareto-optimal curve for FastViTHD ... is significantly better than that of FastViT,' is an empirical observation over trained (Resolution, LLM) pairs, not a quantity forced by definition. The self-defined Avg-5 aggregation is an evaluation summary, not an input to the models or to the Pareto analysis; it averages externally evaluated benchmark scores and is used as a reporting convention. Self-citations to FastViT [82] and MobileCLIP [83] are load-bearing only as reusable architecture and pretraining infrastructure, and both are published, code-released prior works, so they constitute independent support rather than an unverified self-citation chain. The paper itself discloses the main benchmarking limitation in Sec. 4: 'we report latency only for models that are publicly available and in a format favorable to MLX [31],' and Appendix Table 10 marks several baselines as '-' because they are 'difficult to export.' This is a genuine external-validity and selection-bias concern that could affect whether the reported 3.2x/2.3x speedups transfer to other hardware or to excluded baselines; similarly, the 85x headline comparison mixes a 256-token static-resolution FastVLM with LLaVA-OneVision's 7,290-token dynamic-resolution configuration, so LLM prefill token count accounts for much of the ratio. These are comparability and generalizability risks, not circularity: they do not make any claimed result equivalent to its own inputs by construction. Overall, the central claims are self-contained empirical measurements, so the circularity score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

This is an empirical systems paper with no derivation whose validity depends on hidden constants. The load-bearing hand-chosen items are the architecture hyperparameters and the evaluation aggregate, and the key domain assumptions are the fairness of the shared training protocol and the Apple-specific latency stack. No new physical or formal entities are postulated.

free parameters (3)
  • FastViTHD stage depths and widths = Depths [2,12,24,4,2]; widths [96,192,384,768,1536]; MLP expansion 4.0
    Hand-chosen scale-up of FastViT (Sec. 3.2, Fig. 2). The claim that FastViTHD improves the Pareto frontier is demonstrated only for this configuration, and Fig. 3 compares it against a single alternative 'naive' scaling.
  • Avg-5 aggregate metric = Equal average of GQA, TextVQA, POPE, DocVQA and SeedBench scores
    Self-defined aggregate on five hand-selected benchmarks is the ranking signal for all Pareto and optimality claims (Sec. 4). Choosing low-variance benchmarks is reasonable, but the aggregate itself is a modeling choice that determines the conclusions.
  • Dynamic-resolution tile grid = 2x2 grid with 1024x1024 tiles
    Hand-chosen configuration for the highest-resolution variants (Sec. C.1); all 2048x2048 results depend on this tile scheme.
assumptions (3)
  • domain assumption LLaVA-1.5 2-stage training with a single epoch per stage and all modules trainable is a fair protocol for every vision encoder compared
    This schedule is used for all ablations (Sec. 4, Tab. 8). If ViT or ConvNeXt backbones converge differently under this fixed budget, the relative accuracy ranks could shift.
  • domain assumption CoreML (vision) and MLX FP16 (LLM) latency on M1 Max is a representative measure of production time-to-first-token
    The benchmarking section states that all models are converted to CoreML or MLX and measured on a MacBook Pro M1 Max; the headline 3.2x and 85x ratios inherit any conversion-quality artifacts.
  • domain assumption CLIP-style pretraining on DataCompDR-1B transfers comparably across encoder architectures for the VLM objective
    FastViTHD is pretrained following the MobileCLIP setup (Sec. 3.2); the comparison assumes each architecture benefits equally from this pretraining before LLaVA-style tuning.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FastVLM: Efficient Vision Encoding for Vision Language Models." pith.science (2026). https://pith.science/paper/Q4FTJ4M2

@misc{pith2026241213303,
  author       = {Pith},
  title        = {Pith review of: FastVLM: Efficient Vision Encoding for Vision Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q4FTJ4M2}},
  note         = {Machine review of arXiv:2412.13303}
}
abstract

Scaling the input image resolution is essential for enhancing the performance of Vision Language Models (VLMs), particularly in text-rich image understanding tasks. However, popular visual encoders such as ViTs become inefficient at high resolutions due to the large number of tokens and high encoding latency caused by stacked self-attention layers. At different operational resolutions, the vision encoder of a VLM can be optimized along two axes: reducing encoding latency and minimizing the number of visual tokens passed to the LLM, thereby lowering overall latency. Based on a comprehensive efficiency analysis of the interplay between image resolution, vision latency, token count, and LLM size, we introduce FastVLM, a model that achieves an optimized trade-off between latency, model size and accuracy. FastVLM incorporates FastViTHD, a novel hybrid vision encoder designed to output fewer tokens and significantly reduce encoding time for high-resolution images. Unlike previous methods, FastVLM achieves the optimal balance between visual token count and image resolution solely by scaling the input image, eliminating the need for additional token pruning and simplifying the model design. In the LLaVA-1.5 setup, FastVLM achieves 3.2$\times$ improvement in time-to-first-token (TTFT) while maintaining similar performance on VLM benchmarks compared to prior works. Compared to LLaVa-OneVision at the highest resolution (1152$\times$1152), FastVLM achieves better performance on key benchmarks like SeedBench, MMMU and DocVQA, using the same 0.5B LLM, but with 85$\times$ faster TTFT and a vision encoder that is 3.4$\times$ smaller. Code and models are available at https://github.com/apple/ml-fastvlm.

Figures

Figures reproduced from arXiv: 2412.13303 by the authors.

Figure 1
Figure 1. FastVLM is more than 3× faster than prior work. Comparison of commonly used vision encoders for VLMs with (a) Qwen2 [86] 0.5B LLM and (b) Vicuna 7B [98] LLM. All the vision encoders are CLIP [69] pretrained. For a fair comparison all models are trained using LLaVA-1.5 [53] setup with the vision encoders made trainable for resolution adaptation, see Sec. 4 for more details. Marker size for each model corresponds to n… view at source ↗
Figure 2
Figure 2. Overview of the FastVLM architecture. FastVLM consists of our novel vision encoder, FastViTHD, trained using the same setup as LLaVA. The FastViTHD architecture is designed for low latency at high resolution, by utilizing additional self-attention layers, and downsampling to generate 4× fewer tokens than FastViT, and 16× fewer tokens than ViT-L/14 at resolution 336. network extract information at different granulari… view at source ↗
Figure 3
Figure 3. Novel scaling strategy of FastViTHD lowers latency at various image resolutions. FastViT-Naive, a naive scaling of the FastViT architecture, and our proposed FastViTHD have the same number of parameters. ConvNeXt-L is provided for ref￾erence. All models are benchmarked on M1 Macbook Pro and trained with LLaVA-1.5 setup and Vicuna 7B. Note that the y-axis is in log scale. 3.2. FastViTHD: High Resolution Encoder for V… view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Vision latency dominates at high resolution. Break￾down of FastVLM’s time to first token for varying image resolu￾tions. Vision encoder is FastViTHD and LLM is Qwen2-1.5B. 102 103 Time To First Token (ms) 60 61 62 63 64 65 Avg-5 VLM Evals (%) 5122 7682 10242 15362 5122…
Figure 6
Figure 6. Figure 6: Dynamic input resolution (AnyRes) is only optimal at the highest resolution when using fewer tiles (2×2). The vision encoder is FastViTHD. The tile grid size is specified in parenthe￾sis. Training setup is LLaVA-1.5 with Vicuna 7B. Note that the x-axis is in log scale.…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MobileCLIP2: Improving Multi-Modal Reinforced Training

    cs.CV 2025-08 conditional novelty 5.0 of 10

    MobileCLIP2 combines DFN-trained teachers, a fine-tuned CoCa captioner, and new 5-stage FastViT variants to set state-of-the-art ImageNet-1k zero-shot accuracy at low latency.

  2. Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    A framework for DP fine-tuning of MLLMs that prunes visual tokens before training and selectively applies noisy gradient updates to blocks with the largest norms, reporting modest utility and memory gains over DP-SGD.

Reference graph

Works this paper leans on

99 extracted references · 35 canonical work pages · cited by 2 Pith papers

  1. [1]

    Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Men- sch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj B...

  2. [2]

    Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els

    Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els. arXiv preprint...

  3. [3]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xi- aodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Day- iheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfe...

  4. [4]

    Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan 9 Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 202k. 2, 7

  5. [5]

    Introducing our multimodal models, 2023

    Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sa ˘gnak Tas ¸ırlar. Introducing our multimodal models, 2023. 2

  6. [6]

    Paligemma: A versatile 3b vlm for trans- fer, 2024

    Lucas Beyer, Andreas Steiner, Andr ´e Susano Pinto, Alexan- der Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisensch- los, Rishabh Kabra, Matthi...

  7. [7]

    Matryoshka multimodal models

    Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee. Matryoshka multimodal models. arXiv preprint arXiv:2405.17430, 2024. 3, 6

  8. [8]

    ”an augmented benchmark dataset for geometric question answering through dual parallel text encoding”

    Jie ”Cao and Jing” Xiao. ”an augmented benchmark dataset for geometric question answering through dual parallel text encoding”. In ”Proceedings of the 29th International Con- ference on Computational Linguistics”, ”2022”. 4

Show all 99 references
  1. [9]

    Honeybee: Locality-enhanced projector for multimodal llm

    Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh. Honeybee: Locality-enhanced projector for multimodal llm. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  2. [10]

    Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

    Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3558–3568, 2021. 3

  3. [11]

    Florence-vl: Enhanc- ing vision-language models with generative vision encoder and depth-breadth fusion

    Jiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li, Jian- feng Gao, Tianyi Zhou, and Bin Xiao. Florence-vl: Enhanc- ing vision-language models with generative vision encoder and depth-breadth fusion. arXiv preprint arXiv:2412.04424,

  4. [12]

    Vitamin: Designing scalable vision models in the vision-language era

    Jieneng Chen, Qihang Yu, Xiaohui Shen, Alan Yuille, and Liang-Chieh Chen. Vitamin: Designing scalable vision models in the vision-language era. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 4, 5, 7, 1, 2

  5. [13]

    Sharegpt4v: Improving large multi-modal models with better captions

    Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023. 7, 2

  6. [14]

    An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

    Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024. 6

  7. [15]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv prepri...

  8. [16]

    How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites

    Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024. 2

  9. [17]

    MOAT: Alternating mobile convolution and attention brings strong vision models

    Chenglin Yang et al. MOAT: Alternating mobile convolution and attention brings strong vision models. In ICLR, 2023. 4

  10. [18]

    Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices

    Xiangxiang Chu, Limeng Qiao, Xinyang Lin, Shuang Xu, Yang Yang, Yiming Hu, Fei Wei, Xinyu Zhang, Bo Zhang, Xiaolin Wei, et al. Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices. arXiv preprint arXiv:2312.16886, 2023. 3, 7, 2

  11. [19]

    Mobilevlm v2: Faster and stronger baseline for vision language model

    Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, et al. Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv preprint arXiv:2402.03766, 2024. 7, 2

  12. [20]

    Instructblip: Towards general- purpose vision-language models with instruction tuning,

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general- purpose vision-language models with instruction tuning,

  13. [21]

    Deepseek llm: Scaling open-source language models with longtermism

    DeepSeek-AI. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954,

  14. [22]

    Smith, Hannaneh Ha- jishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kem- bhavi

    Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tri- pathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison- Bur...

  15. [23]

    Unveiling encoder-free vision-language models

    Haiwen Diao, Yufeng Cui, Xiaotong Li, Yueze Wang, Huchuan Lu, and Xinlong Wang. Unveiling encoder-free vision-language models. arXiv preprint arXiv:2406.11832,

  16. [24]

    An image is worth 16x16 words: Trans- formers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...

  17. [25]

    Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

    Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, 10 Jiaqi Wang, et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Multimedia ,

  18. [26]

    Data fil- tering networks

    Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. Data fil- tering networks. arXiv preprint arXiv:2309.17425, 2023. 3

  19. [27]

    Dat- acomp: In search of the next generation of multimodal datasets

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Dat- acomp: In search of the next generation of multimodal datasets. arXiv preprint arXiv:2304.14108, 2023. 4, 5

  20. [28]

    Con- vllava: Hierarchical backbones as visual encoder for large multimodal models, 2024

    Chunjiang Ge, Sijie Cheng, Ziming Wang, Jiale Yuan, Yuan Gao, Jun Song, Shiji Song, Gao Huang, and Bo Zheng. Con- vllava: Hierarchical backbones as visual encoder for large multimodal models, 2024. 2, 3, 7, 8

  21. [29]

    Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913...

  22. [30]

    Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale

    Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, and Xi- ang Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. arXiv preprint arXiv:2412.05237,

  23. [31]

    MLX: Efficient and flexible machine learn- ing on apple silicon, 2023

    Awni Hannun, Jagrit Digani, Angelos Katharopoulos, and Ronan Collobert. MLX: Efficient and flexible machine learn- ing on apple silicon, 2023. 7, 8, 3

  24. [32]

    Matryoshka query trans- former for large vision-language models, 2024

    Wenbo Hu, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng, and Kai-Wei Chang. Matryoshka query trans- former for large vision-language models, 2024. 3, 6

  25. [33]

    Dynamic-llava: Efficient multimodal large language models via dynamic vision-language context sparsification, 2024

    Wenxuan Huang, Zijie Zhai, Yunhang Shen, Shaoshen Cao, Fei Zhao, Xiangfeng Xu, Zheyu Ye, and Shaohui Lin. Dynamic-llava: Efficient multimodal large language models via dynamic vision-language context sparsification, 2024. 6

  26. [34]

    Gqa: A new dataset for real-world visual reasoning and compositional question answering

    Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 6700–6709, 2019. 8, 5

  27. [35]

    Dvqa: Understanding data visualizations via ques- tion answering

    Kushal Kafle, Scott Cohen, Brian Price, and Christopher Kanan. Dvqa: Understanding data visualizations via ques- tion answering. In CVPR, 2018. 4

  28. [36]

    Prismatic vlms: Investigating the design space of visually-conditioned language models

    Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models. In International Conference on Machine Learning (ICML), 2024. 3

  29. [37]

    A diagram is worth a dozen images, 2016

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images, 2016. 4

  30. [38]

    Ocr-free document understanding transformer

    Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sang- doo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. In European Confer- ence on Computer Vision (ECCV), 2022. 4

  31. [39]

    Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 4

  32. [40]

    Shamma, Michael S

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. ...

  33. [41]

    Building and better understanding vision- language models: insights and future directions., 2024

    Hugo Laurenc ¸on, Andr´es Marafioti, Victor Sanh, and L ´eo Tronchon. Building and better understanding vision- language models: insights and future directions., 2024. 4

  34. [42]

    Seed-bench: Benchmarking mul- timodal llms with generative comprehension

    Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yix- iao Ge, and Ying Shan. Seed-bench: Benchmarking mul- timodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. 8

  35. [43]

    Llava-next: What else influences visual instruction tun- ing beyond data?, 2024

    Bo Li, Hao Zhang, Kaichen Zhang, Dong Guo, Yuanhan Zhang, Renrui Zhang, Feng Li, Ziwei Liu, and Chunyuan Li. Llava-next: What else influences visual instruction tun- ing beyond data?, 2024. 6, 1, 3

  36. [44]

    Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024

    Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li. Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024. 8

  37. [45]

    Llava-onevision: Easy visual task transfer

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 2, 7, 8, 3, 4

  38. [46]

    Flex- attention for efficient high-resolution vision-language mod- els

    Junyan Li, Delin Chen, Tianle Cai, Peihao Chen, Yining Hong, Zhenfang Chen, Yikang Shen, and Chuang Gan. Flex- attention for efficient high-resolution vision-language mod- els. In European Conference on Computer Vision , pages 286–302. Springer, 2025. 7

  39. [47]

    Li, Sachin Goyal, Joao D

    Kevin Y . Li, Sachin Goyal, Joao D. Semedo, and J. Zico Kolter. Inference optimal vlms need only one visual token but larger models, 2024. 4

  40. [48]

    Evaluating object hallucina- tion in large vision-language models

    Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucina- tion in large vision-language models. arXiv preprint arXiv:2305.10355, 2023. 8

  41. [49]

    Mini-gemini: Mining the potential of multi-modality vision language models

    Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814,

  42. [50]

    Vila: On pre-training for visual language models, 2023

    Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models, 2023. 2, 7

  43. [51]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755,

  44. [52]

    Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models

    Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, 11 et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023. 2, 6

  45. [53]

    Improved baselines with visual instruction tuning, 2023

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023. 1, 2, 3, 5, 6, 7, 4

  46. [54]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. 2, 6, 8

  47. [55]

    Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024. 2

  48. [56]

    Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xucheng Yin, Cheng lin Liu, Lianwen Jin, and Xiang Bai. Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024. 4

  49. [57]

    A convnet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2022. 3

  50. [58]

    Deepseek-vl: Towards real-world vision- language understanding, 2024

    Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. Deepseek-vl: Towards real-world vision- language understanding, 2024. 7, 2

  51. [59]

    Learn to explain: Multimodal reasoning via thought chains for science question answering

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 2022. 8, 4

  52. [60]

    Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts. InIn- ternational Conference on Learning Represen...

  53. [61]

    Feast your eyes: Mixture-of- resolution adaptation for multimodal large language models

    Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xi- aoshuai Sun, and Rongrong Ji. Feast your eyes: Mixture-of- resolution adaptation for multimodal large language models. arXiv preprint arXiv:2403.03003, 2024. 2

  54. [62]

    Smolvlm: Redefining small and efficient multimodal models

    Andr ´es Marafioti, Orr Zohar, Miquel Farr ´e, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Vaibhav Srivastav, Joshua Lochner, Hugo Larcher, Mathieu Morlon, Lewis Tun- stall, Leandro von Werra, and Thomas Wolf. Smolvlm: ...

  55. [63]

    ChartQA: A benchmark for question answering about charts with visual and logical reasoning

    Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Ena- mul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In ”Find- ings of the Association for Computational Linguistics: ACL 2022”, ”2022”. 4, 5, 6

  56. [64]

    V Jawahar

    Minesh Mathew, Viraj Bagal, Rub `en P´erez Tito, Dimosthe- nis Karatzas, Ernest Valveny, and C. V Jawahar. Infograph- icvqa, 2021. 4

  57. [65]

    Docvqa: A dataset for vqa on document images

    Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceed- ings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021. 8, 4, 6

  58. [66]

    Mm1: Methods, analysis & insights from multimodal llm pre- training, 2024

    Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xi- anzhi Du, Futang Peng, Floris Weers, Anton Belyi, Haotian Zhang, Karanjeet Singh, Doug Kang, Ankur Jain, Hongyu H`e, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang...

  59. [67]

    Ocr-vqa: Visual question answering by reading text in images

    Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In ICDAR, 2019. 4

  60. [68]

    Gpt-4 technical report

    OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 2

  61. [69]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  62. [70]

    Llava-prumerge: Adaptive token reduction for efficient large multimodal models

    Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388,

  63. [71]

    Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

    Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In Pro- ceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...

  64. [72]

    When do we not need larger vision models? In European Conference on Computer Vision (ECCV), 2024

    Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, and Trevor Darrell. When do we not need larger vision models? In European Conference on Computer Vision (ECCV), 2024. 2

  65. [73]

    Eagle: Ex- ploring the design space for multimodal llms with mixture of encoders

    Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, and Guilin Liu. Eagle: Ex- ploring the design space for multimodal llms with mixture o...

  66. [74]

    Towards vqa models that can read

    Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019. 8, 4

  67. [75]

    Eva-clip: Improved training techniques for clip at scale

    Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 3

  68. [76]

    Chameleon: Mixed-modal early-fusion foundation models, 2024

    Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models, 2024. 2

  69. [77]

    Qwen2.5: A party of foundation models, 2024

    Qwen Team. Qwen2.5: A party of foundation models, 2024. 2, 8

  70. [78]

    Cambrian-1: A fully open, vision-centric exploration of multimodal llms,

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, 12 Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms,

  71. [79]

    Llama: Open and efficient foundation lan- guage models, 2023

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aure- lien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. Llama: Open and efficient foundation la...

  72. [80]

    Multimodal few-shot learning with frozen language models

    Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Es- lami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Conference on Neu- ral Information Processing Systems (NeurIPS), 2021. 2

  73. [81]

    Mobileone: An im- proved one millisecond mobile backbone

    Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan. Mobileone: An im- proved one millisecond mobile backbone. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 1

  74. [82]

    Fastvit: A fast hybrid vision transformer using structural reparameterization

    Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, On- cel Tuzel, and Anurag Ranjan. Fastvit: A fast hybrid vision transformer using structural reparameterization. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision (ICCV), 2023. 2, 3, 4, 1

  75. [83]

    Mobile- clip: Fast image-text models through multi-modal reinforced training

    Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, and Oncel Tuzel. Mobile- clip: Fast image-text models through multi-modal reinforced training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. ...

  76. [84]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017. 2

  77. [85]

    Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S. Ryoo, Shrikant Kendre, Jieyu Zhang, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin ...

  78. [86]

    Qwen2 technical report

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jian- wei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzhe...

  79. [87]

    Visionzip: Longer is better but not necessary in vision language models, 2024

    Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia. Visionzip: Longer is better but not necessary in vision language models, 2024. 6

  80. [88]

    mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024

    Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024. 2

  81. [89]

    mplug-owl: Modularization empowers large lan- guage models with multimodality, 2023

    Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chaoya Jiang, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. mplug-owl: Modularization empowers large lan- guage models...

  82. [90]

    mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration, 2023

    Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration, 2023. 2

  83. [91]

    Mm-vet: Evaluating large multimodal models for integrated capabilities

    Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023. 8

  84. [92]

    Metaformer baselines for vision

    Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang. Metaformer baselines for vision. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 4

  85. [93]

    Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Ren- liang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A...

  86. [94]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. International Conference on Computer Vision (ICCV), 2023. 2, 3, 8

  87. [95]

    Mm1.5: Methods, analysis & insights from multimodal llm fine-tuning, 2024

    Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean- Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, Zirui Wang, Afs...

  88. [96]

    Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

    Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Lmms- eval: Reality check on the evaluation of large multimodal models, 2024. 8, 9

  89. [97]

    Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference

    Yuan Zhang, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang, Kuan Cheng, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, et al. Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference. arXiv preprint arXiv:2410.04417, 2024. 6 13

  90. [98]

    P Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonza- lez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. 1, 2, 3, 6, 7

  91. [99]

    Vic.” refers to Vicuna [98], “Qw.2

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 2 14 FastVLM: Efficient Vision Encoding for Vision Language Models Supplementar...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.