REVIEW 3 major objections 6 minor 2 cited by
FastVLM: Efficient Vision Encoding for Vision Language Models
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A single hybrid vision encoder, FastViTHD, that downsamples images 64x before self-attention produces far fewer visual tokens and encodes high-resolution images several times faster than ViT, SigLIP, and ConvNeXt encoders, while matching…
desk verdict FastViTHD is a genuinely useful token-efficient encoder with a real Pareto story, but the headline latency ratios rest on Apple-only, selectively reported benchmarks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
FastViTHD is a five-stage hybrid encoder: the first three stages use RepMixer convolutional blocks, and the last two stages use self-attention, with an added patch-embedding layer that downsamples the input by a total factor of 64. This makes self-attention operate on a small tensor (16x16 for a 1024x1024 image), cutting both encoder latency and the number of tokens passed to the LLM, which is what reduces LLM prefilling time and thus time-to-first-token. Multi-scale features pooled from earlier stages with depthwise convolutions add a small accuracy boost. The architecture is pretrained with MobileCLIP's reinforced image-text pipeline and then fine-tuned end-to-end in the LLaVA-1.5 two-stage recipe.
What would settle it
Run the exact same controlled comparison (same LLaVA-1.5 training, same LLM) on a different hardware platform, e.g., an NVIDIA GPU with standard PyTorch or TensorRT, benchmarking all encoders including those excluded from the MLX comparison, and compare TTFT at matched token counts; if FastViTHD's time-to-first-token advantage falls below roughly 2x, the central claim fails.
Extended reading notes
Core claim
The paper's central claim is that a hybrid vision backbone with a 64x downsampled final stage dominates isotropic ViTs and pure convolutional encoders on the accuracy-latency frontier for vision-language models. FastViTHD runs self-attention only on a heavily downsampled feature map, so it encodes a 1024x1024 image in 235 ms on an M1 MacBook Pro while producing only 256 visual tokens, versus ViT-L/14's 576 tokens from a 336x336 image. Across LLMs of 0.5B, 1.5B, and 7B parameters, the Pareto-optimal curve of FastViTHD is over 2.5 points better on the Average-5 metric than the best FastViT curve, and it reaches a target VLM performance up to 3x faster. The paper further shows that scaling input resolution directly beats tiling (AnyRes) except at extreme resolutions, and that a hierarchical encoder with few tokens beats token-pruning methods applied to ViTs.
Load-bearing premise
The claimed speedups are measured only on Apple-silicon devices using CoreML for encoders and MLX for the LLM, and only for models that convert cleanly to those formats, so the latency comparisons could be artifacts of the conversion stack rather than of the architectures themselves.
Editorial extensions
If this is right
- High-resolution VLMs for text-rich images could run at roughly one-third the time-to-first-token of prior ViT-based models, making on-device document and chart understanding practical.
- The need for token-pruning and resampling modules disappears for hierarchical encoders: simply training at lower input resolution yields token counts as low as 16 while outperforming prune-then-feed ViT methods.
- Because the encoder produces far fewer tokens, smaller LLMs (e.g., 0.5B) can handle high-resolution inputs better, and the Pareto analysis shows which (resolution, LLM size) pair is optimal for a given latency budget.
- Scaling visual instruction-tuning data further improves FastVLM, suggesting that the efficient token representation transfers to larger datasets and stronger benchmarks.
- Dynamic-resolution tiling is largely unnecessary with FastViTHD: static resolution scaling is the better accuracy-latency trade-off except at extreme resolutions like 1536x1536.
Reading between the lines
- The headline 85x TTFT improvement versus LLaVA-OneVision mixes several factors at once: it compares a 256-token static-resolution model to a 7,290-token dynamic-resolution configuration with a different LLM, so the speedup is not solely attributable to the vision encoder.
- The benchmark stack (CoreML for encoders, MLX for LLMs, on Apple silicon) may favor hybrid convolutional encoders that convert cleanly; on other hardware with different conversion overheads, the relative gap to ViT and SigLIP could shrink.
- The finding that a 64x-downsampled self-attention stage is enough for competitive VLM accuracy suggests a design rule that could extend to video or multi-image inputs, where token budgets explode even faster.
- A direct test of the mechanism would be to pretrain FastViTHD at even higher downsampling (e.g., 128x) or to scale resolution beyond 1024 in the same setup; if accuracy plateaus while latency keeps falling, the current 64x choice is close to optimal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FastViTHD, a five-stage hybrid vision encoder with an overall 64x downsampling, and FastVLM, a vision-language model that uses it. The central claim is that FastViTHD provides a better accuracy-latency frontier than ViT, SigLIP, and ConvNeXt encoders in the LLaVA-1.5 setup, with reported speedups of 3.2x vs SigLIP-SO400M, 2.3x vs ConvNeXt, and 85x vs LLaVA-OneVision, while being smaller or comparable in accuracy. The paper also presents an analysis of the interplay between resolution, visual token count, LLM size, and TTFT, and argues that static high-resolution scaling outperforms tile-based dynamic resolution except at extreme resolutions. The authors release code and multiple checkpoints, and report variance over three training runs for key ablations.
Significance. If the efficiency claims hold, this is a useful contribution: it demonstrates that a hierarchical hybrid encoder with aggressive downsampling can substantially reduce visual tokens and encoding latency without sacrificing VLM accuracy, and it provides a systematic empirical mapping of the resolution-LLM-token trade-off. The released checkpoints and code make the accuracy claims checkable, and the variance reporting (Sec. D.3) is a strength. However, the efficiency claims rest entirely on measurements on one Apple M1 Max machine using CoreML and MLX conversions, with latency reported only for models that convert favorably to those frameworks. This limits the generality of the headline speedups and of the Pareto-frontier comparisons, and the 85x figure conflates token-count reduction with encoder efficiency.
major comments (3)
- [Sec. 4 (Benchmarking) and Table 10] The latency comparison is restricted to models that are 'publicly available and in a format favorable to MLX' (Sec. 4), and Table 10 marks several baselines (e.g., MM1, ViT-H) as '-' due to export difficulty. This excludes models that convert poorly and biases the Pareto frontier in Fig. 4 and the headline speedups in favor of FastViTHD. Without latency measurements on a neutral platform (e.g., A100 with TensorRT or a standard PyTorch benchmark), or at least a report of failed conversions and their latencies, the central accuracy-latency claim is not established beyond Apple-specific conversion artifacts.
- [Sec. 4.1, Table 6 (R2 vs R4)] The 85x TTFT comparison between FastVLM (R4) and LLaVA-OneVision (R2) is dominated by LLM prefilling, not the vision encoder: Table 10 shows prefill of 11,402.4 ms vs 50.5 ms and encoder latency of 2,721.4 ms vs 116.3 ms. The prefill difference is due to 7,290 vs 256 visual tokens, and the claimed 85x ratio therefore conflates token-count reduction with encoder efficiency. The paper should report encoder-only latency separately, or compare at matched token budgets, to support the claim that the architecture, rather than token count, drives the speedup.
- [Sec. 3.2.1, Fig. 4] The claim that the Pareto-optimal curve for FastViTHD is 'significantly better' with 'an improvement of over 2.5 points on the Average-5 metric' is presented without the underlying per-configuration table. Since the Avg-5 metric is defined in this paper and the figure uses log-scale axes with overlapping points, the reader cannot verify the 2.5-point improvement quantitatively. Please provide the (resolution, LLM, Avg-5, TTFT) values for all points in Fig. 4, or a table, so the Pareto claim is checkable.
minor comments (6)
- [Abstract] The headline '85x faster TTFT' should be qualified as 'on Apple silicon' or 'in our benchmarking setup' to avoid over-generalization to other hardware.
- [Sec. 4 (Benchmarking)] The definition of TTFT as vision encoder latency plus LLM prefill excludes the first decode step; clarify that this is a prefill-only metric rather than the standard time to first generated token.
- [Table 6 caption] The footnote about 'format favorable to MLX' is easy to miss; consider moving the selection-bias caveat to the main text of Sec. 4.
- [Fig. 4] The marker for FastViT at resolution 2048^2 appears to be missing the '2048' label in the figure; please check.
- [Table 10] Rows R3 and R3* are not distinguished by an asterisk in the table; fix the formatting.
- [Sec. 4.1] The sentence 'FastVLM (R40) outperforms Cambrian-1 (R44) ... while being 7.9x faster' relies on TTFT values from Table 10 that are only available for some models; ensure the 7.9x ratio is computed consistently and note any missing entries.
Circularity Check
No circularity: FastVLM's claims are empirical measurements on external VLM benchmarks; self-cited FastViT/MobileCLIP are released prior work, and the MLX-only benchmark selection is a validity risk, not a circular derivation.
full rationale
The paper's derivation chain is architectural design (FastViTHD), CLIP-style pretraining, LLaVA-1.5-style visual instruction tuning, and evaluation on external benchmarks (GQA, TextVQA, POPE, DocVQA, SeedBench, MMMU, etc.). No equation in the paper defines an output in terms of a fitted constant or a target metric, and no 'prediction' is constructed from the benchmark scores it claims to predict. The marginal analysis in Sec. 3.2.1, including the claim that 'the Pareto-optimal curve for FastViTHD ... is significantly better than that of FastViT,' is an empirical observation over trained (Resolution, LLM) pairs, not a quantity forced by definition. The self-defined Avg-5 aggregation is an evaluation summary, not an input to the models or to the Pareto analysis; it averages externally evaluated benchmark scores and is used as a reporting convention. Self-citations to FastViT [82] and MobileCLIP [83] are load-bearing only as reusable architecture and pretraining infrastructure, and both are published, code-released prior works, so they constitute independent support rather than an unverified self-citation chain. The paper itself discloses the main benchmarking limitation in Sec. 4: 'we report latency only for models that are publicly available and in a format favorable to MLX [31],' and Appendix Table 10 marks several baselines as '-' because they are 'difficult to export.' This is a genuine external-validity and selection-bias concern that could affect whether the reported 3.2x/2.3x speedups transfer to other hardware or to excluded baselines; similarly, the 85x headline comparison mixes a 256-token static-resolution FastVLM with LLaVA-OneVision's 7,290-token dynamic-resolution configuration, so LLM prefill token count accounts for much of the ratio. These are comparability and generalizability risks, not circularity: they do not make any claimed result equivalent to its own inputs by construction. Overall, the central claims are self-contained empirical measurements, so the circularity score is 0.
Assumptions & free parameters
free parameters (3)
- FastViTHD stage depths and widths =
Depths [2,12,24,4,2]; widths [96,192,384,768,1536]; MLP expansion 4.0
- Avg-5 aggregate metric =
Equal average of GQA, TextVQA, POPE, DocVQA and SeedBench scores
- Dynamic-resolution tile grid =
2x2 grid with 1024x1024 tiles
assumptions (3)
- domain assumption LLaVA-1.5 2-stage training with a single epoch per stage and all modules trainable is a fair protocol for every vision encoder compared
- domain assumption CoreML (vision) and MLX FP16 (LLM) latency on M1 Max is a representative measure of production time-to-first-token
- domain assumption CLIP-style pretraining on DataCompDR-1B transfers comparably across encoder architectures for the VLM objective
Cite this review
Pith. "Pith review of FastVLM: Efficient Vision Encoding for Vision Language Models." pith.science (2026). https://pith.science/paper/Q4FTJ4M2
@misc{pith2026241213303,
author = {Pith},
title = {Pith review of: FastVLM: Efficient Vision Encoding for Vision Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q4FTJ4M2}},
note = {Machine review of arXiv:2412.13303}
}
abstract
Scaling the input image resolution is essential for enhancing the performance of Vision Language Models (VLMs), particularly in text-rich image understanding tasks. However, popular visual encoders such as ViTs become inefficient at high resolutions due to the large number of tokens and high encoding latency caused by stacked self-attention layers. At different operational resolutions, the vision encoder of a VLM can be optimized along two axes: reducing encoding latency and minimizing the number of visual tokens passed to the LLM, thereby lowering overall latency. Based on a comprehensive efficiency analysis of the interplay between image resolution, vision latency, token count, and LLM size, we introduce FastVLM, a model that achieves an optimized trade-off between latency, model size and accuracy. FastVLM incorporates FastViTHD, a novel hybrid vision encoder designed to output fewer tokens and significantly reduce encoding time for high-resolution images. Unlike previous methods, FastVLM achieves the optimal balance between visual token count and image resolution solely by scaling the input image, eliminating the need for additional token pruning and simplifying the model design. In the LLaVA-1.5 setup, FastVLM achieves 3.2$\times$ improvement in time-to-first-token (TTFT) while maintaining similar performance on VLM benchmarks compared to prior works. Compared to LLaVa-OneVision at the highest resolution (1152$\times$1152), FastVLM achieves better performance on key benchmarks like SeedBench, MMMU and DocVQA, using the same 0.5B LLM, but with 85$\times$ faster TTFT and a vision encoder that is 3.4$\times$ smaller. Code and models are available at https://github.com/apple/ml-fastvlm.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 2 Pith papers
-
MobileCLIP2: Improving Multi-Modal Reinforced Training
MobileCLIP2 combines DFN-trained teachers, a fine-tuned CoCa captioner, and new 5-stage FastViT variants to set state-of-the-art ImageNet-1k zero-shot accuracy at low latency.
-
Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models
A framework for DP fine-tuning of MLLMs that prunes visual tokens before training and selectively applies noisy gradient updates to blocks with the largest norms, reporting modest utility and memory gains over DP-SGD.
Reference graph
Works this paper leans on
-
[1]
Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar ´en Simonyan
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Men- sch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj B...
2022
-
[2]
Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt. Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els. arXiv preprint...
arXiv 2023
-
[3]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xi- aodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Day- iheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfe...
arXiv 2023
-
[4]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan 9 Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 202k. 2, 7
-
[5]
Introducing our multimodal models, 2023
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, and Sa ˘gnak Tas ¸ırlar. Introducing our multimodal models, 2023. 2
2023
-
[6]
Paligemma: A versatile 3b vlm for trans- fer, 2024
Lucas Beyer, Andreas Steiner, Andr ´e Susano Pinto, Alexan- der Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisensch- los, Rishabh Kabra, Matthi...
2024
-
[7]
Mu Cai, Jianwei Yang, Jianfeng Gao, and Yong Jae Lee. Matryoshka multimodal models. arXiv preprint arXiv:2405.17430, 2024. 3, 6
arXiv 2024
-
[8]
”an augmented benchmark dataset for geometric question answering through dual parallel text encoding”
Jie ”Cao and Jing” Xiao. ”an augmented benchmark dataset for geometric question answering through dual parallel text encoding”. In ”Proceedings of the 29th International Con- ference on Computational Linguistics”, ”2022”. 4
2022
Show all 99 references
-
[9]
Honeybee: Locality-enhanced projector for multimodal llm
Junbum Cha, Wooyoung Kang, Jonghwan Mun, and Byungseok Roh. Honeybee: Locality-enhanced projector for multimodal llm. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[10]
Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3558–3568, 2021. 3
2021
-
[11]
Florence-vl: Enhanc- ing vision-language models with generative vision encoder and depth-breadth fusion
Jiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li, Jian- feng Gao, Tianyi Zhou, and Bin Xiao. Florence-vl: Enhanc- ing vision-language models with generative vision encoder and depth-breadth fusion. arXiv preprint arXiv:2412.04424,
-
[12]
Vitamin: Designing scalable vision models in the vision-language era
Jieneng Chen, Qihang Yu, Xiaohui Shen, Alan Yuille, and Liang-Chieh Chen. Vitamin: Designing scalable vision models in the vision-language era. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 4, 5, 7, 1, 2
2024
-
[13]
Sharegpt4v: Improving large multi-modal models with better captions
Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023. 7, 2
2023 arXiv
-
[14]
An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024
Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024. 6
2024
-
[15]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv prepri...
2023 arXiv
-
[16]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024. 2
2024 arXiv
-
[17]
MOAT: Alternating mobile convolution and attention brings strong vision models
Chenglin Yang et al. MOAT: Alternating mobile convolution and attention brings strong vision models. In ICLR, 2023. 4
2023
-
[18]
Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices
Xiangxiang Chu, Limeng Qiao, Xinyang Lin, Shuang Xu, Yang Yang, Yiming Hu, Fei Wei, Xinyu Zhang, Bo Zhang, Xiaolin Wei, et al. Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices. arXiv preprint arXiv:2312.16886, 2023. 3, 7, 2
2023 arXiv
-
[19]
Mobilevlm v2: Faster and stronger baseline for vision language model
Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, Shuang Xu, Fei Wei, Yang Yang, Xiaofei Sun, Yiming Hu, Xinyang Lin, Bo Zhang, et al. Mobilevlm v2: Faster and stronger baseline for vision language model. arXiv preprint arXiv:2402.03766, 2024. 7, 2
2024 arXiv
-
[20]
Instructblip: Towards general- purpose vision-language models with instruction tuning,
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general- purpose vision-language models with instruction tuning,
-
[21]
Deepseek llm: Scaling open-source language models with longtermism
DeepSeek-AI. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954,
-
[22]
Smith, Hannaneh Ha- jishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kem- bhavi
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tri- pathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison- Bur...
2024 arXiv
-
[23]
Unveiling encoder-free vision-language models
Haiwen Diao, Yufeng Cui, Xiaotong Li, Yueze Wang, Huchuan Lu, and Xinlong Wang. Unveiling encoder-free vision-language models. arXiv preprint arXiv:2406.11832,
-
[24]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...
2010 arXiv
-
[25]
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, 10 Jiaqi Wang, et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Multimedia ,
-
[26]
Data fil- tering networks
Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. Data fil- tering networks. arXiv preprint arXiv:2309.17425, 2023. 3
2023 arXiv
-
[27]
Dat- acomp: In search of the next generation of multimodal datasets
Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Dat- acomp: In search of the next generation of multimodal datasets. arXiv preprint arXiv:2304.14108, 2023. 4, 5
2023 arXiv
-
[28]
Con- vllava: Hierarchical backbones as visual encoder for large multimodal models, 2024
Chunjiang Ge, Sijie Cheng, Ziming Wang, Jiale Yuan, Yuan Gao, Jun Song, Shiji Song, Gao Huang, and Bo Zheng. Con- vllava: Hierarchical backbones as visual encoder for large multimodal models, 2024. 2, 3, 7, 8
2024
-
[29]
Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Ba- tra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913...
2017
-
[30]
Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale
Jarvis Guo, Tuney Zheng, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Yizhi Li, Graham Neubig, Wenhu Chen, and Xi- ang Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. arXiv preprint arXiv:2412.05237,
-
[31]
MLX: Efficient and flexible machine learn- ing on apple silicon, 2023
Awni Hannun, Jagrit Digani, Angelos Katharopoulos, and Ronan Collobert. MLX: Efficient and flexible machine learn- ing on apple silicon, 2023. 7, 8, 3
2023
-
[32]
Matryoshka query trans- former for large vision-language models, 2024
Wenbo Hu, Zi-Yi Dou, Liunian Harold Li, Amita Kamath, Nanyun Peng, and Kai-Wei Chang. Matryoshka query trans- former for large vision-language models, 2024. 3, 6
2024
-
[33]
Dynamic-llava: Efficient multimodal large language models via dynamic vision-language context sparsification, 2024
Wenxuan Huang, Zijie Zhai, Yunhang Shen, Shaoshen Cao, Fei Zhao, Xiangfeng Xu, Zheyu Ye, and Shaohui Lin. Dynamic-llava: Efficient multimodal large language models via dynamic vision-language context sparsification, 2024. 6
2024
-
[34]
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 6700–6709, 2019. 8, 5
2019
-
[35]
Dvqa: Understanding data visualizations via ques- tion answering
Kushal Kafle, Scott Cohen, Brian Price, and Christopher Kanan. Dvqa: Understanding data visualizations via ques- tion answering. In CVPR, 2018. 4
2018
-
[36]
Prismatic vlms: Investigating the design space of visually-conditioned language models
Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models. In International Conference on Machine Learning (ICML), 2024. 3
2024
-
[37]
A diagram is worth a dozen images, 2016
Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images, 2016. 4
2016
-
[38]
Ocr-free document understanding transformer
Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sang- doo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. In European Confer- ence on Computer Vision (ECCV), 2022. 4
2022
-
[39]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 4
2023 arXiv
-
[40]
Shamma, Michael S
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalan- tidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. ...
2017
-
[41]
Building and better understanding vision- language models: insights and future directions., 2024
Hugo Laurenc ¸on, Andr´es Marafioti, Victor Sanh, and L ´eo Tronchon. Building and better understanding vision- language models: insights and future directions., 2024. 4
2024
-
[42]
Seed-bench: Benchmarking mul- timodal llms with generative comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yix- iao Ge, and Ying Shan. Seed-bench: Benchmarking mul- timodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. 8
2023 arXiv
-
[43]
Llava-next: What else influences visual instruction tun- ing beyond data?, 2024
Bo Li, Hao Zhang, Kaichen Zhang, Dong Guo, Yuanhan Zhang, Renrui Zhang, Feng Li, Ziwei Liu, and Chunyuan Li. Llava-next: What else influences visual instruction tun- ing beyond data?, 2024. 6, 1, 3
2024
-
[44]
Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024
Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li. Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024. 8
2024
-
[45]
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 2, 7, 8, 3, 4
2024 arXiv
-
[46]
Flex- attention for efficient high-resolution vision-language mod- els
Junyan Li, Delin Chen, Tianle Cai, Peihao Chen, Yining Hong, Zhenfang Chen, Yikang Shen, and Chuang Gan. Flex- attention for efficient high-resolution vision-language mod- els. In European Conference on Computer Vision , pages 286–302. Springer, 2025. 7
2025
-
[47]
Li, Sachin Goyal, Joao D
Kevin Y . Li, Sachin Goyal, Joao D. Semedo, and J. Zico Kolter. Inference optimal vlms need only one visual token but larger models, 2024. 4
2024
-
[48]
Evaluating object hallucina- tion in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucina- tion in large vision-language models. arXiv preprint arXiv:2305.10355, 2023. 8
2023 arXiv
-
[49]
Mini-gemini: Mining the potential of multi-modality vision language models
Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814,
-
[50]
Vila: On pre-training for visual language models, 2023
Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models, 2023. 2, 7
2023
-
[51]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, pages 740–755,
-
[52]
Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models
Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, 11 et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023. 2, 6
2023 arXiv
-
[53]
Improved baselines with visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023. 1, 2, 3, 5, 6, 7, 4
2023
-
[54]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAdvances in Neural Information Processing Systems (NeurIPS), 2023. 2, 6, 8
2023
-
[55]
Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024. 2
2024
-
[56]
Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xucheng Yin, Cheng lin Liu, Lianwen Jin, and Xiang Bai. Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024. 4
2024
-
[57]
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feicht- enhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2022. 3
2022
-
[58]
Deepseek-vl: Towards real-world vision- language understanding, 2024
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. Deepseek-vl: Towards real-world vision- language understanding, 2024. 7, 2
2024
-
[59]
Learn to explain: Multimodal reasoning via thought chains for science question answering
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 2022. 8, 4
2022
-
[60]
Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts
Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts. InIn- ternational Conference on Learning Represen...
-
[61]
Feast your eyes: Mixture-of- resolution adaptation for multimodal large language models
Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xi- aoshuai Sun, and Rongrong Ji. Feast your eyes: Mixture-of- resolution adaptation for multimodal large language models. arXiv preprint arXiv:2403.03003, 2024. 2
2024 arXiv
-
[62]
Smolvlm: Redefining small and efficient multimodal models
Andr ´es Marafioti, Orr Zohar, Miquel Farr ´e, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Vaibhav Srivastav, Joshua Lochner, Hugo Larcher, Mathieu Morlon, Lewis Tun- stall, Leandro von Werra, and Thomas Wolf. Smolvlm: ...
2025 arXiv
-
[63]
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Ena- mul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In ”Find- ings of the Association for Computational Linguistics: ACL 2022”, ”2022”. 4, 5, 6
2022
-
[64]
V Jawahar
Minesh Mathew, Viraj Bagal, Rub `en P´erez Tito, Dimosthe- nis Karatzas, Ernest Valveny, and C. V Jawahar. Infograph- icvqa, 2021. 4
2021
-
[65]
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceed- ings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021. 8, 4, 6
2021
-
[66]
Mm1: Methods, analysis & insights from multimodal llm pre- training, 2024
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xi- anzhi Du, Futang Peng, Floris Weers, Anton Belyi, Haotian Zhang, Karanjeet Singh, Doug Kang, Ankur Jain, Hongyu H`e, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang...
2024
-
[67]
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. Ocr-vqa: Visual question answering by reading text in images. In ICDAR, 2019. 4
2019
-
[68]
Gpt-4 technical report
OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 2
2023 arXiv
-
[69]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[70]
Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388,
-
[71]
Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning. In Pro- ceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Paper...
2018
-
[72]
When do we not need larger vision models? In European Conference on Computer Vision (ECCV), 2024
Baifeng Shi, Ziyang Wu, Maolin Mao, Xin Wang, and Trevor Darrell. When do we not need larger vision models? In European Conference on Computer Vision (ECCV), 2024. 2
2024
-
[73]
Eagle: Ex- ploring the design space for multimodal llms with mixture of encoders
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, and Guilin Liu. Eagle: Ex- ploring the design space for multimodal llms with mixture o...
2024 arXiv
-
[74]
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8317–8326, 2019. 8, 4
2019
-
[75]
Eva-clip: Improved training techniques for clip at scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023. 3
2023 arXiv
-
[76]
Chameleon: Mixed-modal early-fusion foundation models, 2024
Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models, 2024. 2
2024
-
[77]
Qwen2.5: A party of foundation models, 2024
Qwen Team. Qwen2.5: A party of foundation models, 2024. 2, 8
2024
-
[78]
Cambrian-1: A fully open, vision-centric exploration of multimodal llms,
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, 12 Rob Fergus, Yann LeCun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms,
-
[79]
Llama: Open and efficient foundation lan- guage models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aure- lien Rodriguez, Armand Joulin, Edouard Grave, and Guil- laume Lample. Llama: Open and efficient foundation la...
2023
-
[80]
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Es- lami, Oriol Vinyals, and Felix Hill. Multimodal few-shot learning with frozen language models. Conference on Neu- ral Information Processing Systems (NeurIPS), 2021. 2
2021
-
[81]
Mobileone: An im- proved one millisecond mobile backbone
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel, and Anurag Ranjan. Mobileone: An im- proved one millisecond mobile backbone. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 1
2023
-
[82]
Fastvit: A fast hybrid vision transformer using structural reparameterization
Pavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, On- cel Tuzel, and Anurag Ranjan. Fastvit: A fast hybrid vision transformer using structural reparameterization. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision (ICCV), 2023. 2, 3, 4, 1
2023
-
[83]
Mobile- clip: Fast image-text models through multi-modal reinforced training
Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, and Oncel Tuzel. Mobile- clip: Fast image-text models through multi-modal reinforced training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. ...
2024
-
[84]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017. 2
2017
-
[85]
Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S. Ryoo, Shrikant Kendre, Jieyu Zhang, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin ...
2024
-
[86]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jian- wei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzhe...
2024 arXiv
-
[87]
Visionzip: Longer is better but not necessary in vision language models, 2024
Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang, Jingyao Li, Bei Yu, and Jiaya Jia. Visionzip: Longer is better but not necessary in vision language models, 2024. 6
2024
-
[88]
mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024
Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024. 2
2024
-
[89]
mplug-owl: Modularization empowers large lan- guage models with multimodality, 2023
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chaoya Jiang, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. mplug-owl: Modularization empowers large lan- guage models...
2023
-
[90]
mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration, 2023
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl2: Revolutionizing multi-modal large lan- guage model with modality collaboration, 2023. 2
2023
-
[91]
Mm-vet: Evaluating large multimodal models for integrated capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023. 8
2023 arXiv
-
[92]
Metaformer baselines for vision
Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, and Xinchao Wang. Metaformer baselines for vision. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 4
2024
-
[93]
Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi
Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Ren- liang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A...
2024
-
[94]
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. International Conference on Computer Vision (ICCV), 2023. 2, 3, 8
2023
-
[95]
Mm1.5: Methods, analysis & insights from multimodal llm fine-tuning, 2024
Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean- Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, Zirui Wang, Afs...
2024
-
[96]
Lmms- eval: Reality check on the evaluation of large multimodal models, 2024
Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Lmms- eval: Reality check on the evaluation of large multimodal models, 2024. 8, 9
2024
-
[97]
Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference
Yuan Zhang, Chun-Kai Fan, Junpeng Ma, Wenzhao Zheng, Tao Huang, Kuan Cheng, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer, et al. Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference. arXiv preprint arXiv:2410.04417, 2024. 6 13
-
[98]
P Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonza- lez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. 1, 2, 3, 6, 7
2023
-
[99]
Vic.” refers to Vicuna [98], “Qw.2
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mo- hamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 2 14 FastVLM: Efficient Vision Encoding for Vision Language Models Supplementar...
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.