REVIEW 4 major objections 8 minor 137 references
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
T0 review · 4 major / 8 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read With 0.5B training pairs and 1.8B activated parameters, a single-network design matches a modular 2B MLLM and beats 8B monoliths.
desk verdict Solid efficiency recipe for monolithic MLLMs, but the 'leading modular MLLMs' claim is overstated and should be corrected. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a multimodal mixture-of-experts in which visual tokens and textual tokens are routed to separate expert networks inside every Transformer layer: new visual FFN experts in the feed-forward blocks and, in Mono-InternVL-1.5, visual attention experts in the query, key, and value projections. A hard routing rule sends each token only to its modality's expert, so the original language pathway stays mathematically untouched during delta tuning. Around this architecture the paper wraps Endogenous Visual Pre-training (EViP and its improved EViP++), a coarse-to-fine schedule of concept learning on 250M noisy image-text pairs, semantic learning on 150M synthetic captions, then alignment learning on 150M task-related samples, and finally instruction tuning. A fused CUDA kernel completes the design by executing the two expert branches jointly, recovering GPU parallelism lost to modality separation.
What would settle it
Train Mono-InternVL-1.5's reduced recipe (250M concept plus 150M semantic pairs) but replace the LLaVA-665k instruction-tuning set with a held-out recipe such as the full InternVL-1.5 instruction mixture, then compare against the 1.1B-data model on tasks outside the paper's saturation-benchmark list; if the smaller model loses ground on those tasks, the saturation-based data-efficiency claim fails.
Extended reading notes
Core claim
The paper's central claim is that monolithic MLLMs can reach performance comparable to leading modular MLLMs and superior efficiency, provided vision is added as a separate, frozen-out parameter space inside the LLM. Concretely, Mono-InternVL-1.5, with 1.8B activated parameters and 0.5B pre-training image-text pairs, matches the modular InternVL-1.5-2B on average across 15 benchmarks, reduces first-token latency by up to 69.3% at high resolution, and improves average monolithic-benchmark score by +2.8% over Emu3. The discovery is that delta tuning—training only newly inserted visual expert networks while keeping the pre-trained LLM frozen—lets the model learn visual knowledge from noisy data without catastrophic forgetting, and that adding visual attention experts plus reorganizing data toward smaller high-quality sets makes the pre-training markedly more data-efficient.
Load-bearing premise
The paper cuts pre-training data from 922M plus 258M pairs to 250M plus 150M pairs because saturation curves measured on a limited set of benchmarks and a single instruction-tuning recipe appear to flatten, and the claim that the reduced data transfers to other tasks and other tuning recipes rests entirely on those curves.
Editorial extensions
If this is right
- Removing the dedicated vision encoder makes deployment of MLLMs simpler and cheaper; the paper reports first-token latency drops of 65.7–69.3% depending on input resolution.
- Training budgets for monolithic MLLMs can shrink by roughly 58% (1.1B to 0.5B samples) without a downstream penalty, if the measured saturation behavior holds.
- A 1.8B-parameter monolithic model can outperform 8B-parameter monolithic rivals, including a +2.8% average gain over Emu3, suggesting the bottleneck was training design rather than parameter count.
- With the fused CUDA kernel, the modality-specific MoE runs at near single-branch speed, giving up to 2.32x speedup over PyTorch for linear MoE blocks and up to 48% throughput gains at high resolution.
- Separate visual and text experts preserve language ability: the unshared architecture scores substantially higher on NLP benchmarks than a shared-expert version.
Reading between the lines
- A testable extension would vary the instruction-tuning set, not just the pre-training set: if the saturation curves are an artifact of LLaVA-665k, the reduced 0.5B recipe could underperform when tuned with other instruction mixtures.
- Delta tuning keeps the base language model frozen, which implies visual expert capacity could be scaled much larger without retraining the backbone; the paper does not explore that scaling direction.
- The fused kernel's block-level early-exit trick should transfer to any hard-routing sparse MoE, not only vision/text splits, so it is a generic inference-speed contribution.
- Because the largest gains over prior monoliths appear on OCR-heavy and chart benchmarks, the recipe may be especially suited to document understanding; a stress test on natural-photo-only benchmarks would clarify how general the advantage is.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Mono-InternVL-1.5, a monolithic multimodal LLM built on InternLM2-1.8B by adding visual experts to both feed-forward and attention layers in a multimodal mixture-of-experts architecture, trained with an improved endogenous visual pre-training scheme (EViP++) that reduces pre-training data from about 1.1B/1.33B to 0.5B samples, and served with a custom fused CUDA kernel for the modality-specific MoE. The paper reports that Mono-InternVL-1.5 matches or improves on its modular counterpart InternVL-1.5-2B across 15 benchmarks while reducing time-to-first-token by up to 69.3%, and that it outperforms larger monolithic models such as Emu3. It also presents ablations of the visual-expert designs, data scaling, and kernel speedups.
Significance. If the results hold, the paper makes a useful contribution: it demonstrates that a 1.8B-activated-parameter monolithic MLLM can be trained with substantially less data and served with lower latency than a matched modular model, and it provides careful kernel-level and end-to-end efficiency measurements. The extensive ablations, the released code and models, and the direct latency/throughput tables in Sections V.D and V.C are concrete strengths. The significance is tempered, however, by the narrow set of modular baselines used to support the headline conclusion and by the untested generalization of the data-reduction curves in Fig. 5.
major comments (4)
- [Section VI, Abstract, Table III] The conclusion that monolithic MLLMs can reach 'comparable performance and superior efficiency to leading modular MLLMs' is not supported by the paper's own comparisons. The efficiency comparison in Table XIII is only against InternVL-1.5-2B, and the performance table includes newer modular baselines that outperform Mono-InternVL-1.5 on most shared benchmarks: InternVL-2.5-2B leads on MMB (74.7 vs 64.0), MMVet (60.8 vs 54.0), MMMU (43.6 vs 39.1), MathVista (51.3 vs 42.3), and HallusionBench (42.6 vs 32.5), while Qwen2VL-2B leads on MMB, MMMU, MathVista, and OCRBench. The evidence supports 'comparable to InternVL-1.5-2B-class modular models,' not 'leading modular MLLMs.' Please restrict the claim or add direct comparisons with models matched on LLM backbone and training data.
- [Section IV.A, Fig. 5] The data-efficiency claim rests on the saturation curves in Fig. 5, which are measured on a limited set of captioning and VQA benchmarks and with LLaVA-665k as the only instruction-tuning recipe. The decision to cut S1.1 from 922M to 250M and S1.2 from 258M to 150M assumes these curves generalize to other tasks, other high-quality data mixtures, and other instruction-tuning sets. This assumption is load-bearing: if the reduced data fails to transfer under a different tuning recipe, the 'cheaper' claim would not hold in general. Please validate with at least one held-out task or an alternative instruction-tuning recipe, or explicitly scope the data-efficiency claim to the tested configuration.
- [Table I, Section IV.A, Fig. 3] The reported 58% training-data reduction is internally inconsistent. Table I lists 1.1B vs 0.5B, which is a 54.5% reduction, while the stage-level totals in Fig. 3 and Section IV.A (922+258+143+7 = 1.33B vs 250+150+150+7 = 0.557B) give 58.1%. Please correct the table entries or the percentage and define precisely which stages are included in 'training data'.
- [Section V.C, Table V] The zero-shot pre-training comparison in Table V is compromised by training-set overlap. According to Table II, COCO caption data is used in S1.2 for both Mono-InternVL and Mono-InternVL-1.5, and VQAv2 is used in S1.3, yet Table V reports COCO Caps and VQAv2 as zero-shot without marking the S1.2 COCO rows or the VQAv2 row as 'seen.' The existing gray marking only covers the COCO row for Mono-InternVL-S1.3. Please mark all seen benchmarks or use truly held-out tasks for the zero-shot comparison with Flamingo, MM1, and Chameleon.
minor comments (8)
- [Section IV.B] The word 'paractice' should be 'practice.'
- [Section V.B] The word 'utulize' should be 'utilize.'
- [Section III.B] The phrase 'total total' should be 'total.'
- [Section I] The text says 'as shown in Fig. II' but the paper has no figure numbered II; please correct the reference.
- [Section V.C] The sentence 'Compared to Mono-InternVL, Mono-InternVL-1.5 shows comparable or even better performance on multiple benchmarks' appears twice in the comparisons with monolithic MLLMs; one occurrence should be removed.
- [Fig. 5] The caption uses 'V1' and 'V1.5' without defining them; please spell out Mono-InternVL and Mono-InternVL-1.5.
- [Table XIII] The latency and throughput measurements are reported as single numbers without repeated trials or error bars; please state the measurement variability or the number of runs used.
- [Abstract, Section VI, Table XIII] The statement 'reducing first-token latency by up to 69%' implicitly attributes the 69.3% figure to Mono-InternVL-1.5 as a whole, but that number in Table XIII is for the fused-kernel variant; the base Mono-InternVL-1.5 gives at most 66.7% at the 2048-token setting. Please attribute the number precisely.
Circularity Check
No circular derivation: the new efficiency and performance claims rest on fresh experiments and measurements; self-citation is transparent and not load-bearing.
full rationale
The paper is an empirical systems paper rather than a derivation chain. The central claims about Mono-InternVL-1.5 (0.5B pre-training data, +1.9% over Mono-InternVL on the MLLM benchmark average, 69.3% first-token latency reduction, and fused-kernel speedups) are established by new experiments and measurements in Tables I, III, IV, XII, and XIII and in Fig. 5, not by reusing the authors' prior conclusions. The explicit self-citation, 'This paper is built upon our work published in CVPR 2025 [11]', is transparent and describes the architectural base; it is not used as evidence for the new 1.5 results. The decision to cut S1.1/S1.2 data from 922M/258M to 250M/150M is justified by the empirical scaling curves in Fig. 5, which are an extrapolation from controlled LLaVA-665k instruction-tuned checkpoints; this is a stated empirical assumption rather than a fitted parameter renamed as a prediction, and the final model is evaluated independently on 15 benchmarks. Comparisons with InternVL-1.5-2B use a released external baseline and shared data and prompts, so the comparison is controlled rather than circular. The Section VI claim that monolithic MLLMs reach 'comparable performance and superior efficiency to leading modular MLLMs' overstates what the tables show relative to stronger modular baselines such as InternVL-2.5-2B and Qwen2VL-2B, but that is a claim-scope or correctness issue, not circular reasoning. No step in the paper reduces by construction to its own input.
Assumptions & free parameters
free parameters (4)
- S1.1 data scale =
250M (vs 922M in v1)
- S1.2 data scale =
150M (vs 258M in v1)
- S1.3 data scale =
150M (slightly increased from 143M)
- Visual expert parameter count =
1.2B
assumptions (4)
- domain assumption Delta tuning preserves pre-trained knowledge while allowing new modality learning.
- domain assumption InternLM2-1.8B contains transferable knowledge useful for visual tasks.
- domain assumption Synthetic captions from InternVL2-8B are high-quality and low-noise.
- ad hoc to paper Saturation curves in Fig. 5, measured on a limited benchmark set, generalize to other tasks and data.
invented entities (2)
-
Visual expert (FFNv)
-
Visual attention expert (V-QKVProj)
Cite this review
Pith. "Pith review of Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models." pith.science (2026). https://pith.science/paper/TMBRF6KZ
@misc{pith2026250712566,
author = {Pith},
title = {Pith review of: Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/TMBRF6KZ}},
note = {Machine review of arXiv:2507.12566}
}
read the original abstract
This paper focuses on monolithic Multimodal Large Language Models (MLLMs), which integrate visual encoding and language decoding into a single model. Existing structures and pre-training strategies for monolithic MLLMs often suffer from unstable optimization and catastrophic forgetting. To address these challenges, our key idea is to embed a new visual parameter space into a pre-trained LLM, enabling stable learning of visual knowledge from noisy data via delta tuning. Based on this principle, we first introduce Mono-InternVL, an advanced monolithic MLLM that incorporates a set of visual experts through a multimodal mixture-of-experts architecture. In addition, we design an innovative Endogenous Visual Pre-training (EViP) for Mono-InternVL to maximize its visual capabilities via progressive learning. Mono-InternVL achieves competitive performance against existing MLLMs but also leads to relatively expensive data cost. Therefore, we further present Mono-InternVL-1.5, a cheaper and stronger monolithic MLLM equipped with an improved EViP (EViP++). EViP++ introduces additional visual attention experts to Mono-InternVL-1.5 and re-organizes the pre-training process in an efficient manner. During inference, it includes a fused CUDA kernel to speed up its MoE operations. With these designs, Mono-InternVL-1.5 significantly reduces training and inference costs, while still maintaining competitive performance with Mono-InternVL. To evaluate our approach, we conduct extensive experiments across 15 benchmarks. Results demonstrate that Mono-InternVL outperforms existing monolithic MLLMs on 12 out of 15 benchmarks, e.g., +114-point improvement over Emu3 on OCRBench. Compared to its modular counterpart, i.e., InternVL-1.5, Mono-InternVL-1.5 achieves similar multimodal performance while reducing first-token latency by up to 69%. Code and models are released at https://github.com/OpenGVLab/Mono-InternVL.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
J. Bai, S. Bai, Y . Chu, Z. Cui, K. Dang, X. Deng, Y . Fan, W. Ge, Y . Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y . Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y . Zhang, ...
arXiv 2023
-
[3]
Z. Cai, M. Cao, H. Chen, K. Chen, K. Chen, X. Chen, X. Chen, Z. Chen, Z. Chen, P. Chu et al. , “Internlm2 technical report,” arXiv preprint arXiv:2403.17297, 2024. 1, 2, 3, 8, 11
arXiv 2024
-
[4]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML, vol. 139, 2021, pp. 8748–8763. 1, 3
2021
-
[5]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” in NeurIPS, 2023. 1, 3, 6
2023
-
[6]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,
Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma et al. , “How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites,” arXiv:2404.16821, 2024. 1, 2, 3, 5, 6, 7, 8
arXiv 2024
-
[7]
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML, vol. 202, 2023, pp. 19 730–19 742. 1, 3
2023
-
[8]
Introducing our multimodal models,
R. Bavishi, E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Ta¸ sırlar, “Introducing our multimodal models,” 2023. [Online]. Available: https://www.adept.ai/blog/fuyu-8b 1, 3, 7, 8
2023
Show all 137 references
-
[9]
Unveiling encoder-free vision-language models,
H. Diao, Y . Cui, X. Li, Y . Wang, H. Lu, and X. Wang, “Unveiling encoder-free vision-language models,” arXiv preprint arXiv:2406.11832,
-
[10]
A single transformer for scalable vision-language modeling,
Y . Chen, X. Wang, H. Peng, and H. Ji, “A single transformer for scalable vision-language modeling,” arXiv preprint arXiv:2407.06438 , 2024. 1, 3, 7, 8
2024 arXiv
-
[11]
Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,
G. Luo, X. Yang, W. Dou, Z. Wang, J. Liu, J. Dai, Y . Qiao, and X. Zhu, “Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training,” in CVPR, 2025. 1, 2
2025
-
[12]
Chameleon: Mixed-modal early-fusion foundation models,
ChameleonTeam, “Chameleon: Mixed-modal early-fusion foundation models,” arXiv preprint arXiv:2405.09818 , 2024. 1, 3, 4, 7, 8, 9, 11
2024 arXiv
-
[13]
Investigating the catastrophic forgetting in multimodal large language models,
Y . Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y . J. Lee, and Y . Ma, “Investigating the catastrophic forgetting in multimodal large language models,” arXiv preprint arXiv:2309.10313 , 2023. 1
2023 arXiv
-
[14]
Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models,
N. Ding, Y . Qin, G. Yang, F. Wei, Z. Yang, Y . Su, S. Hu, Y . Chen, C.-M. Chan, W. Chen et al., “Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models,” arXiv preprint arXiv:2203.06904, 2022. 1, 4
2022 arXiv
-
[15]
Qwen-vl: A frontier large vision-language model with versatile abilities,
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-vl: A frontier large vision-language model with versatile abilities,” arXiv preprint arXiv:2308.12966 , 2023. 1, 3
2023 arXiv
-
[16]
Lima: Less is more for alignment,
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y . Mao, X. Ma, A. Efrat, P. Yu, L. Yu et al., “Lima: Less is more for alignment,” Advances in Neural Information Processing Systems , vol. 36, pp. 55 006–55 021, 2023. 2, 6, 7
2023
-
[17]
Emu3: Next-token prediction is all you need,
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y . Cui, J. Wang, F. Zhang, Y . Wang, Z. Li, Q. Yu, Y . Zhao, Y . Ao, X. Min, T. Li, B. Wu, B. Zhao, B. Zhang, L. Wang, G. Liu, Z. He, X. Yang, J. Liu, Y . Lin, T. Huang, and Z. Wang, “Emu3: Next-token prediction is all you need,” arXiv: 2409...
2024 arXiv
-
[18]
The dawn of lmms: Preliminary explorations with gpt-4v (ision),
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang, “The dawn of lmms: Preliminary explorations with gpt-4v (ision),” arXiv: 2309.17421, vol. 9, 2023. 3
2023 arXiv
-
[19]
Gemini: a family of highly capable multimodal models,
G. Team, R. Anil, S. Borgeaud, Y . Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al., “Gemini: a family of highly capable multimodal models,” arXiv: 2312.11805, 2023. 3
2023 arXiv
-
[20]
Improved baselines with visual instruction tuning,
H. Liu, C. Li, Y . Li, and Y . J. Lee, “Improved baselines with visual instruction tuning,” arXiv: 2310.03744, 2023. 3, 9
2023 arXiv
-
[21]
Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models,
G. Luo, Y . Zhou, Y . Zhang, X. Zheng, X. Sun, and R. Ji, “Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models,” arXiv preprint arXiv:2403.03003 , 2024. 3
2024 arXiv
-
[22]
Parameter-inverted image pyramid networks for visual perception and multimodal understanding,
Z. Wang, X. Zhu, X. Yang, G. Luo, H. Li, C. Tian, W. Dou, J. Ge, L. Lu, Y . Qiao, and J. Dai, “Parameter-inverted image pyramid networks for visual perception and multimodal understanding,” arXiv preprint arXiv:2501.07783, 2025. 3
2025 arXiv
-
[23]
Synergen-vl: Towards synergistic image understanding and generation with vision experts and token folding,
H. Li, C. Tian, J. Shao, X. Zhu, Z. Wang, J. Zhu, W. Dou, X. Wang, H. Li, L. Lu et al., “Synergen-vl: Towards synergistic image understanding and generation with vision experts and token folding,” arXiv preprint arXiv:2412.09604, 2024. 3
2024 arXiv
-
[24]
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,
J. Li, D. Li, C. Xiong, and S. C. H. Hoi, “BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICLR, vol. 162, 2022, pp. 12 888–12 900. 3
2022
-
[25]
Instructblip: Towards general-purpose vision-language models with instruction tuning,
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. C. H. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” in NeurIPS, 2023. 3
2023
-
[26]
Llava-next: Improved reasoning, ocr, and world knowledge,
H. Liu, C. Li, Y . Li, B. Li, Y . Zhang, S. Shen, and Y . J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/ 3
2024
-
[27]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution,
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge et al., “Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution,” arXiv preprint arXiv:2409.12191, 2024. 3, 7, 8
2024 arXiv
-
[28]
Qwen2. 5-vl technical report,
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang et al., “Qwen2. 5-vl technical report,” arXiv preprint arXiv:2502.13923, 2025. 3
2025 arXiv
-
[29]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, B. Li, P. Luo, T. Lu, Y . Qiao, and J. Dai, “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” arXiv: 2312.14238, 2023. 3
2023 arXiv
-
[31]
Enhancing the reasoning ability of multimodal MONO-INTERNVL-1.5 14 large language models via mixed preference optimization,
W. Wang, Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, J. Zhu, X. Zhu, L. Lu, Y . Qiao et al., “Enhancing the reasoning ability of multimodal MONO-INTERNVL-1.5 14 large language models via mixed preference optimization,” arXiv preprint arXiv:2411.10442, 2024. 3
2024 arXiv
-
[32]
Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,
Z. Gao, Z. Chen, E. Cui, Y . Ren, W. Wang, J. Zhu, H. Tian, S. Ye, J. He, X. Zhu et al. , “Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% performance,” Visual Intelligence, vol. 2, no. 1, pp. 1–17, 2024. 3
2024
-
[33]
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” arXiv: 2302.13971, 2023. 3
2023 arXiv
-
[34]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv: 2307.09288, 2023. 3
2023 arXiv
-
[35]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR, 2021. 3
2021
-
[36]
Show-o: One single transformer to unify multimodal understanding and generation,
J. Xie, W. Mao, Z. Bai, D. J. Zhang, W. Wang, K. Q. Lin, Y . Gu, Z. Chen, Z. Yang, and M. Z. Shou, “Show-o: One single transformer to unify multimodal understanding and generation,” arXiv preprint arXiv:2408.12528, 2024. 3
2024 arXiv
-
[37]
Transfusion: Predict the next token and diffuse images with one multi-modal model,
C. Zhou, L. Yu, A. Babu, K. Tirumala, M. Yasunaga, L. Shamis, J. Kahn, X. Ma, L. Zettlemoyer, and O. Levy, “Transfusion: Predict the next token and diffuse images with one multi-modal model,” arXiv preprint arXiv:2408.11039, 2024. 3
2024 arXiv
-
[38]
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,” Advances in Neural Information Processing Systems, vol. 35, pp. 32 897–32 912, 2022. 3
2022
-
[39]
Image as a foreign language: Beit pretraining for all vision and vision-language tasks,
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, and F. Wei, “Image as a foreign language: Beit pretraining for all vision and vision-language tasks,” arXiv: 2208.10442, 2022. 3
2022 arXiv
-
[40]
Scaling vision-language models with sparse mixture of experts,
S. Shen, Z. Yao, C. Li, T. Darrell, K. Keutzer, and Y . He, “Scaling vision-language models with sparse mixture of experts,” arXiv preprint arXiv:2303.07226, 2023. 3
2023 arXiv
-
[41]
Twenty years of mixture of experts,
S. E. Yuksel, J. N. Wilson, and P. D. Gader, “Twenty years of mixture of experts,” IEEE transactions on neural networks and learning systems , vol. 23, no. 8, pp. 1177–1193, 2012. 3
2012
-
[42]
Moma: Efficient early-fusion pre-training with mixture of modality-aware experts,
X. V . Lin, A. Shrivastava, L. Luo, S. Iyer, M. Lewis, G. Gosh, L. Zettlemoyer, and A. Aghajanyan, “Moma: Efficient early-fusion pre-training with mixture of modality-aware experts,” arXiv preprint arXiv:2407.21770, 2024. 3
2024 arXiv
-
[43]
Mixture-of-depths: Dynamically allocating compute in transformer-based language models,
D. Raposo, S. Ritter, B. Richards, T. Lillicrap, P. C. Humphreys, and A. Santoro, “Mixture-of-depths: Dynamically allocating compute in transformer-based language models,” arXiv preprint arXiv:2404.02258 ,
-
[44]
Aria: An open multimodal native mixture-of- experts model,
D. Li, Y . Liu, H. Wu, Y . Wang, Z. Shen, B. Qu, X. Niu, F. Zhou, C. Huang, Y . Li et al., “Aria: An open multimodal native mixture-of- experts model,” arXiv preprint arXiv:2410.05993 , 2024. 3
2024 arXiv
-
[45]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS, 2017, pp. 5998–6008. 4
2017
-
[46]
Root mean square layer normalization,
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019. 4
2019
-
[47]
Laion-5b: An open large-scale dataset for training next generation image-text models,
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS, vol. 35, pp. 25 278–25 294, 2022. 5, 6, 9
2022
-
[48]
Coyo-700m: Image-text pair dataset,
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset,
-
[49]
Segment anything,
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, P. Dollár, and R. B. Girshick, “Segment anything,” arXiv: 2304.02643, 2023. 6
2023 arXiv
-
[50]
Kosmos-2: Grounding multimodal large language models to the world,
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei, “Kosmos-2: Grounding multimodal large language models to the world,” arXiv preprint arXiv:2306.14824 , 2023. 6
2023 arXiv
-
[51]
Microsoft coco captions: Data collection and evaluation server,
X. Chen, H. Fang, T.-Y . Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick, “Microsoft coco captions: Data collection and evaluation server,” arXiv preprint arXiv:1504.00325 , 2015. 6
2015 arXiv
-
[52]
Textcaps: A dataset for image captioning with reading comprehension,
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh, “Textcaps: A dataset for image captioning with reading comprehension,” in ECCV, vol. 12347, 2020, pp. 742–758. 6
2020
-
[53]
Objects365: A large-scale, high-quality dataset for object detection,
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” in ICCV, 2019, pp. 8430–8439. 6
2019
-
[54]
The all-seeing project: Towards panoptic visual recognition and understanding of the open world,
W. Wang, M. Shi, Q. Li, W. Wang, Z. Huang, L. Xing, Z. Chen, H. Li, X. Zhu, Z. Cao et al., “The all-seeing project: Towards panoptic visual recognition and understanding of the open world,” in ICLR, 2024. 6
2024
-
[55]
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,
J. Gu, X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang et al., “Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,” NeurIPS, vol. 35, pp. 26 418– 26 431, 2022. 6
2022
-
[56]
Laion coco: 600m synthetic captions from laion2b-en
C. Schuhmann, A. Köpf, R. Vencu, T. Coombes, and R. Beau- mont, “Laion coco: 600m synthetic captions from laion2b-en.” https://laion.ai/blog/laion-coco/, 2022. 6
2022
-
[57]
Mmc: Advancing multimodal chart understanding with large- scale instruction tuning,
F. Liu, X. Wang, W. Yao, J. Chen, K. Song, S. Cho, Y . Yacoob, and D. Yu, “Mmc: Advancing multimodal chart understanding with large- scale instruction tuning,” arXiv preprint arXiv:2311.10774 , 2023. 6
2023 arXiv
-
[58]
Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,
Y . Sun, Z. Ni, C.-K. Chng, Y . Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas et al., “Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,” in ICDAR, 2019, pp. 1557–1562. 6
2019
-
[59]
Scene text visual question answering,
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas, “Scene text visual question answering,” in ICCV, 2019, pp. 4291–4301. 6
2019
-
[60]
Icdar2017 competition on reading chinese text in the wild (rctw-17),
B. Shi, C. Yao, M. Liao, M. Yang, P. Xu, L. Cui, S. Belongie, S. Lu, and X. Bai, “Icdar2017 competition on reading chinese text in the wild (rctw-17),” in ICDAR, vol. 1, 2017, pp. 1429–1434. 6
2017
-
[61]
Icdar 2019 robust reading challenge on reading chinese text on signboard,
R. Zhang, Y . Zhou, Q. Jiang, Q. Song, N. Li, K. Zhou, L. Wang, D. Wang, M. Liao, M. Yang et al., “Icdar 2019 robust reading challenge on reading chinese text on signboard,” in ICDAR, 2019, pp. 1577–1581. 6
2019
-
[62]
Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,
C. K. Chng, Y . Liu, Y . Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al., “Icdar2019 robust reading challenge on arbitrary- shaped text-rrc-art,” in ICDAR, 2019, pp. 1571–1576. 6
2019
-
[63]
Ocr-free document understanding transformer,
G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park, “Ocr-free document understanding transformer,” in ECCV, 2022. 6
2022
-
[64]
Coco-text: Dataset and benchmark for text detection and recognition in natural images,
A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie, “Coco-text: Dataset and benchmark for text detection and recognition in natural images,” arXiv preprint arXiv:1601.07140 , 2016. 6
2016 arXiv
-
[65]
Chartqa: A benchmark for question answering about charts with visual and logical reasoning,
A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque, “Chartqa: A benchmark for question answering about charts with visual and logical reasoning,” in ACL, 2022, pp. 2263–2279. 6, 8
2022
-
[66]
A large chinese text dataset in the wild,
T.-L. Yuan, Z. Zhu, K. Xu, C.-J. Li, T.-J. Mu, and S.-M. Hu, “A large chinese text dataset in the wild,” Journal of Computer Science and Technology, vol. 34, pp. 509–521, 2019. 6
2019
-
[67]
Simple and effective multi-paragraph reading comprehension,
C. Clark and M. Gardner, “Simple and effective multi-paragraph reading comprehension,” in ACL, 2018, pp. 845–855. 6, 8
2018
-
[68]
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner, “Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,” in CVPR, 2021, pp. 8802–8812. 6
2021
-
[69]
Plotqa: Reasoning over scientific plots,
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, “Plotqa: Reasoning over scientific plots,” in WACV, 2020, pp. 1527–1536. 6
2020
-
[70]
Infographicvqa,
M. Mathew, V . Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar, “Infographicvqa,” in WACV, 2022, pp. 1697–1706. 6, 8
2022
-
[71]
Making the V in VQA matter: Elevating the role of image understanding in visual question answering,
Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, “Making the V in VQA matter: Elevating the role of image understanding in visual question answering,” in CVPR, 2017, pp. 6325–6334. 6
2017
-
[72]
GQA: A new dataset for real-world visual reasoning and compositional question answering,
D. A. Hudson and C. D. Manning, “GQA: A new dataset for real-world visual reasoning and compositional question answering,” in CVPR, 2019, pp. 6700–6709. 6, 8
2019
-
[73]
Ok-vqa: A visual question answering benchmark requiring external knowledge,
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in CVPR, 2019, pp. 3195–3204. 6
2019
-
[74]
Visual spatial reasoning,
F. Liu, G. Emerson, and N. Collier, “Visual spatial reasoning,” TACL, vol. 11, pp. 635–651, 2023. 6
2023
-
[75]
Visual dialog,
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual dialog,” in CVPR, 2017, pp. 326–335. 6
2017
-
[76]
A diagram is worth a dozen images,
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “A diagram is worth a dozen images,” in ECCV, 2016, pp. 235–251. 6, 8
2016
-
[77]
Learn to explain: Multimodal reasoning via thought chains for science question answering,
P. Lu, S. Mishra, T. Xia, L. Qiu, K. Chang, S. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” in NeurIPS, 2022. 6, 8
2022
-
[78]
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Ha- jishirzi, “Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension,” in CVPR, 2017, pp. 4999–5007. 6
2017
-
[79]
Dvqa: Understanding data visualizations via question answering,
K. Kafle, B. Price, S. Cohen, and C. Kanan, “Dvqa: Understanding data visualizations via question answering,” in CVPR, 2018, pp. 5648–5656. 6 MONO-INTERNVL-1.5 15
2018
-
[80]
Aligning large multi-modal model with robust instruction tuning,
F. Liu, K. Lin, L. Li, J. Wang, Y . Yacoob, and L. Wang, “Aligning large multi-modal model with robust instruction tuning,” arXiv preprint arXiv:2306.14565, 2023. 6
2023 arXiv
-
[81]
An augmented benchmark dataset for geometric question answering through dual parallel text encoding,
J. Cao and J. Xiao, “An augmented benchmark dataset for geometric question answering through dual parallel text encoding,” in COLING, 2022, pp. 1511–1520. 6
2022
-
[82]
Dynamic prompt learning via policy gradient for semi- structured mathematical reasoning,
P. Lu, L. Qiu, K.-W. Chang, Y . N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan, “Dynamic prompt learning via policy gradient for semi- structured mathematical reasoning,” arXiv preprint arXiv:2209.14610 ,
-
[83]
Metamath: Bootstrap your own mathematical questions for large language models,
L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y . Zhang, J. T. Kwok, Z. Li, A. Weller, and W. Liu, “Metamath: Bootstrap your own mathematical questions for large language models,” arXiv preprint arXiv:2309.12284 ,
-
[84]
Clevr-math: A dataset for compositional language, visual and mathematical reasoning,
A. D. Lindström and S. S. Abraham, “Clevr-math: A dataset for compositional language, visual and mathematical reasoning,” arXiv preprint arXiv:2208.05358, 2022. 6
2022 arXiv
-
[85]
Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,
Z. Li, X. Wang, E. Stengel-Eskin, A. Kortylewski, W. Ma, B. Van Durme, and A. L. Yuille, “Super-clevr: A virtual benchmark to diagnose domain robustness in visual reasoning,” in CVPR, 2023, pp. 14 963–14 973. 6
2023
-
[86]
Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning,
P. Lu, R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu, “Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning,” arXiv preprint arXiv:2105.04165 , 2021. 6
2021 arXiv
-
[87]
Kvqa: Knowledge- aware visual question answering,
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar, “Kvqa: Knowledge- aware visual question answering,” in AAAI, vol. 33, no. 01, 2019, pp. 8876–8884. 6
2019
-
[88]
A-okvqa: A benchmark for visual question answering using world knowledge,
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual question answering using world knowledge,” in ECCV, 2022, pp. 146–162. 6
2022
-
[89]
Viquae, a dataset for knowledge- based visual question answering about named entities,
P. Lerner, O. Ferret, C. Guinaudeau, H. Le Borgne, R. Besançon, J. G. Moreno, and J. Lovón Melgarejo, “Viquae, a dataset for knowledge- based visual question answering about named entities,” in SIGIR, 2022, pp. 3108–3120. 6
2022
-
[90]
Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models,
C. He, Z. Jin, C. Xu, J. Qiu, B. Wang, W. Li, H. Yan, J. Wang, and D. Lin, “Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models,” arXiv preprint arXiv:2308.10755 ,
-
[91]
Ocr-vqa: Visual question answering by reading text in images,
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty, “Ocr-vqa: Visual question answering by reading text in images,” in ICDAR, 2019, pp. 947–952. 6
2019
-
[92]
Towards VQA models that can read,
A. Singh, V . Natarajan, M. Shah, Y . Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards VQA models that can read,” in CVPR,
-
[93]
Modeling context in referring expressions,
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV, vol. 9906, 2016, pp. 69–85. 6
2016
-
[94]
Generation and comprehension of unambiguous object descriptions,
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy, “Generation and comprehension of unambiguous object descriptions,” in CVPR, 2016, pp. 11–20. 6
2016
-
[95]
Visual genome: Connecting language and vision using crowdsourced dense image annotations,
R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” IJCV, vol. 123, no. 1, pp. 32–73, 2017. 6
2017
-
[96]
To see is to believe: Prompting gpt-4v for better visual instruction tuning,
J. Wang, L. Meng, Z. Weng, B. He, Z. Wu, and Y .-G. Jiang, “To see is to believe: Prompting gpt-4v for better visual instruction tuning,” arXiv preprint arXiv:2311.07574, 2023. 6
2023 arXiv
-
[97]
Allava: Harnessing gpt4v-synthesized data for a lite vision-language model,
G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang, “Allava: Harnessing gpt4v-synthesized data for a lite vision-language model,” arXiv preprint arXiv:2402.11684,
-
[98]
Gpt-4v dataset,
LAION, “Gpt-4v dataset,” https://huggingface.co/datasets/laion/ gpt4v-dataset, LAION, 2023. 6
2023
-
[99]
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” NeurIPS, vol. 36, 2024. 6
2024
-
[100]
SVIT: scaling up visual instruction tuning,
B. Zhao, B. Wu, and T. Huang, “SVIT: scaling up visual instruction tuning,” arXiv: 2307.04087, 2023. 6
2023 arXiv
-
[101]
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,
Teknium, “Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,” https://huggingface.co/datasets/teknium/ OpenHermes-2.5, HuggingFace, 2023. 6
2023
-
[102]
Alpaca: A strong, replicable instruction- following model,
R. Taori, I. Gulrajani, T. Zhang, Y . Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Alpaca: A strong, replicable instruction- following model,” Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html , vol. 3, no. 6, p. 7,
2023
-
[103]
Coig-cqia: Quality is all you need for chinese instruction fine-tuning,
Y . Bai, X. Du, Y . Liang, Y . Jin, Z. Liu, J. Zhou, T. Zheng, X. Zhang, N. Ma, Z. Wang et al., “Coig-cqia: Quality is all you need for chinese instruction fine-tuning,” arXiv preprint arXiv:2403.18058 , 2024. 6
2024 arXiv
-
[104]
Egotaskqa: Understanding human tasks in egocentric videos,
B. Jia, T. Lei, S.-C. Zhu, and S. Huang, “Egotaskqa: Understanding human tasks in egocentric videos,” Advances in Neural Information Processing Systems, vol. 35, pp. 3343–3360, 2022. 6
2022
-
[105]
Mementos: A comprehensive benchmark for multimodal large language model reasoning over image sequences,
X. Wang, Y . Zhou, X. Liu, H. Lu, Y . Xu, F. He, J. Yoon, T. Lu, G. Bertasius, M. Bansal et al., “Mementos: A comprehensive benchmark for multimodal large language model reasoning over image sequences,” arXiv preprint arXiv:2401.10529 , 2024. 6
2024 arXiv
-
[106]
Star: A benchmark for situated reasoning in real-world videos,
B. Wu, S. Yu, Z. Chen, J. B. Tenenbaum, and C. Gan, “Star: A benchmark for situated reasoning in real-world videos,” in Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS) , 2021. 6
2021
-
[107]
Ntu rgb+ d: A large scale dataset for 3d human activity analysis,
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 1010–1019. 6
2016
-
[108]
Videochat: Chat-centric video understanding,
K. Li, Y . He, Y . Wang, Y . Li, W. Wang, P. Luo, Y . Wang, L. Wang, and Y . Qiao, “Videochat: Chat-centric video understanding,” arXiv preprint arXiv:2305.06355, 2023. 6
2023 arXiv
-
[109]
Movie description,
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” International Journal of Computer Vision , 2017. [Online]. Available: http://link.springer.com/article/10.1007/s11263-016-0987-1? wt_mc=Internal.Event.1.SE...
2017 doi
-
[110]
Icdar2019 competition on scanned receipt ocr and information extraction,
Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. Jawa- har, “Icdar2019 competition on scanned receipt ocr and information extraction,” in 2019 International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 2019, pp. 1516–1520. 6
2019
-
[111]
Funsd: A dataset for form understanding in noisy scanned documents,
J.-P. T. Guillaume Jaume, Hazim Kemal Ekenel, “Funsd: A dataset for form understanding in noisy scanned documents,” in Accepted to ICDAR-OST, 2019. 6
2019
-
[112]
Visual information extraction in the wild: practical dataset and end-to- end solution,
J. Kuang, W. Hua, D. Liang, M. Yang, D. Jiang, B. Ren, and X. Bai, “Visual information extraction in the wild: practical dataset and end-to- end solution,” in International Conference on Document Analysis and Recognition. Springer, 2023, pp. 36–53. 6
2023
-
[113]
Mobilevlm v2: Faster and stronger baseline for vision language model,
X. Chu, L. Qiao, X. Zhang, S. Xu, F. Wei, Y . Yang, X. Sun, Y . Hu, X. Lin, B. Zhang et al., “Mobilevlm v2: Faster and stronger baseline for vision language model,” arXiv preprint arXiv:2402.03766 , 2024. 7, 8
2024 arXiv
-
[114]
Mini-gemini: Mining the potential of multi-modality vision language models,
Y . Li, Y . Zhang, C. Wang, Z. Zhong, Y . Chen, R. Chu, S. Liu, and J. Jia, “Mini-gemini: Mining the potential of multi-modality vision language models,” arXiv: 2403.18814, 2024. 7, 8
2024 arXiv
-
[116]
Deepseek-vl: Towards real-world vision-language understanding,
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, Y . Sun et al., “Deepseek-vl: Towards real-world vision-language understanding,” arXiv preprint arXiv:2403.05525 , 2024. 7, 8
2024 arXiv
-
[117]
Paligemma: A versatile 3b vlm for transfer,
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello et al. , “Paligemma: A versatile 3b vlm for transfer,” arXiv preprint arXiv:2407.07726, 2024. 7, 8
2024 arXiv
-
[118]
Minicpm-v: A gpt-4v level mllm on your phone,
Y . Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He et al. , “Minicpm-v: A gpt-4v level mllm on your phone,” arXiv preprint arXiv:2408.01800 , 2024. 7, 8
2024 arXiv
-
[119]
Expanding performance boundaries of open- source multimodal models with model, data, and test-time scaling,
Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu et al., “Expanding performance boundaries of open- source multimodal models with model, data, and test-time scaling,” arXiv preprint arXiv:2412.05271 , 2024. 7, 8
2024 arXiv
-
[120]
Vision as lora,
H. Wang, Y . Ye, B. Li, Y . Nie, J. Lu, J. Tang, Y . Wang, and C. Huang, “Vision as lora,” arXiv preprint arXiv:2503.20680 , 2025. 7, 8, 9
2025 arXiv
-
[121]
Mmbench: Is your multi-modal model an all-around player?
Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin, “Mmbench: Is your multi-modal model an all-around player?” arXiv: 2307.06281, 2023. 8
2023 arXiv
-
[122]
Mm-vet: Evaluating large multimodal models for integrated capabilities,
W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang, “Mm-vet: Evaluating large multimodal models for integrated capabilities,” arXiv: 2308.02490, 2023. 8
2023 arXiv
-
[123]
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,
X. Yue, Y . Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y . Sun et al., “Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,” arXiv: 2311.16502, 2023. 8 MONO-INTERNVL-1.5 16
2023 arXiv
-
[124]
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao, “Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts,” arXiv: 2310.02255,
-
[125]
Seed-bench: Benchmarking multimodal llms with generative comprehension,
B. Li, R. Wang, G. Wang, Y . Ge, Y . Ge, and Y . Shan, “Seed-bench: Benchmarking multimodal llms with generative comprehension,” arXiv: 2307.16125, 2023. 8
2023 arXiv
-
[126]
On the hidden mystery of ocr in large multimodal models,
Y . Liu, Z. Li, H. Li, W. Yu, M. Huang, D. Peng, M. Liu, M. Chen, C. Li, L. Jin et al., “On the hidden mystery of ocr in large multimodal models,” arXiv preprint arXiv:2305.07895 , 2023. 8
2023 arXiv
-
[127]
Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models,
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y . Yacoobet al., “Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models,” arXiv: 2310.14566, 2023. 8
-
[128]
Measuring massive multitask language understanding,
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” arXiv preprint arXiv:2009.03300 , 2020. 8
2009 arXiv
-
[129]
Cmmlu: Measuring massive multitask language understanding in chinese,
H. Li, Y . Zhang, F. Koto, Y . Yang, H. Zhao, Y . Gong, N. Duan, and T. Baldwin, “Cmmlu: Measuring massive multitask language understanding in chinese,” arXiv preprint arXiv:2306.09212 , 2023. 8
2023 arXiv
-
[130]
Agieval: A human-centric benchmark for evaluating foundation models,
W. Zhong, R. Cui, Y . Guo, Y . Liang, S. Lu, Y . Wang, A. Saied, W. Chen, and N. Duan, “Agieval: A human-centric benchmark for evaluating foundation models,” arXiv preprint arXiv:2304.06364 , 2023. 8
2023 arXiv
-
[131]
Measuring mathematical problem solving with the math dataset,
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the math dataset,” arXiv preprint arXiv:2103.03874 , 2021. 8
2021 arXiv
-
[132]
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,
H. Duan, J. Yang, Y . Qiao, X. Fang, L. Chen, Y . Liu, X. Dong, Y . Zang, P. Zhang, J. Wang, D. Lin, and K. Chen, “Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,” 2024. [Online]. Available: https://arxiv.org/abs/2407.11691 8
2024 arXiv
-
[133]
Opencompass: A universal evaluation platform for foundation models,
Contributors, “Opencompass: A universal evaluation platform for foundation models,” https://github.com/open-compass/opencompass,
-
[134]
Flamingo: a visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” NeurIPS, vol. 35, pp. 23 716– 23 736, 2022. 9
2022
-
[135]
Mm1: Methods, anal- ysis & insights from multimodal llm pre-training,
B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers et al. , “Mm1: Methods, anal- ysis & insights from multimodal llm pre-training,” arXiv preprint arXiv:2403.09611, 2024. 9
2024 arXiv
-
[136]
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” TACL, vol. 2, pp. 67–78, 2014. 10
2014
-
[137]
Lmdeploy: A toolkit for compressing, deploy- ing, and serving llm,
LMDeployContributors, “Lmdeploy: A toolkit for compressing, deploy- ing, and serving llm,” https://github.com/InternLM/lmdeploy, 2023. 11, 12
2023
-
[138]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778. 12
2016
-
[2024]
1, 3, 4, 7, 8, 9, 11
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.