REVIEW 3 major objections 7 minor 1 cited by
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that instruction tuning of multimodal language models with a text-heavy mixture (75% text-only, 25% vision-language) under a fixed 186,000-instance budget matches or outperforms vision-heavy mixtures on both text and…
desk verdict Useful controlled sweep on modality ratios with token accounting; the vision-parity claim needs seeds and a validation split before it's believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is task-level semantic alignment between text-only and vision-language instruction data. The paper samples 100,000 instruction prompts from Super-NaturalInstructions and Vision-Flan, embeds them with a pretrained sentence transformer, and reports a significantly non-negative mean cosine similarity between the two modalities, reasoning that tasks are defined by their instructions and that shared task semantics transfer once vision-language pretraining aligns image tokens with text tokens. The method itself is a controlled training recipe: a fixed 186,000-instance budget, FLAN-style formatting, a CLIP-ViT-L/14@336 visual encoder with a two-layer MLP projector, and a fully unfrozen LLM, with only the data composition varied across experiments.
What would settle it
Evaluate MLAN on vision tasks with no close text-only analogue, such as OCR, fine-grained object grounding, or object counting in images, under the same 186,000-instance budget; if the 75% text-heavy mixture falls clearly below the vision-heavy Cambrian-1 mixture on those benchmarks, the claim that language-based tuning generally preserves and transfers vision knowledge is falsified.
Extended reading notes
Core claim
The central discovery is that instruction tuning of a multimodal LLM can be re-oriented around text without losing vision performance. Under a fixed budget of 186,000 training instances, the MLAN mixture of 75% text-only and 25% vision-language instructions yields held-out zero-shot performance on both modalities that is on par with or better than vision-heavy recipes (MIX-LLaVA-1.5 with 6% text and MIX-Cambrian-1 with 25% text). On the language benchmarks, MLAN averages 64.50 versus 63.18 for MIX-Cambrian-1 on Llama-3.2-3B and 71.57 versus 67.76 on Llama-3.1-8B; on vision benchmarks it averages 59.13 versus 58.57 on the 3B model and trails MIX-Cambrian-1 by 0.33 points on the 8B model. The text-heavy model processes 60.1 million training tokens, compared with 101.5 million for the Cambrian-1 mix and 117.2 million for the LLaVA-1.5 mix, a reduction of roughly 40% to 49%. The paper argues the transfer is possible because text-only and vision-language instructions share task-level semantics, and its controlled ablation shows that even 12.5% text-only data sharply raises both text and vision scores.
Load-bearing premise
The transfer mechanism assumes that cosine similarity between embedded text-only and vision-language instruction prompts is a reliable proxy for whether abilities learned from text-only data will transfer to image-grounded tasks; if that similarity is not the right proxy, the motivation for the text-heavy mixture weakens even if the fixed-budget empirical results still hold.
Editorial extensions
If this is right
- A fixed-budget instruction tuning set can be made 75% text-only without sacrificing held-out vision performance, on both a 3B and an 8B Llama-based multimodal model.
- Vision-heavy instruction mixtures (6% to 25% text) erode language knowledge on datasets like CommonsenseQA and CosmosQA by up to 20 percentage points, while the text-heavy mixture largely avoids that degradation.
- Because CLIP converts each image into 576 visual tokens, replacing vision instances with text instances at the same instance budget cuts the number of training tokens processed by roughly half.
- The mixture ratio is not arbitrary: 12.5% text-only data already produces a sharp gain on both text and vision axes, and vision performance peaks at a moderate language share, showing that neither pure modality is sufficient.
Reading between the lines
- If text-heavy tuning generalizes beyond the two models tested here, the cost bottleneck of multimodal instruction tuning shifts from collecting image-text pairs to assembling diverse text-only task mixtures; a testable extension is to select text-only tasks whose instruction embeddings are most similar to a target vision benchmark and measure the resulting vision score.
- The cosine-similarity analysis implies a practical data-selection tool: embed candidate text-only and vision-language instructions, then choose cross-modally similar subsets instead of fixing ratios by hand.
- The asymmetric forgetting pattern — text abilities erode under vision-heavy tuning while vision abilities improve under text tuning — suggests language knowledge is the fragile resource in multimodal models, a prediction that could be checked on other architectures and other pretraining corpora.
- The paper's own limitation section notes that OCR, captioning, and other specialized out-of-distribution vision tasks were not evaluated; a natural next test is whether the 75/25 mixture holds up when those tasks are added to the benchmark suite.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MLAN, a visual instruction tuning strategy that replaces most vision-language instruction data with text-only data (75% text, 25% vision) under a fixed 186k-instance budget after multimodal pretraining. Using LLaVA-style models built on Llama-3.2-3B and Llama-3.1-8B, the authors compare MLAN against simplified mixtures with LLaVA-1.5-style (6% text) and Cambrian-1-style (25% text) ratios on seven text-only and five vision-language held-out benchmarks. The headline finding is that MLAN matches or exceeds these vision-heavy proxy mixtures on vision benchmarks, clearly improves text benchmarks, and processes roughly 40-50% fewer training tokens. Additional experiments explore the language-data ratio curve, task diversity, pretraining data, and instruction-tuned versus base backbones.
Significance. If the result is robust to seed variation and does not come from selection on the evaluation benchmarks, this is a valuable and cost-saving empirical contribution. The paper addresses an underexplored design choice (modality composition in MLLM instruction tuning) with controlled fixed-budget comparisons at two model scales, includes token-level efficiency accounting, and provides several ablations. It is also candid about limitations, including the narrow architecture scope and the absence of specialized tasks such as OCR and captioning. However, the central 'matches or exceeds' claim currently rests on small vision-side margins from single runs, and the chosen 75% ratio appears to have been selected from the same evaluation curves used in the main tables. These issues need to be addressed before the quantitative conclusion can be considered reliable.
major comments (3)
- [§3.1, Tables 1–2] All results in the main comparison are single runs, with no error bars, standard deviations, or repeated-seed analysis reported anywhere in the manuscript. The vision-side difference that supports 'matches or exceeds' is extremely small for the larger model: MLAN's vision average is 62.25 versus 62.58 for MIX-Cambrian-1 in Table 1, a 0.33-point deficit, while individual benchmarks trade by several points (MMMU 34.44 vs 36.00; MMBench 72.51 vs 73.50; POPE 81.84 vs 82.57). Under sampling or optimization noise, that deficit could easily become a multi-point gap, which would change the conclusion from 'matches or exceeds' to 'slightly worse on vision with large text gains.' Please provide at least three seeds with error bars or confidence intervals, or explicitly restrict the claim to 'comparable on vision' with the uncertainty stated. The Limitations section does not currently address this single-run variance.
- [§3.3, Figure 3] The paper selects the 75% text-only ratio for MLAN after inspecting the knowledge-transfer curves in Figure 3, which are computed on the same held-out evaluation benchmarks used in Tables 1 and 2; no separate validation split is described. If the ratio was chosen by peeking at these curves, the reported comparison is optimistically selected rather than a prediction, and the 0.33-point vision deficit could reflect overfitting to the chosen evaluation suite. Please either document that the ratio was fixed before evaluation, or introduce a validation split (for example, hold out a subset of the 12 benchmarks for ratio selection and report the remaining benchmarks as the headline results).
- [§3.1, Table 3 and Appendix D.1] The baselines called 'MIX-LLaVA-1.5' and 'MIX-Cambrian-1' are not the actual LLaVA-1.5 or Cambrian-1 instruction tuning recipes: they are simplified mixtures sampled from the same two datasets (Super-NaturalInstructions and Vision-Flan) with 6% and 25% text-only ratios at a fixed 186k-instance budget. Actual LLaVA-1.5 uses 665k instances and Cambrian-1 uses millions of instances with different data sources, as Table 8 itself shows. Therefore, the conclusion that MLAN 'matches or better performance' on downstream vision-language tasks compared with these state-of-the-art recipes overstates what is directly supported. Please either compare against the real recipes (ideally at matched token budgets) or rename the baselines to 'proxy mixtures with LLaVA-1.5/Cambrian-1 text ratios' and avoid the state-of-the-art claim.
minor comments (7)
- [Table 1, Appendix A, Table 7] The treatment of ScienceQA is inconsistent: the Table 1 note says it is 'included in Vision-Flan but excluded in experiments,' whereas Appendix A says it is removed from the training set for evaluation, and Table 7 lists it as an evaluation benchmark. Please clarify whether it was excluded from training or from evaluation, or both, and correct the wording accordingly.
- [§2.1, Figure 4] The statistical tests for cosine similarity report only p-values; please also report effect sizes and confidence intervals, since with 100k samples a 'significantly non-negative mean' is a weak statement that does not quantify how similar the instructions actually are.
- [§3.3, Figure 3] The knowledge-transfer curve would be more informative with per-benchmark values and error bars; currently the text 'peaks and then slightly declines' cannot be quantitatively verified from the figure.
- [§3.4, Table 4] Because most production MLLMs use an instruction-tuned chat backbone, the 'Instruct LLM' row is an important caveat: the main comparison in Tables 1–2 uses non-instruction base models, so the practical significance of MLAN for standard recipes remains unclear. Please discuss whether the main conclusions hold with chat backbones.
- [Appendix D.1, Table 8] The Cambrian-1 row in Table 8 is garbled ('Cambrian-1 (Tong et al., 2024) – Cambrian-7M 1.68M ∼7M 23.8%'); please fix the formatting so the dataset size and text-only percentage are unambiguous.
- [§2.3, Tables 6–7] The text states '12 comprehensive benchmarks' while Tables 6–7 list 13 datasets because ARC-E and ARC-C are reported separately; please align the count.
- [Appendix B / Data Availability] No code, configuration, or checkpoint release is mentioned; given that the headline depends on exact data sampling and the 75% ratio, releasing the data mixture and training configuration would substantially improve reproducibility.
Circularity Check
No significant circularity: MLAN is an empirical data-mixture study judged against external held-out benchmarks, and no derivation reduces to its own inputs.
full rationale
The paper's central claim is empirical: under a fixed 186,000-instance budget, a 75% text-only / 25% vision-language instruction tuning mixture matches or exceeds vision-heavy mixtures on held-out text and vision benchmarks. This claim is tested against external datasets (ARC, MMLU, POPE, MMMU, MME, MMBench, etc.) and against two baseline mixtures, MIX-LLaVA-1.5 and MIX-Cambrian-1, which are reproduced under the same training protocol. MLAN is not derived from the evaluation scores; the 75% ratio is a dataset-composition choice reported as a method, and the main results are direct benchmark measurements. The cosine-similarity analysis in Section 2.1 is motivational rather than derivational: it supports the hypothesis that instruction semantics are shared across modalities, but the transfer claim is not computed from those similarity scores. There is no equation in which the output is identical to an input by construction, and no parameter is fitted to a subset of data and then renamed as a prediction. The paper does not rely on a load-bearing self-citation or an imported uniqueness theorem; its architecture and training recipe follow standard LLaVA practice with external references. The most legitimate concern is that the 75% ratio appears to have been selected after inspecting the knowledge-transfer curve in Figure 3 on the same evaluation benchmarks used in the main tables, with no validation split or repeated seeds described. That is a selection-bias and reproducibility concern about the strength of the 'matches or exceeds' claim, not a definitional or constructional circularity, because the reported numbers are not forced to equal the curve by construction. Accordingly, the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- Text-only ratio for MLAN =
75% text-only, 25% vision-language
- Instruction tuning instance budget =
186,000
assumptions (2)
- domain assumption Strong visual pretraining alignment makes image tokens functionally similar to text tokens for instruction following
- domain assumption Cosine similarity of instruction embeddings (all-mpnet-base-v2) is a valid measure of cross-modal task transferability
Cite this review
Pith. "Pith review of MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models." pith.science (2026). https://pith.science/paper/TMM7B6CX
@misc{pith2026241110557,
author = {Pith},
title = {Pith review of: MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/TMM7B6CX}},
note = {Machine review of arXiv:2411.10557}
}
read the original abstract
We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base. Existing work lacks sufficient experimentation on the importance of each modality in the instruction tuning stage, often using a majority of vision-language data while keeping text-only data limited and fixing mixtures of modalities. By incorporating diverse text-only data in the visual instruction tuning stage, we vary vision-language data in various controlled experiments to investigate the importance of modality in visual instruction tuning. Our comprehensive evaluation shows that the text-heavy instruction tuning approach is able to perform on-par with traditional vision-heavy mixtures on both modalities across 12 general datasets while using as low as half the total training tokens. We find that simply increasing sufficiently diverse text-only data enables transfer of instruction following ability and domain knowledge across modalities while being more efficient than the vision-language approach.
Figures
Forward citations
Cited by 1 Pith paper
-
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
POINTS-Seeker-8B is an 8B multimodal model trained from scratch for agentic search that uses seeding and visual-space history folding to outperform prior models on six visual reasoning benchmarks.
Reference graph
Works this paper leans on
-
[1]
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, Soham Ghosh, Am \'e lie H \'e liou, Paul Jacob, Albert Q. Jiang, Kartik Khandelwal, Timoth \'e e Lacroix, Guillaume Lample, Diego Las Casas, Thibaut Lavril, and 23 others. 2024. https://...
-
[2]
Menick, Sebastian Borgeaud, and 8 others
Jean - Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, and 8 others. 2022. http://papers.nips.cc/paper\_files/paper...
2022
-
[3]
Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. 2024. https://arxiv.org/abs/2404.03413 Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens . Preprint, arXiv:2404.03413
arXiv 2024
-
[4]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. https://doi.org/10.48550/ARXIV.2308.12966 Qwen-vl: A frontier large vision-language model with versatile abilities . CoRR, abs/2308.12966
-
[5]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025. https://doi.org/10.48550/arXiv.2502.13923 Qwen2.5- VL Technical Report . Preprint, arXiv:2502.13923
-
[6]
Sumithra Bhakthavatsalam, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, and Peter Clark. 2021. https://arxiv.org/abs/2102.03315 Think you have solved direct-answer question answering? try arc-da, the direct-answer AI2 reasoning challenge . CoRR, abs/2102.03315
arXiv 2021
-
[7]
Jinhe Bi, Yifan Wang, Danqi Yan, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. https://doi.org/10.48550/arXiv.2502.12119 PRISM : Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection . Preprint, arXiv:2502.12119
-
[8]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://arxiv.org/abs/1911.11641 PIQA : Reasoning about physical commonsense in natural language . In Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence
arXiv 2020
Show all 82 references
-
[9]
Cheng Chen, Junchen Zhu, Xu Luo, Hengtao Shen, Lianli Gao, and Jingkuan Song. 2024 a . Coin: A benchmark of continual instruction tuning for multimodel large language model. arXiv preprint arXiv:2403.08350
2024 arXiv
-
[10]
Delong Chen, Jianfeng Liu, Wenliang Dai, and Baoyuan Wang. 2024 b . https://doi.org/10.1609/AAAI.V38I16.29727 Visual instruction tuning with polite flamingo . In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applicat...
2024 doi
- [11]
-
[12]
Ruibo Chen, Yihan Wu, Lichang Chen, Guodong Liu, Qi He, Tianyi Xiong, Chenxi Liu, Junfeng Guo, and Heng Huang. 2024 c . Your vision-language model itself is a strong filter: Towards high-quality instruction tuning with data selection. arXiv preprint arXiv:2402.12501
2024 arXiv
- [13]
- [14]
-
[15]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\
2023
-
[16]
Christopher Clark, Kenton Lee, Ming - Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. https://doi.org/10.18653/V1/N19-1300 Boolq: Exploring the surprising difficulty of natural yes/no questions . In Proceedings of the 2019 Conference of the North Ame...
2019 doi
- [17]
-
[18]
Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuoling Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. https://arxiv.org/abs/2409.11402 Nvlm: Open frontier-class multimodal llms . Preprint, arXiv:2409.11402
2024 arXiv
-
[19]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/9a6a435e75419a836fe47ab6793623e6-Abstract-Conference.html Instructblip: Towards gener...
2023
- [20]
-
[21]
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint...
2023
-
[22]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
- [23]
-
[24]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300
2020 arXiv
-
[25]
Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. https://doi.org/10.18653/V1/D19-1243 Cosmos QA: machine reading comprehension with contextual commonsense reasoning . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing...
2019 doi
-
[26]
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Nils Johan Bertil Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. 2023. http://papers.nips...
2023
-
[27]
Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. 2024. https://openreview.net/forum?id=6FXtu8clyp Prismatic vlms: Investigating the design space of visually-conditioned language models . In Forty-first International Conference...
2024
-
[28]
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard H. Hovy. 2017. https://doi.org/10.18653/V1/D17-1082 RACE: large-scale reading comprehension dataset from examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP ...
2017 doi
- [29]
-
[30]
Jaewoo Lee, Boyang Li, and Sung Ju Hwang. 2024. Concept-skill transferability-based data selection for large vision-language models. arXiv preprint arXiv:2406.10995
2024 arXiv
-
[31]
Bo Li, Hao Zhang, Kaichen Zhang, Dong Guo, Yuanhan Zhang, Renrui Zhang, Feng Li, Ziwei Liu, and Chunyuan Li. 2024 a . https://llava-vl.github.io/blog/2024-05-25-llava-next-ablations/ Llava-next: What else influences visual instruction tuning beyond data?
2024
- [32]
-
[33]
Chen Li, Yixiao Ge, Dian Li, and Ying Shan. 2024 c . Vision-language instruction tuning: A review and analysis. Transactions on Machine Learning Research
2024
-
[34]
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023 a . https://arxiv.org/abs/2306.00890 Llava-med: Training a large language-and-vision assistant for biomedicine in one day . Preprint, arXiv:2306.00890
2023 arXiv
-
[35]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023 b . https://proceedings.mlr.press/v202/li23q.html BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models . In International Conference on Machine Learning, ICML 20...
2023
-
[36]
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji - Rong Wen. 2023 c . https://doi.org/10.18653/V1/2023.EMNLP-MAIN.20 Evaluating object hallucination in large vision-language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...
2023 doi
-
[37]
https://doi.org/10.48550/arXiv.2501.14818 Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Zhiqi Li, Guo Chen, Shilong Liu, Shihao Wang, Vibashan VS, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Subhashree Radhakrishnan, Nadine Chang, Karan Sapra, Amala Sanjay Deshmukh, Tuomas Rintamaki, Matthieu Le, Ilia Karmanov, Lukas Voegtle, Philipp Fischer, De-An Huang, and 8 ...
- [38]
- [39]
-
[40]
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024 a . https://llava-vl.github.io/blog/2024-01-30-llava-next/ Llava-next: Improved reasoning, ocr, and world knowledge
2024
-
[41]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 b . http://papers.nips.cc/paper\_files/paper/2023/hash/6dcf277ea32ce3288914faf369fe6de0-Abstract-Conference.html Visual instruction tuning . In Advances in Neural Information Processing Systems 36: Annual Conference...
2023
- [42]
-
[43]
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Yike Yuan, Wangbo Zhao, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2024 b . Mmbench: Is your multi-modal model an all-around player? In Computer Vision--ECCV 2024: 18th European Conference, Milan, I...
2024
- [44]
-
[45]
Zikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao, Yaliang Li, and Ji-Rong Wen. 2024 c . Less is more: Data value estimation for visual instruction tuning. arXiv preprint arXiv:2403.09559
2024 arXiv
-
[46]
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai - Wei Chang, Song - Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. http://papers.nips.cc/paper\_files/paper/2022/hash/11332b6b6cf4485b84afadb1352d3a9a-Abstract-Conference.html Learn to explain: Multimodal rea...
2022
-
[47]
Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen, Xiaoshuai Sun, and Rongrong Ji. 2024. Cheap and quick: Efficient vision-language instruction tuning for large language models. Advances in Neural Information Processing Systems, 36
2024
- [48]
-
[49]
Meta AI . 2024. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
2024
-
[50]
OpenAI. 2024. Hello gpt-4
2024
-
[51]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[52]
Maxime Oquab, Timoth \'e e Darcet, Th \'e o Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, and 1 others. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193
2023 arXiv
-
[53]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing sys...
2022
-
[54]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. http://proceedings.mlr.press/v139/radford21a.html Learning transferable visual models...
2021
-
[55]
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Padfield, James Qin, Danny Rozenberg, Tara Sainath, Johan Schalkwyk, Matt Sharif...
2023 arXiv
-
[56]
Patel, and Shao-Yuan Lo
Bardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel, and Shao-Yuan Lo. 2025. https://doi.org/10.48550/arXiv.2503.07591 Filter Images First , Generate Instructions Later : Pre-Instruction Data Selection for Visual Instruction Tuning . Preprint, arXiv:2503.07591
-
[57]
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie - Yan Liu. 2020. https://proceedings.neurips.cc/paper/2020/hash/c3a690be93aa602ee2dc0ccab5b7b67e-Abstract.html Mpnet: Masked and permuted pre-training for language understanding . In Advances in Neural Information Processing S...
2020
-
[58]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/V1/N19-1421 Commonsenseqa: A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North American Chapter of the Association...
2019 doi
-
[59]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, and 1 others. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805
2023 arXiv
- [60]
- [61]
-
[62]
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 a . Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560
2022 arXiv
-
[63]
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, and 1 others. 2022 b . Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp ...
2022
-
[64]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. https://openreview.net/forum?id=gEZrGCozdqR Finetuned language models are zero-shot learners . In The Tenth International Conference on Learning Repr...
2022
-
[65]
Lai Wei, Zihao Jiang, Weiran Huang, and Lichao Sun. 2023. Instructiongpt-4: A 200-instruction paradigm for fine-tuning minigpt-4. arXiv preprint arXiv:2308.12067
2023 arXiv
-
[66]
Junda Wu, Xintong Li, Tong Yu, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Jingbo Shang, and Julian McAuley. 2024. Commit: Coordinated instruction tuning for multimodal large language models. arXiv preprint arXiv:2407.20454
2024 arXiv
-
[68]
Zhiyang Xu, Chao Feng, Rulin Shao, Trevor Ashby, Ying Shen, dingnan jin, Yu Cheng, Qifan Wang, and Lifu Huang. 2024. https://api.semanticscholar.org/CorpusID:267750488 Vision-flan: Scaling human-labeled tasks in visual instruction tuning . In Annual Meeting of the Association ...
2024
-
[69]
Zhiyang Xu, Ying Shen, and Lifu Huang. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.641 Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2023 doi
-
[70]
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, and 1 others. 2023. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178
2023 arXiv
-
[71]
Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang. 2024. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2024
-
[72]
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549
2023 arXiv
-
[73]
Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, and 3 others. 2024. https://doi.org/10.110...
2024
-
[74]
Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong, and Ruihua Song. 2024. https://doi.org/10.18653/V1/2024.NAACL-LONG.440 What matters in training a gpt4-style language model with multimodal inputs? In Proceedings of the 2024 Conf...
2024 doi
-
[75]
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023 a . https://doi.org/10.18653/v1/2023.findings-emnlp.1055 S peech GPT : Empowering large language models with intrinsic cross-modal conversational abilities . In Findings of the Associati...
2023 doi
- [76]
-
[77]
Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023 b . Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199
2023 arXiv
-
[78]
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and 1 others. 2023 c . Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792
2023
-
[79]
Yi-Kai Zhang, Shiyin Lu, Yang Li, Yanqing Ma, Qingguo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, and Han-Jia Ye. 2024 b . Wings: Learning multimodal llms without text-only forgetting. Advances in Neural Information Processing Systems, 37:31828--31853
2024
-
[80]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstr...
2023
-
[81]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592
2023 arXiv
-
[82]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[83]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.