Pith. sign in

REVIEW 3 major objections 7 minor 1 cited by

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models

T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that instruction tuning of multimodal language models with a text-heavy mixture (75% text-only, 25% vision-language) under a fixed 186,000-instance budget matches or outperforms vision-heavy mixtures on both text and…

desk verdict Useful controlled sweep on modality ratios with token accounting; the vision-parity claim needs seeds and a validation split before it's believed. read the letter →

arxiv 2411.10557 v3 pith:TMM7B6CX submitted 2024-11-15 cs.CL

classification cs.CL
keywords instructiontuningmultimodallargelanguagemodelszero-shotgeneralizationtext-onlydatavision-languagetransfertrainingefficiencymixturecatastrophicforgetting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that, after multimodal pretraining, the instruction tuning stage of a multimodal language model does not need to be dominated by image-text data. It proposes MLAN, a fixed-budget mixture that is 75% text-only and 25% vision-language instruction data, trained on 186,000 instances sampled from Super-NaturalInstructions and Vision-Flan. Across 12 held-out text and vision benchmarks, this text-heavy mixture matches or slightly beats the vision-heavy mixtures used by LLaVA-1.5 and Cambrian-1 on two Llama-based models, while seeing fewer than half the images and processing up to about half the training tokens. The paper's explanation is that instruction-following ability and domain knowledge are mostly language-borne once vision-language alignment is done, so a small amount of vision data suffices for grounding. If correct, this would make multimodal instruction tuning substantially cheaper and more reliant on diverse text-only task supervision.

What carries the argument

The load-bearing mechanism is task-level semantic alignment between text-only and vision-language instruction data. The paper samples 100,000 instruction prompts from Super-NaturalInstructions and Vision-Flan, embeds them with a pretrained sentence transformer, and reports a significantly non-negative mean cosine similarity between the two modalities, reasoning that tasks are defined by their instructions and that shared task semantics transfer once vision-language pretraining aligns image tokens with text tokens. The method itself is a controlled training recipe: a fixed 186,000-instance budget, FLAN-style formatting, a CLIP-ViT-L/14@336 visual encoder with a two-layer MLP projector, and a fully unfrozen LLM, with only the data composition varied across experiments.

What would settle it

Evaluate MLAN on vision tasks with no close text-only analogue, such as OCR, fine-grained object grounding, or object counting in images, under the same 186,000-instance budget; if the 75% text-heavy mixture falls clearly below the vision-heavy Cambrian-1 mixture on those benchmarks, the claim that language-based tuning generally preserves and transfers vision knowledge is falsified.

Watch

Extended reading notes

Core claim

The central discovery is that instruction tuning of a multimodal LLM can be re-oriented around text without losing vision performance. Under a fixed budget of 186,000 training instances, the MLAN mixture of 75% text-only and 25% vision-language instructions yields held-out zero-shot performance on both modalities that is on par with or better than vision-heavy recipes (MIX-LLaVA-1.5 with 6% text and MIX-Cambrian-1 with 25% text). On the language benchmarks, MLAN averages 64.50 versus 63.18 for MIX-Cambrian-1 on Llama-3.2-3B and 71.57 versus 67.76 on Llama-3.1-8B; on vision benchmarks it averages 59.13 versus 58.57 on the 3B model and trails MIX-Cambrian-1 by 0.33 points on the 8B model. The text-heavy model processes 60.1 million training tokens, compared with 101.5 million for the Cambrian-1 mix and 117.2 million for the LLaVA-1.5 mix, a reduction of roughly 40% to 49%. The paper argues the transfer is possible because text-only and vision-language instructions share task-level semantics, and its controlled ablation shows that even 12.5% text-only data sharply raises both text and vision scores.

Load-bearing premise

The transfer mechanism assumes that cosine similarity between embedded text-only and vision-language instruction prompts is a reliable proxy for whether abilities learned from text-only data will transfer to image-grounded tasks; if that similarity is not the right proxy, the motivation for the text-heavy mixture weakens even if the fixed-budget empirical results still hold.

Editorial extensions

If this is right

  • A fixed-budget instruction tuning set can be made 75% text-only without sacrificing held-out vision performance, on both a 3B and an 8B Llama-based multimodal model.
  • Vision-heavy instruction mixtures (6% to 25% text) erode language knowledge on datasets like CommonsenseQA and CosmosQA by up to 20 percentage points, while the text-heavy mixture largely avoids that degradation.
  • Because CLIP converts each image into 576 visual tokens, replacing vision instances with text instances at the same instance budget cuts the number of training tokens processed by roughly half.
  • The mixture ratio is not arbitrary: 12.5% text-only data already produces a sharp gain on both text and vision axes, and vision performance peaks at a moderate language share, showing that neither pure modality is sufficient.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If text-heavy tuning generalizes beyond the two models tested here, the cost bottleneck of multimodal instruction tuning shifts from collecting image-text pairs to assembling diverse text-only task mixtures; a testable extension is to select text-only tasks whose instruction embeddings are most similar to a target vision benchmark and measure the resulting vision score.
  • The cosine-similarity analysis implies a practical data-selection tool: embed candidate text-only and vision-language instructions, then choose cross-modally similar subsets instead of fixing ratios by hand.
  • The asymmetric forgetting pattern — text abilities erode under vision-heavy tuning while vision abilities improve under text tuning — suggests language knowledge is the fragile resource in multimodal models, a prediction that could be checked on other architectures and other pretraining corpora.
  • The paper's own limitation section notes that OCR, captioning, and other specialized out-of-distribution vision tasks were not evaluated; a natural next test is whether the 75/25 mixture holds up when those tasks are added to the benchmark suite.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper proposes MLAN, a visual instruction tuning strategy that replaces most vision-language instruction data with text-only data (75% text, 25% vision) under a fixed 186k-instance budget after multimodal pretraining. Using LLaVA-style models built on Llama-3.2-3B and Llama-3.1-8B, the authors compare MLAN against simplified mixtures with LLaVA-1.5-style (6% text) and Cambrian-1-style (25% text) ratios on seven text-only and five vision-language held-out benchmarks. The headline finding is that MLAN matches or exceeds these vision-heavy proxy mixtures on vision benchmarks, clearly improves text benchmarks, and processes roughly 40-50% fewer training tokens. Additional experiments explore the language-data ratio curve, task diversity, pretraining data, and instruction-tuned versus base backbones.

Significance. If the result is robust to seed variation and does not come from selection on the evaluation benchmarks, this is a valuable and cost-saving empirical contribution. The paper addresses an underexplored design choice (modality composition in MLLM instruction tuning) with controlled fixed-budget comparisons at two model scales, includes token-level efficiency accounting, and provides several ablations. It is also candid about limitations, including the narrow architecture scope and the absence of specialized tasks such as OCR and captioning. However, the central 'matches or exceeds' claim currently rests on small vision-side margins from single runs, and the chosen 75% ratio appears to have been selected from the same evaluation curves used in the main tables. These issues need to be addressed before the quantitative conclusion can be considered reliable.

major comments (3)
  1. [§3.1, Tables 1–2] All results in the main comparison are single runs, with no error bars, standard deviations, or repeated-seed analysis reported anywhere in the manuscript. The vision-side difference that supports 'matches or exceeds' is extremely small for the larger model: MLAN's vision average is 62.25 versus 62.58 for MIX-Cambrian-1 in Table 1, a 0.33-point deficit, while individual benchmarks trade by several points (MMMU 34.44 vs 36.00; MMBench 72.51 vs 73.50; POPE 81.84 vs 82.57). Under sampling or optimization noise, that deficit could easily become a multi-point gap, which would change the conclusion from 'matches or exceeds' to 'slightly worse on vision with large text gains.' Please provide at least three seeds with error bars or confidence intervals, or explicitly restrict the claim to 'comparable on vision' with the uncertainty stated. The Limitations section does not currently address this single-run variance.
  2. [§3.3, Figure 3] The paper selects the 75% text-only ratio for MLAN after inspecting the knowledge-transfer curves in Figure 3, which are computed on the same held-out evaluation benchmarks used in Tables 1 and 2; no separate validation split is described. If the ratio was chosen by peeking at these curves, the reported comparison is optimistically selected rather than a prediction, and the 0.33-point vision deficit could reflect overfitting to the chosen evaluation suite. Please either document that the ratio was fixed before evaluation, or introduce a validation split (for example, hold out a subset of the 12 benchmarks for ratio selection and report the remaining benchmarks as the headline results).
  3. [§3.1, Table 3 and Appendix D.1] The baselines called 'MIX-LLaVA-1.5' and 'MIX-Cambrian-1' are not the actual LLaVA-1.5 or Cambrian-1 instruction tuning recipes: they are simplified mixtures sampled from the same two datasets (Super-NaturalInstructions and Vision-Flan) with 6% and 25% text-only ratios at a fixed 186k-instance budget. Actual LLaVA-1.5 uses 665k instances and Cambrian-1 uses millions of instances with different data sources, as Table 8 itself shows. Therefore, the conclusion that MLAN 'matches or better performance' on downstream vision-language tasks compared with these state-of-the-art recipes overstates what is directly supported. Please either compare against the real recipes (ideally at matched token budgets) or rename the baselines to 'proxy mixtures with LLaVA-1.5/Cambrian-1 text ratios' and avoid the state-of-the-art claim.
minor comments (7)
  1. [Table 1, Appendix A, Table 7] The treatment of ScienceQA is inconsistent: the Table 1 note says it is 'included in Vision-Flan but excluded in experiments,' whereas Appendix A says it is removed from the training set for evaluation, and Table 7 lists it as an evaluation benchmark. Please clarify whether it was excluded from training or from evaluation, or both, and correct the wording accordingly.
  2. [§2.1, Figure 4] The statistical tests for cosine similarity report only p-values; please also report effect sizes and confidence intervals, since with 100k samples a 'significantly non-negative mean' is a weak statement that does not quantify how similar the instructions actually are.
  3. [§3.3, Figure 3] The knowledge-transfer curve would be more informative with per-benchmark values and error bars; currently the text 'peaks and then slightly declines' cannot be quantitatively verified from the figure.
  4. [§3.4, Table 4] Because most production MLLMs use an instruction-tuned chat backbone, the 'Instruct LLM' row is an important caveat: the main comparison in Tables 1–2 uses non-instruction base models, so the practical significance of MLAN for standard recipes remains unclear. Please discuss whether the main conclusions hold with chat backbones.
  5. [Appendix D.1, Table 8] The Cambrian-1 row in Table 8 is garbled ('Cambrian-1 (Tong et al., 2024) – Cambrian-7M 1.68M ∼7M 23.8%'); please fix the formatting so the dataset size and text-only percentage are unambiguous.
  6. [§2.3, Tables 6–7] The text states '12 comprehensive benchmarks' while Tables 6–7 list 13 datasets because ARC-E and ARC-C are reported separately; please align the count.
  7. [Appendix B / Data Availability] No code, configuration, or checkpoint release is mentioned; given that the headline depends on exact data sampling and the 75% ratio, releasing the data mixture and training configuration would substantially improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: MLAN is an empirical data-mixture study judged against external held-out benchmarks, and no derivation reduces to its own inputs.

full rationale

The paper's central claim is empirical: under a fixed 186,000-instance budget, a 75% text-only / 25% vision-language instruction tuning mixture matches or exceeds vision-heavy mixtures on held-out text and vision benchmarks. This claim is tested against external datasets (ARC, MMLU, POPE, MMMU, MME, MMBench, etc.) and against two baseline mixtures, MIX-LLaVA-1.5 and MIX-Cambrian-1, which are reproduced under the same training protocol. MLAN is not derived from the evaluation scores; the 75% ratio is a dataset-composition choice reported as a method, and the main results are direct benchmark measurements. The cosine-similarity analysis in Section 2.1 is motivational rather than derivational: it supports the hypothesis that instruction semantics are shared across modalities, but the transfer claim is not computed from those similarity scores. There is no equation in which the output is identical to an input by construction, and no parameter is fitted to a subset of data and then renamed as a prediction. The paper does not rely on a load-bearing self-citation or an imported uniqueness theorem; its architecture and training recipe follow standard LLaVA practice with external references. The most legitimate concern is that the 75% ratio appears to have been selected after inspecting the knowledge-transfer curve in Figure 3 on the same evaluation benchmarks used in the main tables, with no validation split or repeated seeds described. That is a selection-bias and reproducibility concern about the strength of the 'matches or exceeds' claim, not a definitional or constructional circularity, because the reported numbers are not forced to equal the curve by construction. Accordingly, the appropriate finding is no significant circularity.

Assumptions & free parameters 2 free parameters · 2 assumptions · 0 invented entities

MLAN is a data mixture recipe rather than a new physical or architectural entity. The central result rests on public datasets and standard LLaVA-style components; the main assumed inputs are the transfer hypothesis and the fixed training budget.

free parameters (2)
  • Text-only ratio for MLAN = 75% text-only, 25% vision-language
    The headline mixture was selected after observing the ratio sweep in Figure 3 on evaluation benchmarks, not on a separate validation set.
  • Instruction tuning instance budget = 186,000
    A fixed budget chosen by the authors to make comparisons fair; the absolute size is arbitrary and may not transfer to larger training regimes.
assumptions (2)
  • domain assumption Strong visual pretraining alignment makes image tokens functionally similar to text tokens for instruction following
    Stated in Sections 2.1 and 2.2 as the reason text-only instruction tuning can transfer to vision; not directly measured except through downstream results.
  • domain assumption Cosine similarity of instruction embeddings (all-mpnet-base-v2) is a valid measure of cross-modal task transferability
    Used in Section 2.1 and Figure 2 to motivate MLAN; similarity of instruction text is treated as evidence about transfer of learned abilities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models." pith.science (2026). https://pith.science/paper/TMM7B6CX

@misc{pith2026241110557,
  author       = {Pith},
  title        = {Pith review of: MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TMM7B6CX}},
  note         = {Machine review of arXiv:2411.10557}
}
read the original abstract

We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base. Existing work lacks sufficient experimentation on the importance of each modality in the instruction tuning stage, often using a majority of vision-language data while keeping text-only data limited and fixing mixtures of modalities. By incorporating diverse text-only data in the visual instruction tuning stage, we vary vision-language data in various controlled experiments to investigate the importance of modality in visual instruction tuning. Our comprehensive evaluation shows that the text-heavy instruction tuning approach is able to perform on-par with traditional vision-heavy mixtures on both modalities across 12 general datasets while using as low as half the total training tokens. We find that simply increasing sufficiently diverse text-only data enables transfer of instruction following ability and domain knowledge across modalities while being more efficient than the vision-language approach.

Figures

Figures reproduced from arXiv: 2411.10557 by the authors.

Figure 1
Figure 1. Overview of MLAN. (a) MLAN represents a shift in perspective towards text during instruction tuning. After vision-language pretraining, we include diverse text-only data in our instruction tuning mixture spanning many tasks. We emphasize including text-only data to show the transferability of instruction tuning across modalities. For evaluation, we select ample text-only and vision-language datasets, allowing us to … view at source ↗
Figure 2
Figure 2. Similarity between text-only and vision-language instruction tuning data shown both (a) quantita￾tively with similarity scores and (b) qualitatively with examples. 100k instructions are sampled from the Super-NaturalInstructions (Wang et al., 2022b) and Vision-Flan (Xu et al., 2024) datasets and embedded by a pretrained sentenceTransformer, all-mpnet-base-v2 (Song et al., 2020). The red vertical line denotes the mea… view at source ↗
Figure 3
Figure 3. Average scores on Llama-3.2-3B based MLLMs with respect to the percentage of language data mixed in. The percentage denotes the amount of language data. Base LLM Variant Text Avg. Vision Avg. Llama-3.2-3B +MLAN 64.50 59.13 +Instruct LLM 67.74 60.98 -25% tasks 65.68 55.71 -50% tasks 65.98 55.73 -75% tasks 66.35 56.79 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Distribution of the cosine similarity of [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    POINTS-Seeker-8B is an 8B multimodal model trained from scratch for agentic search that uses seeding and visual-space history folding to outperform prior models on six visual reasoning benchmarks.

Reference graph

Works this paper leans on

82 extracted references · 15 canonical work pages · cited by 1 Pith paper

  1. [1]

    Jiang, Kartik Khandelwal, Timoth \'e e Lacroix, Guillaume Lample, Diego Las Casas, Thibaut Lavril, and 23 others

    Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, Soham Ghosh, Am \'e lie H \'e liou, Paul Jacob, Albert Q. Jiang, Kartik Khandelwal, Timoth \'e e Lacroix, Guillaume Lample, Diego Las Casas, Thibaut Lavril, and 23 others. 2024. https://...

  2. [2]

    Menick, Sebastian Borgeaud, and 8 others

    Jean - Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L. Menick, Sebastian Borgeaud, and 8 others. 2022. http://papers.nips.cc/paper\_files/paper...

  3. [3]

    Kirolos Ataallah, Xiaoqian Shen, Eslam Abdelrahman, Essam Sleiman, Deyao Zhu, Jian Ding, and Mohamed Elhoseiny. 2024. https://arxiv.org/abs/2404.03413 Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens . Preprint, arXiv:2404.03413

  4. [4]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. https://doi.org/10.48550/ARXIV.2308.12966 Qwen-vl: A frontier large vision-language model with versatile abilities . CoRR, abs/2308.12966

  5. [5]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025. https://doi.org/10.48550/arXiv.2502.13923 Qwen2.5- VL Technical Report . Preprint, arXiv:2502.13923

  6. [6]

    Sumithra Bhakthavatsalam, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, and Peter Clark. 2021. https://arxiv.org/abs/2102.03315 Think you have solved direct-answer question answering? try arc-da, the direct-answer AI2 reasoning challenge . CoRR, abs/2102.03315

  7. [7]

    Jinhe Bi, Yifan Wang, Danqi Yan, Xun Xiao, Artur Hecker, Volker Tresp, and Yunpu Ma. 2025. https://doi.org/10.48550/arXiv.2502.12119 PRISM : Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection . Preprint, arXiv:2502.12119

  8. [8]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://arxiv.org/abs/1911.11641 PIQA : Reasoning about physical commonsense in natural language . In Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence

Show all 82 references
  1. [9]

    Cheng Chen, Junchen Zhu, Xu Luo, Hengtao Shen, Lianli Gao, and Jingkuan Song. 2024 a . Coin: A benchmark of continual instruction tuning for multimodel large language model. arXiv preprint arXiv:2403.08350

  2. [10]

    Delong Chen, Jianfeng Liu, Wenliang Dai, and Baoyuan Wang. 2024 b . https://doi.org/10.1609/AAAI.V38I16.29727 Visual instruction tuning with polite flamingo . In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applicat...

  3. [11]

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023 a . https://doi.org/10.48550/ARXIV.2311.12793 Sharegpt4v: Improving large multi-modal models with better captions . CoRR, abs/2311.12793

  4. [12]

    Ruibo Chen, Yihan Wu, Lichang Chen, Guodong Liu, Qi He, Tianyi Xiong, Chenxi Liu, Junfeng Guo, and Heng Huang. 2024 c . Your vision-language model itself is a strong filter: Towards high-quality instruction tuning with data selection. arXiv preprint arXiv:2402.12501

  5. [13]

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, and 23 others. 2025. https://doi.org/10.48550/arXiv...

  6. [14]

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2023 b . https://doi.org/10.48550/ARXIV.2312.14238 Internvl: Scaling up vision foundation models and alignin...

  7. [15]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\

  8. [16]

    Christopher Clark, Kenton Lee, Ming - Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. https://doi.org/10.18653/V1/N19-1300 Boolq: Exploring the surprising difficulty of natural yes/no questions . In Proceedings of the 2019 Conference of the North Ame...

  9. [17]

    Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2025. https://doi.org/10.48550/arXiv.2503.15621 LLaVA-MORE : A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning . Prepri...

  10. [18]

    Wenliang Dai, Nayeon Lee, Boxin Wang, Zhuoling Yang, Zihan Liu, Jon Barker, Tuomas Rintamaki, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. https://arxiv.org/abs/2409.11402 Nvlm: Open frontier-class multimodal llms . Preprint, arXiv:2409.11402

  11. [19]

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/9a6a435e75419a836fe47ab6793623e6-Abstract-Conference.html Instructblip: Towards gener...

  12. [20]

    Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, Taira Anderson, Erin Bransom, Kiana Ehsani, Huong Ngo, YenSung Chen, Ajay Patel, Mark Yatskar, Chris Callison-Burch ...

  13. [21]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint...

  14. [22]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  15. [23]

    Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, and Rongrong Ji. 2023. https://doi.org/10.48550/ARXIV.2306.13394 MME: A comprehensive evaluation benchmark for multimodal large language mo...

  16. [24]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300

  17. [25]

    Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. https://doi.org/10.18653/V1/D19-1243 Cosmos QA: machine reading comprehension with contextual commonsense reasoning . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing...

  18. [26]

    Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Nils Johan Bertil Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. 2023. http://papers.nips...

  19. [27]

    Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. 2024. https://openreview.net/forum?id=6FXtu8clyp Prismatic vlms: Investigating the design space of visually-conditioned language models . In Forty-first International Conference...

  20. [28]

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard H. Hovy. 2017. https://doi.org/10.18653/V1/D17-1082 RACE: large-scale reading comprehension dataset from examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP ...

  21. [29]

    Hugo Lauren c on, L \' e o Tronchon, Matthieu Cord, and Victor Sanh. 2024. https://doi.org/10.48550/ARXIV.2405.02246 What matters when building vision-language models? CoRR, abs/2405.02246

  22. [30]

    Jaewoo Lee, Boyang Li, and Sung Ju Hwang. 2024. Concept-skill transferability-based data selection for large vision-language models. arXiv preprint arXiv:2406.10995

  23. [31]

    Bo Li, Hao Zhang, Kaichen Zhang, Dong Guo, Yuanhan Zhang, Renrui Zhang, Feng Li, Ziwei Liu, and Chunyuan Li. 2024 a . https://llava-vl.github.io/blog/2024-05-25-llava-next-ablations/ Llava-next: What else influences visual instruction tuning beyond data?

  24. [32]

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024 b . https://doi.org/10.48550/arXiv.2408.03326 LLaVA-OneVision : Easy Visual Task Transfer . Preprint, arXiv:2408.03326

  25. [33]

    Chen Li, Yixiao Ge, Dian Li, and Ying Shan. 2024 c . Vision-language instruction tuning: A review and analysis. Transactions on Machine Learning Research

  26. [34]

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023 a . https://arxiv.org/abs/2306.00890 Llava-med: Training a large language-and-vision assistant for biomedicine in one day . Preprint, arXiv:2306.00890

  27. [35]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023 b . https://proceedings.mlr.press/v202/li23q.html BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models . In International Conference on Machine Learning, ICML 20...

  28. [36]

    Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji - Rong Wen. 2023 c . https://doi.org/10.18653/V1/2023.EMNLP-MAIN.20 Evaluating object hallucination in large vision-language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Langua...

  29. [37]

    https://doi.org/10.48550/arXiv.2501.14818 Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

    Zhiqi Li, Guo Chen, Shilong Liu, Shihao Wang, Vibashan VS, Yishen Ji, Shiyi Lan, Hao Zhang, Yilin Zhao, Subhashree Radhakrishnan, Nadine Chang, Karan Sapra, Amala Sanjay Deshmukh, Tuomas Rintamaki, Matthieu Le, Ilia Karmanov, Lukas Voegtle, Philipp Fischer, De-An Huang, and 8 ...

  30. [38]

    Ji Lin, Hongxu Yin, Wei Ping, Yao Lu, Pavlo Molchanov, Andrew Tao, Huizi Mao, Jan Kautz, Mohammad Shoeybi, and Song Han. 2023. https://doi.org/10.48550/ARXIV.2312.07533 VILA: on pre-training for visual language models . CoRR, abs/2312.07533

  31. [39]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023 a . https://doi.org/10.48550/ARXIV.2310.03744 Improved baselines with visual instruction tuning . CoRR, abs/2310.03744

  32. [40]

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024 a . https://llava-vl.github.io/blog/2024-01-30-llava-next/ Llava-next: Improved reasoning, ocr, and world knowledge

  33. [41]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 b . http://papers.nips.cc/paper\_files/paper/2023/hash/6dcf277ea32ce3288914faf369fe6de0-Abstract-Conference.html Visual instruction tuning . In Advances in Neural Information Processing Systems 36: Annual Conference...

  34. [42]

    Yiyang Liu, James Chenhao Liang, Ruixiang Tang, Yugyung Lee, Majid Rabbani, Sohail Dianat, Raghuveer Rao, Lifu Huang, Dongfang Liu, Qifan Wang, and Cheng Han. 2025 a . https://doi.org/10.48550/arXiv.2503.00723 Re- Imagining Multimodal Instruction Tuning : A Representation View...

  35. [43]

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Yike Yuan, Wangbo Zhao, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2024 b . Mmbench: Is your multi-modal model an all-around player? In Computer Vision--ECCV 2024: 18th European Conference, Milan, I...

  36. [44]

    Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Vishwesh Nath, Jinyi Hu, Sifei Liu, and 8 others. 2025 b . https://doi....

  37. [45]

    Zikang Liu, Kun Zhou, Wayne Xin Zhao, Dawei Gao, Yaliang Li, and Ji-Rong Wen. 2024 c . Less is more: Data value estimation for visual instruction tuning. arXiv preprint arXiv:2403.09559

  38. [46]

    Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai - Wei Chang, Song - Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. http://papers.nips.cc/paper\_files/paper/2022/hash/11332b6b6cf4485b84afadb1352d3a9a-Abstract-Conference.html Learn to explain: Multimodal rea...

  39. [47]

    Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen, Xiaoshuai Sun, and Rongrong Ji. 2024. Cheap and quick: Efficient vision-language instruction tuning for large language models. Advances in Neural Information Processing Systems, 36

  40. [48]

    Brandon McKinzie, Zhe Gan, Jean - Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, Anton Belyi, Haotian Zhang, Karanjeet Singh, Doug Kang, Ankur Jain, Hongyu H \` e , Max Schwarzer, Tom Gunter, Xiang Kong, and 13 ...

  41. [49]

    Meta AI . 2024. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

  42. [50]

    OpenAI. 2024. Hello gpt-4

  43. [51]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  44. [52]

    Maxime Oquab, Timoth \'e e Darcet, Th \'e o Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, and 1 others. 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193

  45. [53]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing sys...

  46. [54]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. http://proceedings.mlr.press/v139/radford21a.html Learning transferable visual models...

  47. [55]

    Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Padfield, James Qin, Danny Rozenberg, Tara Sainath, Johan Schalkwyk, Matt Sharif...

  48. [56]

    Patel, and Shao-Yuan Lo

    Bardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel, and Shao-Yuan Lo. 2025. https://doi.org/10.48550/arXiv.2503.07591 Filter Images First , Generate Instructions Later : Pre-Instruction Data Selection for Visual Instruction Tuning . Preprint, arXiv:2503.07591

  49. [57]

    Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie - Yan Liu. 2020. https://proceedings.neurips.cc/paper/2020/hash/c3a690be93aa602ee2dc0ccab5b7b67e-Abstract.html Mpnet: Masked and permuted pre-training for language understanding . In Advances in Neural Information Processing S...

  50. [58]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/V1/N19-1421 Commonsenseqa: A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North American Chapter of the Association...

  51. [59]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, and 1 others. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805

  52. [60]

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie. 2024. https://doi.org/10.48550/ARXIV.2406.16860 Cambrian-1: A fully open, visi...

  53. [61]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton - Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu,...

  54. [62]

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 a . Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560

  55. [63]

    Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, and 1 others. 2022 b . Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp ...

  56. [64]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. https://openreview.net/forum?id=gEZrGCozdqR Finetuned language models are zero-shot learners . In The Tenth International Conference on Learning Repr...

  57. [65]

    Lai Wei, Zihao Jiang, Weiran Huang, and Lichao Sun. 2023. Instructiongpt-4: A 200-instruction paradigm for fine-tuning minigpt-4. arXiv preprint arXiv:2308.12067

  58. [66]

    Junda Wu, Xintong Li, Tong Yu, Yu Wang, Xiang Chen, Jiuxiang Gu, Lina Yao, Jingbo Shang, and Julian McAuley. 2024. Commit: Coordinated instruction tuning for multimodal large language models. arXiv preprint arXiv:2407.20454

  59. [68]

    Zhiyang Xu, Chao Feng, Rulin Shao, Trevor Ashby, Ying Shen, dingnan jin, Yu Cheng, Qifan Wang, and Lifu Huang. 2024. https://api.semanticscholar.org/CorpusID:267750488 Vision-flan: Scaling human-labeled tasks in visual instruction tuning . In Annual Meeting of the Association ...

  60. [69]

    Zhiyang Xu, Ying Shen, and Lifu Huang. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.641 Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...

  61. [70]

    Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, and 1 others. 2023. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178

  62. [71]

    Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan, Anwen Hu, Haowei Liu, Qi Qian, Ji Zhang, and Fei Huang. 2024. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  63. [72]

    Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. 2023. A survey on multimodal large language models. arXiv preprint arXiv:2306.13549

  64. [73]

    Xiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, and 3 others. 2024. https://doi.org/10.110...

  65. [74]

    Yan Zeng, Hanbo Zhang, Jiani Zheng, Jiangnan Xia, Guoqiang Wei, Yang Wei, Yuchen Zhang, Tao Kong, and Ruihua Song. 2024. https://doi.org/10.18653/V1/2024.NAACL-LONG.440 What matters in training a gpt4-style language model with multimodal inputs? In Proceedings of the 2024 Conf...

  66. [75]

    Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023 a . https://doi.org/10.18653/v1/2023.findings-emnlp.1055 S peech GPT : Empowering large language models with intrinsic cross-modal conversational abilities . In Findings of the Associati...

  67. [76]

    Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, Bowen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, and 4 others. 20...

  68. [77]

    Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. 2023 b . Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199

  69. [78]

    Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and 1 others. 2023 c . Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792

  70. [79]

    Yi-Kai Zhang, Shiyin Lu, Yang Li, Yanqing Ma, Qingguo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, and Han-Jia Ye. 2024 b . Wings: Learning multimodal llms without text-only forgetting. Advances in Neural Information Processing Systems, 37:31828--31853

  71. [80]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstr...

  72. [81]

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592

  73. [82]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  74. [83]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.