REVIEW 2 major objections 1 minor 1 cited by
DataComp-VLM: Improved Open Datasets for Vision-Language Models
T0 review · 2 major / 1 minor · reviewed 2026-07-02 · grok-4.3
Pith's one-line read Instruction-heavy data mixtures scale better than caption-heavy ones for vision-language model training.
desk verdict DCVLM gives a controlled benchmark and 6T-token corpus for VLM data curation, with evidence that instruction-heavy mixing beats filtering at scale, but the 33-task suite may over-weight the very capabilities that mixing targets. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The DCVLM benchmark, which standardizes curation operations (filtering, mixing, formatting, sampling) across fixed model sizes and token budgets and evaluates on a fixed suite of up to 52 downstream tasks in nine domains.
What would settle it
A caption-heavy mixture or a purely filtering-based curation strategy that achieves higher average accuracy than DCVLM-Baseline on the same 33-task core suite when trained at the 8B scale with 200B tokens.
Extended reading notes
Core claim
Data mixing, not filtering, is the dominant factor in building high-quality VLM training sets; instruction-heavy mixtures outperform caption-heavy ones, with the performance gap widening at larger model and data scales. The DCVLM-Baseline mixture derived from these experiments enables an 8B-parameter VLM trained on 200B tokens to reach 63.6 percent average accuracy across the 33-task core suite, a 5.4-point improvement over the previous state-of-the-art open VLM dataset FineVision.
Load-bearing premise
The selected downstream benchmarks adequately represent the full range of general VLM capabilities.
Editorial extensions
If this is right
- Instruction-heavy mixtures deliver increasing returns as model size and token count grow.
- Filtering alone yields smaller gains than careful composition of data types.
- The DCVLM-Baseline dataset can be used directly to train stronger open VLMs without proprietary data.
- Curation effort should prioritize mixing ratios over removal of individual low-quality examples.
Reading between the lines
- Future work could test whether the same mixing preference holds when the evaluation suite is expanded to include more reasoning-heavy or long-context tasks.
- The benchmark design makes it straightforward to measure whether new data sources improve performance mainly by changing the overall mixture balance.
- Practitioners building custom VLM datasets may achieve comparable gains by reweighting existing public collections toward instruction data rather than collecting new filtered captions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DataComp-VLM (DCVLM), a benchmark and corpus of 6T multimodal tokens from 160 datasets across four data types. It enables controlled experiments on curation operations (filtering, mixing, formatting, sampling) for 1B-8B VLMs trained on 6.25B-200B tokens, evaluated on up to 52 downstream benchmarks across 9 domains. The central empirical finding is that data mixing—not filtering—is the dominant factor, with instruction-heavy mixtures outperforming caption-heavy ones (gains widening at larger scales); the resulting DCVLM-Baseline yields an 8B VLM at 63.6% on the 33-task core suite (+5.4pp over FineVision).
Significance. If the results hold, the work supplies the first large-scale, controlled benchmark for VLM data curation and demonstrates that mixing strategies can be systematically optimized, with public release of the corpus, baseline, and evaluation suite as a concrete community resource. The scale of the experiments (multiple model sizes and token budgets) and the quantitative improvement over an existing SOTA open dataset are notable strengths.
major comments (2)
- [Evaluation] Evaluation section (and abstract): the claim that 'instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales' is supported only by relative performance on the 33-task core suite. The manuscript provides no evidence that this suite was constructed with held-out domains, that deltas were tested for sensitivity to task re-weighting, or that performance was measured on underrepresented categories such as pure captioning, OCR, or retrieval; this is load-bearing for the generalization that mixing is the key curation operation.
- [Experiments] Methods / Experiments: the abstract and results report specific accuracy gains (e.g., 63.6% and +5.4pp) without mention of error bars, multiple random seeds, or data-exclusion criteria for the downstream benchmarks. Because the central mixing-vs-filtering conclusion rests on these measured deltas, the absence of statistical characterization weakens verification of the reported improvements.
minor comments (1)
- [Abstract] The abstract states 'up to 52 downstream benchmarks' while the core suite is described as 33 tasks; clarify the exact overlap and selection criteria in the main text.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback. We address the two major comments point by point below, indicating where revisions will be made.
read point-by-point responses
-
Referee: [Evaluation] Evaluation section (and abstract): the claim that 'instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales' is supported only by relative performance on the 33-task core suite. The manuscript provides no evidence that this suite was constructed with held-out domains, that deltas were tested for sensitivity to task re-weighting, or that performance was measured on underrepresented categories such as pure captioning, OCR, or retrieval; this is load-bearing for the generalization that mixing is the key curation operation.
Authors: The 33-task core suite was explicitly selected to ensure coverage across all 9 evaluation domains (including captioning, OCR, retrieval, and VQA), with the full 52-task results provided in the appendix showing consistent trends. We will add an explicit subsection on task selection methodology, domain balance, and sensitivity analysis to re-weighting in the revised manuscript to further support the generalization. revision: partial
-
Referee: [Experiments] Methods / Experiments: the abstract and results report specific accuracy gains (e.g., 63.6% and +5.4pp) without mention of error bars, multiple random seeds, or data-exclusion criteria for the downstream benchmarks. Because the central mixing-vs-filtering conclusion rests on these measured deltas, the absence of statistical characterization weakens verification of the reported improvements.
Authors: We agree that additional statistical detail would strengthen the presentation. All experiments use fixed seeds for reproducibility; the scale of the 8B/200B-token runs made multiple independent seeds computationally prohibitive. We will add a limitations paragraph on this point, report error bars for all smaller-scale ablations, and expand the existing description of benchmark data-exclusion criteria in Section 4. revision: partial
Circularity Check
No circularity: empirical results on external benchmarks
full rationale
The paper's central claims rest on empirical measurements: models trained on curated mixtures are evaluated on a held-out suite of up to 52 downstream benchmarks across 9 domains. No equations, fitted parameters, or self-referential definitions appear in the derivation; the superiority of instruction-heavy mixing is reported as observed performance deltas (e.g., +5.4pp over FineVision), not as a quantity forced by the curation process itself. Self-citations, if present, are not load-bearing for the mixing-vs-filtering conclusion. The evaluation distribution is external to the training data construction, satisfying the condition for a self-contained empirical result.
Assumptions & free parameters
assumptions (1)
- standard math Standard i.i.d. sampling and gradient-based optimization assumptions used in large-scale model training.
Cite this review
Pith. "Pith review of DataComp-VLM: Improved Open Datasets for Vision-Language Models." pith.science (2026). https://pith.science/paper/HLF232VG
@misc{pith2026260628551,
author = {Pith},
title = {Pith review of: DataComp-VLM: Improved Open Datasets for Vision-Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/HLF232VG}},
note = {Machine review of arXiv:2606.28551}
}
read the original abstract
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B-8B models and 6.25B-200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-Baseline, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FineVision, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available at https://www.datacomp.ai/dcvlm/.
Figures
Figures from the paper (35 more)
Forward citations
Cited by 1 Pith paper
-
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
Caption quality can be split into Coverage and Precision; in controlled training runs, Coverage predicts VLM understanding while Precision predicts T2I generation.
Reference graph
Works this paper leans on
-
[1]
SemDeDup: Data-efficient learning at web-scale through semantic deduplication
A. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication.arXiv preprint arXiv:2303.09540, 2023. Cited on page 45
work page Pith review arXiv 2023
- [2]
-
[3]
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, J. Bao, A. Benhaim, M. Cai, V . Chaudhary, C. Chen, et al. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras.arXiv preprint arXiv:2503.01743, 2025. Cited on page 2
work page Pith review arXiv 2025
-
[4]
M. Acharya, K. Kafle, and C. Kanan. TallyQA: Answering complex counting questions. In AAAI Conference on Artificial Intelligence (AAAI), 2019. Cited on pages 53 and 54
work page 2019
-
[5]
L. Agnolucci, L. Galteri, M. Bertini, and A. Del Bimbo. Arniqa: Learning distortion manifold for image quality assessment. InIEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 189–198, 2024. Cited on pages 45 and 75
work page 2024
-
[6]
J. Ainslie, J. Lee-Thorp, M. De Jong, Y . Zemlyanskiy, F. Lebrón, and S. Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. InConference on Empirical Methods in Natural Language Processing (EMNLP), pages 4895–4901, 2023. Cited on page 48
work page 2023
- [7]
-
[8]
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS), 35:23716–23736, 2022. Cited on pages 2 and 45
work page 2022
Show all 299 references
-
[9]
L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíˇcek, A. P. Lajarín, V . Srivastav, et al. Smollm2: When smol goes big–data-centric training of a small language model.arXiv preprint arXiv:2502.02737, 2025. Cited on pages 6 and 46
2025 arXiv
-
[10]
Allen-Zhu and Y
Z. Allen-Zhu and Y . Li. Physics of language models: Part 3.1, knowledge storage and extraction.arXiv preprint arXiv:2309.14316, 2023. Cited on page 8
2023
-
[11]
Amini, S
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y . Choi, and H. Hajishirzi. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In J. Burstein, C. Doran, and T. Solorio, editors,Proceedings of the 2019 Conference of the North American ...
2019 doi
-
[12]
X. An, Y . Xie, K. Yang, W. Zhang, X. Zhao, Z. Cheng, Y . Wang, S. Xu, C. Chen, D. Zhu, et al. Llava-onevision-1.5: Fully open framework for democratized multimodal training.arXiv preprint arXiv:2509.23661, 2025. Cited on pages 2, 3, 9, and 45. 12
2025 arXiv
-
[13]
Ankner, C
Z. Ankner, C. Blakeney, K. Sreenivasan, M. Marion, M. L. Leavitt, and M. Paul. Perplexed by perplexity: Perplexity-based data pruning with small reference models.arXiv preprint arXiv:2405.20541, 2024. Cited on page 5
2024
-
[14]
Awadalla, L
A. Awadalla, L. Xue, O. Lo, M. Shu, H. Lee, E. Guha, M. Jordan, S. Shen, M. Awadalla, S. Savarese, et al. Mint-1t: Scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens.Advances in Neural Information Processing Systems (NeurIPS), 37: 36805–3...
2024
-
[15]
C. Baek, R. P. Monti, D. Schwab, A. Abbas, R. Adiga, C. Blakeney, M. Böther, P. Burstein, A. G. Carranza, A. Deng, et al. The finetuner’s fallacy: When to pretrain with your finetuning data.arXiv preprint arXiv:2603.16177, 2026. Cited on page 8
2026
-
[16]
S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025. Cited on pages 2, 3, and 4
2025 arXiv
-
[17]
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin. Qwen2.5-vl technical report, 2025. URLhttp...
2025 arXiv
-
[18]
Berasi, M
D. Berasi, M. Farina, M. Mancini, and E. Ricci. Linear model merging unlocks simple and scalable multimodal data mixture optimization.arXiv preprint arXiv:2602.04937, 2026. Cited on pages 3 and 45
2026
-
[19]
Bevli, S
A. Bevli, S. Chaybouti, Y . Dahou, H. Hacid, N. D. Huynh, P. H. L. Khac, S. Narayan, W. R. Para, and A. Singh. Falcon perception.arXiv preprint arXiv:2603.27365, 2026. Cited on page 2
2026
-
[20]
L. Beyer. On the speed of ViTs and CNNs.http://lb.eyer.be/a/vit-cnn-speed.html, 2024. Cited on page 68
2024
-
[21]
Beyer, A
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al. Paligemma: A versatile 3b vlm for transfer.arXiv preprint arXiv:2407.07726, 2024. Cited on pages 2, 45, and 46
2024 arXiv
-
[22]
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas. Scene text visual question answering. InIEEE/CVF International Conference on Computer Vision (ICCV), pages 4291–4301, 2019. Cited on pages 53 and 54
2019
-
[23]
Bordt, S
S. Bordt, S. Srinivas, V . Boreiko, and U. V on Luxburg. How much can we forget about data contamination?arXiv preprint arXiv:2410.03249, 2024. Cited on page 46
2024
-
[24]
Breuel and WebDataset Contributors
T. Breuel and WebDataset Contributors. WebDataset: A high-performance Python-based I/O system for large (and small) deep learning problems, with strong support for PyTorch. https://github.com/webdataset/webdataset, 2020. Cited on page 85
2020
-
[25]
A. Z. Broder. On the resemblance and containment of documents. InProceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), pages 21–29. IEEE, 1997. Cited on pages 4, 46, and 69. 13
1997
-
[26]
Cahyawijaya, H
S. Cahyawijaya, H. Lovenia, J. R. A. Moniz, T. H. Wong, M. R. Farhansyah, T. T. Maung, F. Hudi, D. Anugraha, M. R. S. Habibi, M. R. Qorib, et al. Crowdsource, crawl, or generate? creating sea-vl, a multicultural vision-language dataset for southeast asia. InProceedings of the ...
2025
-
[27]
Cao and J
J. Cao and J. Xiao. An augmented benchmark dataset for geometric question answering through dual parallel text encoding. InInternational Conference on Computational Linguistics (COLING), 2022. Cited on pages 53 and 54
2022
-
[28]
Carlini, D
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang. Quantifying memorization across neural language models. InInternational Conference on Learning Representations (ICLR), 2022. Cited on page 8
2022
-
[29]
J. Carter. TextOCR-GPT4V: A re-captioning of TextOCR with GPT-4V, 2024. Hugging Face dataset card,https://huggingface.co/datasets/jimmycarter/textocr-gpt4v. Cited on pages 53 and 54
2024
-
[30]
Chang, D
S. Chang, D. Palzer, J. Li, E. Fosler-Lussier, and N. Xiao. MapQA: A dataset for question answering on choropleth maps.arXiv preprint arXiv:2211.08545, 2022. Cited on pages 53 and 54
2022
-
[31]
Changpinyo, P
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3558–3568, 2021. Cited on page 45
2021
-
[32]
G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang. ALLaV A: Harnessing GPT4V-synthesized data for a lite vision-language model. arXiv preprint arXiv:2402.11684, 2024. Cited on pages 53 and 54
2024 arXiv
-
[33]
J. Chen, T. Li, J. Qin, P. Lu, L. Lin, C. Chen, and X. Liang. UniGeo: Unifying geometry logical reasoning via reformulating mathematical expression. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2022. Cited on pages 53 and 54
2022
-
[34]
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin. ShareGPT4V: Improving large multi-modal models with better captions. InEuropean Conference on Computer Vision (ECCV), pages 370–387. Springer, 2024. Cited on pages 53, 54, and 55
2024
-
[35]
L. Chen, J. Li, X. Dong, P. Zhang, Y . Zang, Z. Chen, H. Duan, J. Wang, Y . Qiao, D. Lin, et al. Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. Cited on page 66
2024
-
[36]
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021. Cited on page 66
2021 arXiv
-
[37]
M. F. Chen, T. Murray, D. Heineman, M. Jordan, H. Hajishirzi, C. Ré, L. Soldaini, and K. Lo. Olmix: A framework for data mixing throughout lm development.arXiv preprint arXiv:2602.12237, 2026. Cited on pages 3, 45, and 93
2026
-
[38]
W. Chen, M. Yin, M. Ku, P. Lu, Y . Wan, X. Ma, J. Xu, X. Wang, and T. Xia. Theoremqa: A theorem-driven question answering dataset. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7889–7901, 2023. Cited on page 66. 14
2023
-
[39]
X. Chen, H. Fang, T.-Y . Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft COCO captions: Data collection and evaluation server.arXiv preprint arXiv:1504.00325, 2015. Cited on pages 53, 54, and 66
2015 arXiv
-
[40]
Y . Chen, S. Qian, H. Tang, X. Lai, Z. Liu, S. Han, and J. Jia. LongloRA: Efficient fine-tuning of long-context large language models. InInternational Conference on Learning Representations (ICLR), 2024. URLhttps://openreview.net/forum?id=6PmJoRfdaK. Cited on pages 53 and 54
2024
-
[41]
Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271, 2024. Cited on pages 1, 4, 5, 47, 50, 52, 54, 6...
2024 arXiv
-
[42]
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. Cited on p...
2024
-
[43]
C. K. Chng, Y . Liu, Y . Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding, et al. ICDAR2019 robust reading challenge on arbitrary-shaped text - RRC-ArT. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53 and 54
2019
-
[44]
J. H. Cho, A. Madotto, E. Mavroudi, T. Afouras, T. Nagarajan, M. Maaz, Y . Song, T. Ma, S. Hu, S. Jain, et al. Perceptionlm: Open-access data and models for detailed visual understanding. arXiv preprint arXiv:2504.13180, 2025. Cited on pages 2 and 86
2025
-
[45]
Clark, J
C. Clark, J. Zhang, Z. Ma, J. S. Park, M. Salehi, R. Tripathi, S. Lee, Z. Ren, C. D. Kim, Y . Yang, et al. Molmo2: Open weights and data for vision-language models with video understanding and grounding.arXiv preprint arXiv:2601.10611, 2026. Cited on pages 3, 45, and 77
2026 arXiv
-
[47]
Cobbe, V
K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. Cited on page 66
2021 arXiv
-
[48]
Conover, M
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin. Free Dolly: Introducing the world’s first truly open instruction- tuned LLM, 2023. Databricks Blog https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially...
2023
-
[49]
Contributors
O. Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023. Cited on page 62
2023
-
[50]
M. R. Costa-Jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, et al. No language left behind: Scaling human-centered machine translation.arXiv preprint arXiv:2207.04672, 2022. Cited on page 59
2022 arXiv
-
[51]
E. Cui, Y . He, Z. Ma, Z. Chen, H. Tian, W. Wang, K. Li, Y . Wang, W. Wang, X. Zhu, L. Lu, T. Lu, Y . Wang, L. Wang, Y . Qiao, and J. Dai. Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024. URLhttps://sharegpt4o.github.io/. Cited on page 3. 15
2024
-
[52]
G. Cui, L. Yuan, N. Ding, G. Yao, B. He, W. Zhu, Y . Ni, G. Xie, R. Xie, Y . Lin, Z. Liu, and M. Sun. UltraFeedback: Boosting language models with scaled AI feedback.International Conference on Machine Learning (ICML), 2024. Cited on pages 53 and 54
2024
-
[53]
D. Dai, Y . Li, Y . Liu, M. Jia, Z. YuanHui, and G. Wang. 15M multimodal facial image-text dataset.arXiv preprint arXiv:2407.08515, 2024. Cited on pages 53 and 54
2024
-
[54]
T. Dao. Flashattention-2: Faster attention with better parallelism and work partitioning.arXiv preprint arXiv:2307.08691, 2023. Cited on page 47
2023 arXiv
-
[55]
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra. Visual dialog. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. Cited on pages 53 and 54
2017
-
[56]
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, et al. Large scale distributed deep networks.Advances in Neural Information Processing Systems (NeurIPS), 25, 2012. Cited on page 50
2012
-
[57]
Deitke, C
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y . Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al. Molmo and pixmo: Open weights and open data for state-of-the-art vision- language models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...
2025
-
[58]
A. S. Deshmukh, K. Chumachenko, T. Rintamaki, M. Le, T. Poon, D. M. Taheri, I. Karmanov, G. Liu, J. Seppanen, G. Chen, et al. Nvidia nemotron nano v2 vl.arXiv preprint arXiv:2511.03929, 2025. Cited on pages 2, 3, 9, 45, and 86
2025
-
[59]
S. Diao, Y . Yang, Y . Fu, X. Dong, D. Su, M. Kliegl, Z. Chen, P. Belcak, Y . Suhara, H. Yin, et al. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training.arXiv preprint arXiv:2504.13161, 2025. Cited on pages 3, 45, and 77
2025 arXiv
-
[60]
N. Ding, Y . Chen, B. Xu, Y . Qin, S. Hu, Z. Liu, M. Sun, and B. Zhou. Enhancing chat language models by scaling high-quality instructional conversations. In H. Bouamor, J. Pino, and K. Bali, editors,Conference on Empirical Methods in Natural Language Processing (EMNLP), pages...
2023
-
[61]
Dodge, M
J. Dodge, M. Sap, A. Marasovi ´c, W. Agnew, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. InProceedings of the 2021 conference on empirical methods in natural language processing, p...
2021
-
[62]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020. Cited on page 47
2010 arXiv
-
[63]
Douze, A
M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou. The faiss library.IEEE Transactions on Big Data, 2025. Cited on page 68
2025
-
[64]
H. Duan, J. Yang, Y . Qiao, X. Fang, L. Chen, Y . Liu, X. Dong, Y . Zang, P. Zhang, J. Wang, et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. InACM International Conference on Multimedia, pages 11198–11201, 2024. Cited on pages 2 and 62. 16
2024
-
[65]
Evans, N
T. Evans, N. Parthasarathy, H. Merzi ´c, and O. J. Henaff. Data curation via joint example selection further accelerates multimodal learning.Advances in Neural Information Processing Systems (NeurIPS), 37:141240–141260, 2024. Cited on pages 5 and 6
2024
-
[66]
L. Fan, D. Krishnan, P. Isola, D. Katabi, and Y . Tian. Improving clip training with language rewrites.Advances in Neural Information Processing Systems (NeurIPS), 36:35544–35575, 2023. Cited on page 45
2023
-
[67]
A. Fang, G. Ilharco, M. Wortsman, Y . Wan, V . Shankar, A. Dave, and L. Schmidt. Data determines distributional robustness in contrastive language image pre-training (clip). In International Conference on Machine Learning (ICML), pages 6216–6234. PMLR, 2022. Cited on page 1
2022
-
[68]
A. Fang, A. M. Jose, A. Jain, L. Schmidt, A. Toshev, and V . Shankar. Data filtering networks. arXiv preprint arXiv:2309.17425, 2023. Cited on pages 5 and 45
2023
-
[69]
A. Fang, H. Pouransari, M. Jordan, A. Toshev, V . Shankar, L. Schmidt, and T. Gunter. Datasets, documents, and repetitions: The practicalities of unequal data quality.arXiv preprint arXiv:2503.07879, 2025. Cited on page 8
2025
-
[70]
L. Feng, G. R. Ghosal, J. M. Springer, Z. Zhong, and A. Raghunathan. Early data exposure improves robustness to subsequent fine-tuning.arXiv preprint arXiv:2605.12705, 2026. Cited on page 8
2026 arXiv
-
[71]
C. Fu, P. Chen, Y . Shen, Y . Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models.arXiv preprint arXiv:2306.13394, 2023. Cited on pages 65 and 66
2023 arXiv
-
[72]
X. Fu, Y . Hu, B. Li, Y . Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W.-C. Ma, and R. Krishna. Blink: Multimodal large language models can see but not perceive. InEuropean Conference on Computer Vision, pages 148–166. Springer, 2024. Cited on page 66
2024
-
[73]
S. Y . Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al. Datacomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Systems (NeurIPS), 36:27092–27112, 2023. Cited o...
2023
-
[74]
L. Gao. An empirical exploration in quality filtering of text data.arXiv preprint arXiv:2109.00698, 2021. Cited on page 7
2021
-
[75]
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020. Cited on page 46
2020 arXiv
-
[76]
Gervais, A
P. Gervais, A. Fadeeva, and A. Maksai. Mathwriting: A dataset for handwritten mathematical expression recognition. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V .2, KDD ’25, page 5459–5469, New York, NY , USA, 2025. Association for Co...
2025 doi
-
[77]
Ghosal, V
D. Ghosal, V . T. Y . Han, C. Y . Ken, and S. Poria. Are language models puzzle prodigies? Algorithmic puzzles unveil serious challenges in multimodal reasoning.arXiv preprint arXiv:2403.03864, 2024. Cited on pages 53 and 54. 17
2024
-
[78]
Ghosh, S
A. Ghosh, S. Dziadzio, A. Prabhu, V . Udandarao, S. Albanie, and M. Bethge. Onebench to test them all: Sample-level benchmarking over open-ended capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pag...
2025
-
[79]
Ghosh, V
A. Ghosh, V . Udandarao, T. Nguyen, M. Farina, M. Cherti, J. Jitsev, S. Oh, E. Ricci, L. Schmidt, and M. Bethge. Concept-aware batch sampling improves language-image pretraining.arXiv preprint arXiv:2511.20643, 2025. Cited on pages 1, 45, 78, and 79
2025
-
[80]
Glaive-Code-Assistant, 2023
Glaive AI. Glaive-Code-Assistant, 2023. https://huggingface.co/datasets/ glaiveai/glaive-code-assistant. Cited on pages 53 and 54
2023
-
[81]
Goyal, P
S. Goyal, P. Maini, Z. C. Lipton, A. Raghunathan, and J. Z. Kolter. Scaling laws for data filtering–data curation cannot be compute agnostic. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22702–22711, 2024. Cited on pages 4, 7, 46, and 82
2024
-
[82]
Goyal, T
Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. Cited on pages 53 and 54
2017
-
[83]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024. Cited on page 1
2024 arXiv
-
[84]
T. Gu, Z. Zhou, K. Huang, D. Liang, Y . Wang, H. Zhao, Y . Yao, X. Qiao, K. Wang, Y . Yang, et al. Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models.Advances in Neural Information Processing Systems, 37:7256–7295, 2024. Cited on page 66
2024
-
[85]
T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y . Yacoob, et al. Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. InProceedings of the IEEE/CVF conference on com...
2024
-
[86]
E. Guha, R. Marten, S. Keh, N. Raoof, G. Smyrnis, H. Bansal, M. Nezhurina, J. Mercat, T. Vu, Z. Sprague, et al. Openthoughts: Data recipes for reasoning models.arXiv preprint arXiv:2506.04178, 2025. Cited on page 65
2025 arXiv
-
[87]
D. Guo, F. Wu, F. Zhu, F. Leng, G. Shi, H. Chen, H. Fan, J. Wang, J. Jiang, J. Wang, et al. Seed1. 5-vl technical report.arXiv preprint arXiv:2505.07062, 2025. Cited on page 2
2025 arXiv
-
[88]
H. Guo, X. Qin, J. Liu, J. Han, J. Liu, and E. Ding. EATEN: Entity-aware attention for single shot visual text extraction. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53 and 54
2019
-
[89]
J. Guo, T. Zheng, Y . Li, Y . Bai, B. Li, Y . Wang, K. Zhu, G. Neubig, W. Chen, and X. Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...
2025
-
[90]
Gupta, A
A. Gupta, A. Vedaldi, and A. Zisserman. Synthetic data for text localisation in natural images. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016. Cited on pages 53 and 54. 18
2016
-
[91]
Gurari, Q
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham. Vizwiz grand challenge: Answering visual questions from blind people. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–3617, 2018. Cited on page 66
2018
-
[92]
Y . Han, C. Zhang, X. Chen, X. Yang, Z. Wang, G. Yu, B. Fu, and H. Zhang. ChartLlama: A multimodal LLM for chart understanding and generation.arXiv preprint arXiv:2311.16483, 2023. Cited on pages 53 and 54
2023
-
[93]
Hanu and Unitary team
L. Hanu and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020. Cited on page 89
2020
-
[94]
C. He, Z. Jin, C. Xu, J. Qiu, B. Wang, W. Li, H. Yan, J. Wang, and D. Lin. Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models.arXiv preprint arXiv:2308.10755, 2023. Cited on pages 4, 53, 54, and 55
2023
-
[95]
M. He, Y . Liu, Z. Yang, S. Zhang, C. Luo, F. Gao, Q. Zheng, Y . Wang, X. Zhang, and L. Jin. ICPR 2018 contest on robust reading for multi-type web images (MTWI). InInternational Conference on Pattern Recognition (ICPR), 2018. Cited on pages 53 and 54
2018
-
[96]
X. He, Y . Zhang, L. Mou, E. Xing, and P. Xie. PathVQA: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020. Cited on pages 53 and 54
2003 arXiv
-
[97]
Heineman, V
D. Heineman, V . Hofmann, I. Magnusson, Y . Gu, N. A. Smith, H. Hajishirzi, K. Lo, and J. Dodge. Signal and noise: A framework for reducing uncertainty in language model evaluation.arXiv preprint arXiv:2508.13144, 2025. Cited on page 5
2025
-
[98]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020. Cited on page 66
2009 arXiv
-
[99]
Hendrycks, C
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset.arXiv preprint arXiv:2103.03874, 2021. Cited on page 66
2021 arXiv
-
[100]
Hernandez, T
D. Hernandez, T. Brown, T. Conerly, N. DasSarma, D. Drain, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, T. Henighan, T. Hume, et al. Scaling laws and interpretability of learning from repeated data.arXiv preprint arXiv:2205.10487, 2022. Cited on page 8
2022 arXiv
-
[101]
Hessel, A
J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y . Choi. Clipscore: A reference-free evaluation metric for image captioning. InConference on Empirical Methods in Natural Language Processing (EMNLP), pages 7514–7528, 2021. Cited on pages 3 and 45
2021
-
[102]
R. Hong, W. Agnew, T. Kohno, and J. Morgenstern. Who’s in and who’s out? a case study of multimodal clip-filtering in datacomp. InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–17, 2024. Cited on page 94
2024
-
[103]
W. Hong, W. Yu, X. Gu, G. Wang, G. Gan, H. Tang, J. Cheng, J. Qi, J. Ji, L. Pan, et al. Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning.arXiv preprint arXiv:2507.01006, 2025. Cited on page 2. 19
2025 arXiv
-
[104]
Honovich, T
O. Honovich, T. Scialom, O. Levy, and T. Schick. Unnatural instructions: Tuning language models with (almost) no human labor. In A. Rogers, J. Boyd-Graber, and N. Okazaki, editors,Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1...
2023 doi
-
[105]
A. Hu, H. Xu, J. Ye, M. Yan, L. Zhang, B. Zhang, J. Zhang, Q. Jin, F. Huang, and J. Zhou. mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.Findings of the Association for Computational Linguistics: EMNLP 2024, pages 3096–3120, 2024. Cited on pag...
2024
-
[106]
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022. Cited on page 49
2022
-
[107]
Huang, Y
Y . Huang, Y . Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y . Zhang, Y . Fu, et al. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.Advances in neural information processing systems, 36:62991–63010, 2023. Cited on page 66
2023
-
[108]
Huang, K
Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. Jawahar. ICDAR2019 competition on scanned receipt OCR and information extraction. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53 and 54
2019
-
[109]
D. A. Hudson and C. D. Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6700–6709, 2019. Cited on pages 53, 54, and 66
2019
-
[110]
Ionescu, H
B. Ionescu, H. Müller, et al. Overview of the ImageCLEF 2024: Multimedia retrieval in medical applications. InInternational Conference of the Cross-Language Evaluation Forum for European Languages, 2024. Cited on pages 53 and 54
2024
-
[111]
Jhamtani and T
H. Jhamtani and T. Berg-Kirkpatrick. Learning to describe differences between pairs of similar images. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2018. Cited on pages 53 and 54
2018
-
[112]
C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning (ICML), pages 4904–4916. PMLR, 2021....
2021
-
[113]
Y . Jia, J. Li, X. Yue, B. Li, P. Nie, K. Zou, and W. Chen. VisualWebInstruct: Scaling up multimodal instruction data through web search. In C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, editors,Conference on Empirical Methods in Natural Language Processing (EM...
2025 doi
-
[114]
Jiang, X
D. Jiang, X. He, H. Zeng, C. Wei, M. Ku, Q. Liu, and W. Chen. Mantis: Interleaved multi- image instruction tuning.arXiv preprint arXiv:2405.01483, 2024. Cited on page 66
2024
-
[115]
Jiang, K
M. Jiang, K. Z. Liu, M. Zhong, R. Schaeffer, S. Ouyang, J. Han, and S. Koyejo. Investigating data contamination for pre-training language models.arXiv preprint arXiv:2401.06059, 2024. Cited on page 46. 20
2024
-
[116]
Joshi, E
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, 2017...
2017
-
[117]
Joshi, H
S. Joshi, H. Yin, R. Adiga, H. Mongstad, A. Deng, A. Carranza, A. Fang, A. Abbas, A. Suri, B. Larsen, et al. 20/20 vision language models: A prescription for better vlms through data curation alone.arXiv preprint arXiv:2605.11405, 2026. Cited on page 2
2026 arXiv
-
[118]
Joshi, H
S. Joshi, H. Yin, R. Adiga, R. Monti, A. Carranza, A. Fang, A. Deng, A. Abbas, B. Larsen, C. Blakeney, et al. Datbench: Discriminative, faithful, and efficient vlm evaluations.arXiv preprint arXiv:2601.02316, 2026. Cited on page 2
2026
-
[119]
Kafle, B
K. Kafle, B. Price, S. Cohen, and C. Kanan. DVQA: Understanding data visualizations via question answering. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5648–5656, 2018. Cited on pages 53 and 54
2018
-
[120]
S. E. Kahou, V . Michalski, A. Atkinson, Á. Kádár, A. Trischler, and Y . Bengio. FigureQA: An annotated figure dataset for visual reasoning.arXiv preprint arXiv:1710.07300, 2017. Cited on pages 53 and 54
2017 arXiv
-
[121]
F. Kang, Y . Sun, B. Wen, S. Chen, D. Song, R. Mahmood, and R. Jia. Autoscale: Scale-aware data mixing for pre-training llms.arXiv preprint arXiv:2407.20177, 2024. Cited on pages 3 and 45
2024
-
[122]
Kantharaj, R
S. Kantharaj, R. T. Leong, X. Lin, A. Masry, M. Thakkar, E. Hoque, and S. Joty. Chart-to-text: A large-scale benchmark for chart summarization. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4005–4023, 2...
2022
-
[123]
Karamcheti, S
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models. InForty- first International Conference on Machine Learning, 2024. Cited on page 7
2024
-
[124]
Karpathy and L
A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3128– 3137, 2015. Cited on page 66
2015
-
[125]
Kazemi, H
M. Kazemi, H. Alvari, A. Anand, J. Wu, X. Chen, and R. Soricut. GeomVerse: A systematic evaluation of large models for geometric reasoning.arXiv preprint arXiv:2312.12241, 2023. Cited on pages 53 and 54
2023
-
[126]
Kazemzadeh, V
S. Kazemzadeh, V . Ordonez, M. Matten, and T. Berg. ReferItGame: Referring to objects in photographs of natural scenes. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2014. Cited on pages 53, 54, and 66
2014
-
[127]
Kembhavi, M
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi. A diagram is worth a dozen images. InEuropean Conference on Computer Vision (ECCV), pages 235–251. Springer, 2016. Cited on pages 53, 54, and 66
2016
-
[128]
Kembhavi, M
A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi. Are you smarter than a sixth grader? Textbook question answering for multimodal machine comprehension. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. Cited on pages 53 and 54. 21
2017
-
[129]
Kiela, H
D. Kiela, H. Firooz, A. Mohan, V . Goswami, A. Singh, P. Ringshia, and D. Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes.Advances in Neural Information Processing Systems (NeurIPS), 2020. Cited on pages 53 and 54
2020
-
[130]
G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park. OCR-free document understanding transformer. InEuropean Conference on Computer Vision (ECCV), 2022. Cited on pages 53 and 54
2022
-
[131]
W. Kim, S. Chun, T. Kim, D. Han, and S. Yun. Hype: Hyperbolic entailment filtering for underspecified images and texts. InEuropean Conference on Computer Vision (ECCV), pages 247–265. Springer, 2024. Cited on page 45
2024
-
[132]
Know-Saraswati-CoT: Chain-of-thought Sanskrit/Hindi reasoning dataset
knowrohit07 and Knowledge Tech Team. Know-Saraswati-CoT: Chain-of-thought Sanskrit/Hindi reasoning dataset. https://huggingface.co/datasets/knowrohit07/ know-saraswati-cot, 2024. Hugging Face dataset card. Cited on pages 53 and 54
2024
-
[133]
Kuang, W
J. Kuang, W. Hua, D. Liang, M. Yang, D. Jiang, B. Ren, and X. Bai. Visual information extraction in the wild: practical dataset and end-to-end solution.International Conference on Document Analysis and Recognition (ICDAR), 2023. Cited on pages 53 and 54
2023
-
[134]
Kumar, A
A. Kumar, A. Raghunathan, R. Jones, T. Ma, and P. Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution.arXiv preprint arXiv:2202.10054, 2022. Cited on page 8
2022
-
[135]
Kuznetsova, H
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al. The Open Images dataset V4: Unified image classification, object detection, and visual relationship detection at scale.International Journal of Comp...
1956
-
[136]
Kwiatkowski, J
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al. Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019. Cite...
2019
-
[137]
G. Lai, Q. Xie, H. Liu, Y . Yang, and E. Hovy. Race: Large-scale reading comprehension dataset from examinations. InProceedings of the 2017 conference on empirical methods in natural language processing, pages 785–794, 2017. Cited on page 66
2017
-
[138]
Lambert, J
N. Lambert, J. Morrison, V . Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V . Miranda, A. Liu, N. Dziri, S. Lyu, et al. Tulu 3: Pushing frontiers in open language model post-training.arXiv preprint arXiv:2411.15124, 2024. Cited on page 68
2024 arXiv
-
[139]
J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific Data, 5(1):1–10, 2018. Cited on pages 53 and 54
2018
-
[140]
Laurençon, L
H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. Rush, D. Kiela, et al. Obelics: An open web-scale filtered dataset of interleaved image-text documents.Advances in Neural Information Processing Systems (NeurIPS), 36:71683–7170...
2023
-
[141]
Laurençon, L
H. Laurençon, L. Tronchon, M. Cord, and V . Sanh. What matters when building vision- language models?Advances in Neural Information Processing Systems (NeurIPS), 37: 87874–87907, 2024. Cited on page 45. 22
2024
-
[142]
Laurençon, L
H. Laurençon, L. Tronchon, M. Cord, and V . Sanh. What matters when building vision- language models?Advances in Neural Information Processing Systems (NeurIPS), 37: 87874–87907, 2024. Cited on page 3
2024
-
[143]
Laurençon, A
H. Laurençon, A. Marafioti, V . Sanh, and L. Tronchon. Building and better understanding vision-language models: insights and future directions, 2024. URL https://arxiv.org/ abs/2408.12637. Cited on pages 53 and 54
2024
-
[144]
Laurençon, L
H. Laurençon, L. Tronchon, and V . Sanh. Unlocking the conversion of web screenshots into html code with the websight dataset, 2024. URLhttps://arxiv.org/abs/2403.09029. Cited on pages 53 and 54
2024
-
[145]
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini. Deduplicating training data makes language models better. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, 2...
2022
-
[146]
Lee and S
S. Lee and S. Hwang. Selective training for large vision language models via visual information gain.arXiv preprint arXiv:2602.17186, 2026. Cited on page 6
2026 arXiv
-
[147]
Lerner, O
P. Lerner, O. Ferret, C. Guinaudeau, H. Le Borgne, R. Besançon, J. G. Moreno, and J. Lovon- Melgarejo. ViQuAE, a dataset for knowledge-based visual question answering about named entities. InACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 202...
2022
-
[148]
B. Li, R. Wang, G. Wang, Y . Ge, Y . Ge, and Y . Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension.arXiv preprint arXiv:2307.16125, 2023. Cited on page 66
2023 arXiv
-
[149]
B. Li, Y . Ge, Y . Chen, Y . Ge, R. Zhang, and Y . Shan. Seed-bench-2-plus: Benchmarking multimodal large language models with text-rich visual comprehension.arXiv preprint arXiv:2404.16790, 2024. Cited on page 66
2024
-
[150]
B. Li, Z. Lin, W. Peng, J. d. D. Nyandwi, D. Jiang, Z. Ma, S. Khanuja, R. Krishna, G. Neubig, and D. Ramanan. Naturalbench: Evaluating vision-language models on natural adversarial samples.Advances in Neural Information Processing Systems, 37:17044–17068, 2024. Cited on page 66
2024
-
[151]
B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liu, et al. Llava-onevision: Easy visual task transfer.arXiv preprint arXiv:2408.03326, 2024. Cited on pages 3 and 87
2024 arXiv
-
[152]
C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao. LLaV A-Med: Training a large language-and-vision assistant for biomedicine in one day. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023. Ci...
2023
-
[153]
H. Li, Y . Zhang, F. Koto, Y . Yang, H. Zhao, Y . Gong, N. Duan, and T. Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese. InFindings of the Association for Computational Linguistics: ACL 2024, pages 11260–11285, 2024. Cited on page 66
2024
-
[154]
H. Li, Y . Chen, S. Miao, Q. Dong, J. Chen, Y . Hu, J. Chen, M. Qin, Y . Wu, Y . Zhou, et al. Legalone: a family of foundation models for reliable legal reasoning.arXiv preprint arXiv:2602.00642, 2026. Cited on page 8. 23
2026
-
[155]
J. Li, D. Li, S. Savarese, and S. Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InInternational Conference on Machine Learning (ICML), pages 19730–19742. PMLR, 2023. Cited on page 45
2023
-
[156]
J. LI, E. Beeching, L. Tunstall, B. Lipkin, R. Soletskyi, S. Huang, K. Rasul, L. Yu, A. Q. Jiang, Z. Shen, Z. Qin, B. Dong, L. Zhou, Y . Fleureau, G. Lample, and S. Polu. NuminaMath 1.5: Second iteration of NuminaMath, 2024. Hugging Face dataset card https://huggingface. co/da...
2024
-
[157]
J. LI, E. Beeching, L. Tunstall, et al. NuminaMath-TIR: Tool-integrated reasoning math dataset, 2024. Hugging Face dataset card https://huggingface.co/datasets/AI-MO/ NuminaMath-TIR. Cited on pages 53 and 54
2024
-
[158]
J. Li, A. Fang, G. Smyrnis, M. Ivgi, M. Jordan, S. Gadre, H. Bansal, E. Guha, S. Keh, K. Arora, et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems (NeurIPS), 37:14200–14282, 2024. Cited o...
2024
-
[159]
J. Li, J. Chen, Y . Qu, S. Xu, Z. Lin, J. Zhu, B. Xu, W. Tan, P. Fu, J. Ju, et al. Xiaomi mimo-vl-miloco technical report.arXiv preprint arXiv:2512.17436, 2025. Cited on page 2
2025
-
[160]
L. Li, Y . Wang, R. Xu, P. Wang, X. Feng, L. Kong, and Q. Liu. Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models. InAnnual Meeting of the Association for Computational Linguistics (ACL), 2024. Cited on pages 53 and 54
2024
-
[161]
Q. Li, Z. Chen, W. Wang, W. Wang, S. Ye, Z. Jin, G. Chen, Y . He, Z. Gao, E. Cui, et al. Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text. arXiv preprint arXiv:2406.08418, 2024. Cited on pages 4, 53, 54, 55, and 76
2024
-
[162]
Li and N
S. Li and N. Tajbakhsh. SciGraphQA: A large-scale synthetic multi-turn question-answering dataset for scientific graphs.arXiv preprint arXiv:2308.03349, 2023. Cited on pages 53 and 54
2023
-
[163]
X. Li, H. Tu, M. Hui, Z. Wang, B. Zhao, J. Xiao, S. Ren, J. Mei, Q. Liu, H. Zheng, et al. What if we recaption billions of web images with llama-3?arXiv preprint arXiv:2406.08478, 2024. Cited on pages 45 and 78
2024
-
[164]
Y . Li, Y . Du, K. Zhou, J. Wang, X. Zhao, and J.-R. Wen. Evaluating object hallucination in large vision-language models. InProceedings of the 2023 conference on empirical methods in natural language processing, pages 292–305, 2023. Cited on pages 65 and 66
2023
-
[165]
W. Lian, G. Wang, B. Goodson, E. Pentland, A. Cook, C. V ong, and "Teknium". Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023. URL https://https://huggingface.co/Open-Orca/SlimOrca. Cited on pages 4, 53, and 54
2023
-
[166]
H. Lin, V . Hosu, and D. Saupe. Kadid-10k: A large-scale artificially distorted iqa database. In2019 Eleventh International Conference on Quality of Multimedia Experience (QoMEX), pages 1–3. IEEE, 2019. Cited on page 75
2019
-
[167]
T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. InEuropean Conference on Computer Vision (ECCV), pages 740–755. Springer, 2014. Cited on page 66. 24
2014
-
[168]
A. D. Lindström and S. S. Abraham. CLEVR-Math: A dataset for compositional language, visual and mathematical reasoning. InInternational Workshop on Neural-Symbolic Learning and Reasoning (NeSy), 2022. Cited on pages 53 and 54
2022
-
[169]
Liu, L.-M
B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y . Yang, and X.-M. Wu. SLAKE: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. InIEEE International Symposium on Biomedical Imaging (ISBI), 2021. Cited on pages 53 and 54
2021
-
[170]
C.-L. Liu, F. Yin, D.-H. Wang, and Q.-F. Wang. CASIA online and offline Chinese handwriting databases.International Conference on Document Analysis and Recognition (ICDAR), 2011. Cited on pages 53 and 54
2011
-
[171]
F. Liu, G. Emerson, and N. Collier. Visual spatial reasoning.Transactions of the Association for Computational Linguistics (TACL), 11:635–651, 2023. Cited on pages 53, 54, and 66
2023
-
[172]
F. Liu, K. Lin, L. Li, J. Wang, Y . Yacoob, and L. Wang. Mitigating hallucination in large multi-modal models via robust instruction tuning.International Conference on Learning Representations (ICLR), 2024. Cited on pages 53 and 54
2024
-
[173]
F. Liu, X. Wang, W. Yao, J. Chen, K. Song, S. Cho, Y . Yacoob, and D. Yu. Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2024
-
[174]
F. Liu, W. Zhou, B. Liu, P. Guo, Z. Wang, B. Zhang, Y . Zhang, Y . Yu, X. Zhou, and T. Wang. Infolaw: Information scaling laws for large language models with quality-weighted mixture data and repetition.arXiv preprint arXiv:2605.02364, 2026. Cited on page 8
2026 arXiv
-
[175]
H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning.Advances in Neural Information Processing Systems (NeurIPS), 36:34892–34916, 2023. Cited on pages 1, 2, 3, 7, 8, 45, 47, and 86
2023
-
[176]
H. Liu, C. Li, Y . Li, and Y . J. Lee. Improved baselines with visual instruction tuning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26296– 26306, 2024. Cited on pages 53 and 54
2024
-
[177]
H. Liu, C. Li, Y . Li, B. Li, Y . Zhang, S. Shen, and Y . J. Lee. Llavanext: Improved reasoning, ocr, and world knowledge, 2024. Cited on pages 2, 4, and 45
2024
-
[178]
J. Liu, T. Ou, Y . Song, Y . Qu, W. Lam, C. Xiong, W. Chen, G. Neubig, and X. Yue. Harnessing webpage uis for text-rich visual understanding, 2024. URLhttps://arxiv.org/abs/2410. 13824. Cited on pages 53 and 54
2024
-
[179]
Q. Liu, X. Zheng, N. Muennighoff, G. Zeng, L. Dou, T. Pang, J. Jiang, and M. Lin. Regmix: Data mixture as regression for language model pre-training, 2024. Cited on pages 3, 45, and 93
2024
-
[180]
X. Liu, Y . Zhu, J. Gu, Y . Lan, C. Yang, and Y . Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. InEuropean Conference on Computer Vision, pages 386–403. Springer, 2024. Cited on page 66. 25
2024
-
[181]
Y . Liu, Y . Cao, Z. Gao, W. Wang, Z. Chen, W. Wang, H. Tian, L. Lu, X. Zhu, T. Lu, et al. Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity. Science China Information Sciences, 67(12):220103, 2024. Cited on pages 53 and 54
2024
-
[182]
Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu, et al. Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision (ECCV), pages 216–233. Springer, 2024. Cited on page 66
2024
-
[183]
Y . Liu, Z. Li, M. Huang, B. Yang, W. Yu, C. Li, X.-C. Yin, C.-L. Liu, L. Jin, and X. Bai. Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 2024. Cited on page 66
2024
-
[184]
Z. Liu, T. Chu, Y . Zang, X. Dong, P. Zhang, Z. Yang, Y . Duan, D. Lin, Y . Wang, and J. Wang. MMDU: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for LVLMs. InAdvances in Neural Information Processing Systems (NeurIPS), 2024. Cited on ...
2024
-
[185]
Chinese OCR dataset
longmaodata. Chinese OCR dataset. https://huggingface.co/datasets/ longmaodata/Chinese-OCR, 2024. Hugging Face dataset card. Cited on pages 53 and 54
2024
-
[186]
Longpre, L
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y . Tay, D. Zhou, Q. V . Le, B. Zoph, J. Wei, and A. Roberts. The flan collection: Designing data and methods for effective instruction tuning. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, edi...
2023
-
[187]
Loshchilov and F
I. Loshchilov and F. Hutter. Sgdr: Stochastic gradient descent with warm restarts.arXiv preprint arXiv:1608.03983, 2016. Cited on page 50
2016 arXiv
-
[188]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017. Cited on pages 4 and 50
2017 arXiv
-
[189]
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding.arXiv preprint arXiv:2403.05525, 2024. Cited on page 2
2024 arXiv
-
[190]
P. Lu, R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu. Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021. Cited on p...
2021
-
[191]
P. Lu, L. Qiu, J. Chen, T. Xia, Y . Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu. IconQA: A new benchmark for abstract diagram understanding and visual language reasoning. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2021. Cit...
2021
-
[192]
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems (NeurIPS), 35:2507–2521, 2022. Cited on pa...
2022
-
[193]
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023. Cited on page 66
2023 arXiv
-
[194]
P. Lu, L. Qiu, K.-W. Chang, Y . N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. In International Conference on Learning Representations (ICLR), 2023. Cited on pages 53 and 54
2023
-
[195]
Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang. WizardCoder: Empowering code large language models with Evol-Instruct.International Conference on Learning Representations (ICLR), 2024. Cited on pages 53 and 54
2024
-
[196]
W. Ma, H. Chen, G. Zhang, Y .-C. Chou, J. Chen, C. de Melo, and A. Yuille. 3dsrbench: A comprehensive 3d spatial reasoning benchmark. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 6924–6934, 2025. Cited on page 66
2025
-
[197]
Madaan, A
L. Madaan, A. K. Singh, R. Schaeffer, A. Poulton, S. Koyejo, P. Stenetorp, S. Narang, and D. Hupkes. Quantifying variance in evaluation benchmarks.arXiv preprint arXiv:2406.10229, 2024. Cited on page 5
2024
-
[198]
Magar and R
I. Magar and R. Schwartz. Data contamination: From memorization to exploitation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 157–165, 2022. Cited on page 46
2022
-
[199]
Mahmoud, M
A. Mahmoud, M. Elhoushi, A. Abbas, Y . Yang, N. Ardalani, H. Leather, and A. S. Morcos. Sieve: Multimodal dataset pruning using image captioning models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22423–22432, 2024. Cited on page 45
2024
-
[200]
Maini, S
P. Maini, S. Goyal, Z. C. Lipton, J. Z. Kolter, and A. Raghunathan. T-mars: Improving visual representations by circumventing text feature learning.arXiv preprint arXiv:2307.03132, 2023. Cited on page 45
2023
-
[201]
Mao, C.-W
C. Mao, C.-W. Xie, C. Zhong, H. Deng, J. Zhao, J. Xiao, J. Xing, J. Zhang, J. Zhou, J. Zhang, et al. Wan-image: Pushing the boundaries of generative visual intelligence.arXiv preprint arXiv:2604.19858, 2026. Cited on page 3
2026 arXiv
-
[202]
H. Mao, M. Cheung, and J. She. Deepart: Learning joint representations of visual arts. In ACM International Conference on Multimedia, pages 1183–1191, 2017. Cited on pages 53 and 54
2017
-
[203]
Marafioti, O
A. Marafioti, O. Zohar, M. Farré, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Tazi, et al. Smolvlm: Redefining small and efficient multimodal models.arXiv preprint arXiv:2504.05299, 2025. Cited on page 2
2025 arXiv
-
[204]
Marino, M
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. Cited on pages 53, 54, and 66
2019
-
[205]
Marti and H
U.-V . Marti and H. Bunke. The IAM-database: an English sentence database for offline handwriting recognition.International Journal on Document Analysis and Recognition (IJDAR), 5:39–46, 2002. Cited on pages 53 and 54. 27
2002
-
[206]
Masry, X
A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguistics: ACL 2022, pages 2263–2279, 2022. Cited on pages 53, 54, and 66
2022
-
[207]
Masry, P
A. Masry, P. Kavehzadeh, D. X. Long, E. Hoque, and S. Joty. UniChart: A universal vision- language pretrained model for chart comprehension and reasoning. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2023. Cited on pages 53 and 54
2023
-
[208]
Masry, M
A. Masry, M. Thakkar, A. Bajaj, A. Kartha, E. Hoque, and S. Joty. Chartgemma: Visual instruction-tuning for chart reasoning in the wild.Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, pages 625–643, 2025. Cited on pages 53 and 54
2025
-
[209]
Mathew, D
M. Mathew, D. Karatzas, and C. Jawahar. Docvqa: A dataset for vqa on document images. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021. Cited on pages 53, 54, and 66
2021
-
[210]
Mathew, V
M. Mathew, V . Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar. Infographicvqa. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 1697–1706, 2022. Cited on pages 53, 54, and 66
2022
-
[211]
KOpen-Hermes-25: Korean translation of OpenHermes-2.5, 2024
Maywell. KOpen-Hermes-25: Korean translation of OpenHermes-2.5, 2024. Hugging Face dataset cardhttps://huggingface.co/datasets/maywell/ko_Ultrafeedback_ binarized. Cited on pages 53 and 54
2024
-
[212]
Mazumder, C
M. Mazumder, C. Banbury, X. Yao, B. Karlaš, W. Gaviria Rojas, S. Diamos, G. Diamos, L. He, A. Parrish, H. R. Kirk, et al. Dataperf: Benchmarks for data-centric ai development.Advances in Neural Information Processing Systems (NeurIPS), 36:5320–5347, 2023. Cited on pages 3 and 45
2023
-
[213]
McKinzie, Z
B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, A. Belyi, et al. Mm1: methods, analysis and insights from multimodal llm pre-training. In European Conference on Computer Vision (ECCV), pages 304–323. Springer, 2024. Cited on pages...
2024
-
[214]
Methani, P
N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar. Plotqa: Reasoning over scientific plots. InIEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 1527–1536, 2020. Cited on pages 53 and 54
2020
-
[215]
Mishra, S
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty. OCR-VQA: Visual question answering by reading text in images. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53, 54, and 66
2019
-
[216]
Mitra, H
A. Mitra, H. Khanpour, C. Rosset, and A. Awadallah. Orca-math: Unlocking the potential of slms in grade school math, 2024. URLhttps://arxiv.org/abs/2402.14830. Cited on pages 53 and 54
2024
-
[217]
Mizrahi, A
D. Mizrahi, A. B. L. Larsen, J. Allardice, S. Petryk, Y . Gorokhov, J. Li, A. Fang, J. Gardner, T. Gunter, and A. Dehghan. Language models improve when pretraining data matches target tasks.arXiv preprint arXiv:2507.12466, 2025. Cited on pages 4, 6, 7, 46, and 82
2025
-
[218]
Mohri, J
C. Mohri, J. Duchi, and T. Hashimoto. A bitter lesson for data filtering.arXiv preprint arXiv:2605.19407, 2026. Cited on page 7. 28
2026 arXiv
-
[219]
Muennighoff, A
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel. Scaling data-constrained language models.Advances in Neural Information Processing Systems (NeurIPS), 36:50358–50376, 2023. Cited on page 8
2023
-
[220]
V . K. Nagaraja, V . I. Morariu, and L. S. Davis. Modeling context between objects for referring expression understanding. InEuropean Conference on Computer Vision (ECCV), pages 792–
-
[221]
Cited on pages 66 and 67
Springer, 2016. Cited on pages 66 and 67
2016
-
[222]
Nezhurina, T
M. Nezhurina, T. Porian, G. Pucceti, T. Kerssies, R. Beaumont, M. Cherti, and J. Jitsev. Scaling laws for robust comparison of open foundation language-vision models and datasets.arXiv preprint arXiv:2506.04598, 2025. Cited on pages 4 and 7
2025
-
[223]
H. Ngo, M. Deitke, M. Bartelds, S. Pratt, J. Gardner, M. Jordan, and L. Schmidt. Olmoasr: Open models and data for training robust speech recognition models.arXiv preprint arXiv:2508.20869, 2025. Cited on page 1
2025
-
[224]
Nguyen, V
H. Nguyen, V . May, H. Raj, M. Nezhurina, Y . Wang, Y . Luo, M. C. Vu, T. Nakamura, K. Tsui, V . K. Nguyen, et al. Mixturevitae: Open web-scale pretraining dataset with high quality instruction and reasoning data built from permissive-first text sources.arXiv preprint arXiv:25...
2025
-
[225]
Nguyen, G
T. Nguyen, G. Ilharco, M. Wortsman, S. Oh, and L. Schmidt. Quality not quantity: On the interaction between dataset design and robustness of clip.Advances in Neural Information Processing Systems (NeurIPS), 35:21455–21469, 2022. Cited on page 1
2022
-
[226]
Nguyen, S
T. Nguyen, S. Y . Gadre, G. Ilharco, S. Oh, and L. Schmidt. Improving multimodal datasets with image captioning.Advances in Neural Information Processing Systems (NeurIPS), 36: 22047–22069, 2023. Cited on page 45
2023
-
[227]
Nguyen, M
T. Nguyen, M. Wallingford, S. Santy, W.-C. Ma, S. Oh, L. Schmidt, P. W. Koh, and R. Krishna. Multilingual diversity improves vision-language representations.Advances in Neural Information Processing Systems (NeurIPS), 37:91430–91459, 2024. Cited on pages 45 and 59
2024
-
[228]
T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y . Gu, S. Huang, M. Jordan, et al. 2 olmo 2 furious.arXiv preprint arXiv:2501.00656, 2024. Cited on pages 3, 5, 45, and 77
2024 arXiv
-
[229]
T. Olmo, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, et al. Olmo 3.arXiv preprint arXiv:2512.13961, 2025. Cited on page 5
2025 arXiv
-
[230]
Chat markup language (ChatML)
OpenAI. Chat markup language (ChatML). https://github.com/openai/ openai-python/blob/main/chatml.md, 2022. Accessed: 29 April 2026. Cited on page 75
2022
-
[231]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023. Cited on pages 45 and 46
2023 arXiv
-
[232]
Parashar, Z
S. Parashar, Z. Lin, T. Liu, X. Dong, Y . Li, D. Ramanan, J. Caverlee, and S. Kong. The neglected tails in vision-language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12988–12997, 2024. Cited on page 60. 29
2024
-
[233]
Penedo, Q
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only.arXiv preprint arXiv:2306.01116, 2023. Cited on page 1
2023 arXiv
-
[234]
Penedo, H
G. Penedo, H. Kydlí ˇcek, A. Lozhkov, M. Mitchell, C. Raffel, L. V on Werra, T. Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems (NeurIPS), 37:30811–30849, 2024. Cited on pages 1, 3, 5, 45,...
2024
-
[235]
Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei. Kosmos-2: Grounding multimodal large language models to the world.arXiv preprint arXiv:2306.14824, 2023. Cited on pages 53, 54, and 55
2023 arXiv
-
[236]
Pizzi, S
E. Pizzi, S. D. Roy, S. N. Ravindra, P. Goyal, and M. Douze. A self-supervised descriptor for image copy detection. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14532–14542, 2022. Cited on pages 4 and 68
2022
-
[237]
Ponomarenko, O
N. Ponomarenko, O. Ieremeiev, V . Lukin, K. Egiazarian, L. Jin, J. Astola, B. V ozel, K. Chehdi, M. Carli, F. Battisti, et al. Color image database tid2013: Peculiarities and preliminary results. InEuropean workshop on visual information processing (EUVIP), pages 106–111. IEEE...
2013
-
[238]
Pouget, L
A. Pouget, L. Beyer, E. Bugliarello, X. Wang, A. P. Steiner, X. Zhai, and I. Alabdulmohsin. No filter: Cultural and socioeconomic diversity in contrastive vision-language models.Advances in Neural Information Processing Systems (NeurIPS), 37:106474–106496, 2024. Cited on page 59
2024
-
[239]
Pramanick, R
S. Pramanick, R. Chellappa, and S. Venugopalan. SPIQA: A dataset for multimodal question answering on scientific papers.Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2024. Cited on pages 53 and 54
2024
-
[240]
R. Qiao, Q. Tan, G. Dong, M. MinhuiWu, C. Sun, X. Song, J. Wang, Z. Gongque, S. Lei, Y . Zhang, et al. We-math: Does your large multimodal model achieve human-like mathematical reasoning? InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics...
2025
-
[241]
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...
2025 arXiv
-
[242]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning (ICML), pages 8748–8763. PmLR, 2021. Cited...
2021
-
[243]
Rajani, L
N. Rajani, L. Tunstall, E. Beeching, N. Lambert, A. M. Rush, and T. Wolf. No Robots, 2023. https://huggingface.co/datasets/HuggingFaceH4/no_robots. Cited on pages 53 and 54
2023
-
[244]
Rajbhandari, J
S. Rajbhandari, J. Rasley, O. Ruwase, and Y . He. Zero: Memory optimizations toward training trillion parameter models. InSC20: international conference for high performance computing, networking, storage and analysis, pages 1–16. IEEE, 2020. Cited on page 50. 30
2020
-
[245]
leetcode
RayBernard. leetcode. https://huggingface.co/datasets/RayBernard/leetcode,
-
[246]
Cited on pages 53 and 54
Hugging Face dataset card. Cited on pages 53 and 54
-
[247]
Rodriguez, X
J. Rodriguez, X. Jian, S. S. Panigrahi, T. Zhang, A. Feizi, A. Puri, A. Kalkunte Suresh, F. Savard, A. Masry, S. Nayak, R. Awal, M. Massoud, A. Abaskohi, Z. Li, S. Wang, P.-A. Noël, M. L. Richter, S. Vadacchino, S. Agarwal, S. Biswas, S. Shanian, Y . Zhang, N. Bolger, K. MacDo...
2025
-
[248]
K. Roth, V . Udandarao, S. Dziadzio, A. Prabhu, M. Cherti, O. Vinyals, O. Hénaff, S. Albanie, M. Bethge, and Z. Akata. A practitioner’s guide to continual multimodal pretraining.arXiv preprint arXiv:2408.14471, 2024. Cited on page 50
2024
-
[249]
Sainz, I
O. Sainz, I. García-Ferrero, A. Jacovi, J. A. Campos, Y . Elazar, E. Agirre, Y . Goldberg, W.-L. Chen, J. Chim, L. Choshen, et al. Data contamination report from the 2024 conda shared task. InProceedings of the 1st Workshop on Data Contamination (CONDA), pages 41–56, 2024. Cit...
2024
-
[250]
Sakaguchi, R
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y . Choi. Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021. Cited on page 66
2021
-
[251]
Saxena, P
R. Saxena, P. Minervini, and F. Keller. PosterSum: A multimodal benchmark for scientific poster summarization. In K. Inui, S. Sakti, H. Wang, D. F. Wong, P. Bhattacharyya, B. Banerjee, A. Ekbal, T. Chakraborty, and D. P. Singh, editors,Proceedings of the 14th International Joi...
2025
-
[252]
Schaeffer, J
R. Schaeffer, J. Kazdan, B. Abbasi, K. Z. Liu, B. Miranda, A. Ahmed, F. Berez, A. Puri, S. Biderman, N. Mireshghallah, et al. Quantifying the effect of test set contamination on generative evaluations.arXiv preprint arXiv:2601.04301, 2026. Cited on page 46
2026
-
[253]
Schuhmann, R
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems (NeurIPS), 35:252...
2022
-
[254]
Schwenk, A
D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowledge. InEuropean Conference on Computer Vision (ECCV), 2022. Cited on pages 53 and 54
2022
-
[255]
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar. KVQA: Knowledge-aware visual question answering. InAAAI Conference on Artificial Intelligence (AAAI), 2019. Cited on pages 53 and 54
2019
-
[256]
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun. Objects365: A large-scale, high-quality dataset for object detection. InIEEE/CVF International Conference on Computer Vision (ICCV), 2019. Cited on pages 53 and 54
2019
-
[257]
Shapourian, K
H. Shapourian, K. Hejazi, O. M. Sule, and B. Millidge. Zaya1-vl-8b technical report.arXiv preprint arXiv:2605.08560, 2026. Cited on page 3. 31
2026 arXiv
-
[258]
N. Shazeer. Glu variants improve transformer.arXiv preprint arXiv:2002.05202, 2020. Cited on page 48
2002 arXiv
-
[259]
W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pa...
2016
-
[260]
Shinoda, K
R. Shinoda, K. Saito, S. Tanaka, T. Hirasawa, and Y . Ushiku. SBS Figures: Pre-training figure QA from stage-by-stage synthesized images.arXiv preprint arXiv:2412.17606, 2024. Cited on pages 53 and 54
2024
-
[261]
Shukor, L
M. Shukor, L. Bethune, D. Busbridge, D. Grangier, E. Fini, A. El-Nouby, and P. Ablin. Scaling laws for optimal data mixtures.arXiv preprint arXiv:2507.09404, 2025. Cited on page 7
2025
-
[262]
Sidorov, R
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh. TextCaps: A dataset for image captioning with reading comprehension. InEuropean Conference on Computer Vision (ECCV), pages 742–758. Springer, 2020. Cited on pages 53 and 54
2020
-
[263]
Singh, V
A. Singh, V . Natarajan, M. Shah, Y . Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach. Towards vqa models that can read. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8317–8326, 2019. Cited on pages 53, 54, and 66
2019
-
[264]
Singh, G
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner. TextOCR: Towards large- scale end-to-end reasoning for arbitrary-shaped scene text. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Cited on pages 53 and 54
2021
-
[265]
e. a. Singh. Persian synthetic OCR dataset (ParSynth-OCR-200K), 2021. Hugging Face dataset card. Cited on pages 53 and 54
2021
-
[266]
Sorscher, R
B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. Morcos. Beyond neural scaling laws: beating power law scaling via data pruning.Advances in Neural Information Processing Systems (NeurIPS), 35:19523–19536, 2022. Cited on page 45
2022
-
[267]
Steiner, A
A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y . Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, et al. Paligemma 2: A family of versatile vlms for transfer.arXiv preprint arXiv:2412.03555, 2024. Cited on page 2
2024 arXiv
-
[268]
Stiennon, L
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. V oss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. InAdvances in Neural Information Processing Systems (NeurIPS), 2020. Cited on pages 53 and 54
2020
-
[269]
D. Su, K. Kong, Y . Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguisti...
2025
-
[270]
J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. Cited on page 48
2024
-
[271]
Sun, D.-W
H.-L. Sun, D.-W. Zhou, Y . Li, S. Lu, C. Yi, Q.-G. Chen, Z. Xu, W. Luo, K. Zhang, D.-C. Zhan, et al. Parrot: Multilingual visual instruction tuning.arXiv preprint arXiv:2406.02539, 2024. Cited on page 66. 32
2024
-
[272]
T. Sun, X. Zhang, Z. He, P. Li, Q. Cheng, X. Liu, H. Yan, Y . Shao, Q. Tang, S. Zhang, et al. MOSS: An open conversational large language model.Machine Intelligence Research, 2024. Cited on pages 53 and 54
2024
-
[273]
Y . Sun, Z. Ni, C.-K. Chng, Y . Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas, et al. ICDAR2019 competition on large-scale street view text with partial labeling - RRC-LSVT. In International Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on ...
2019
-
[274]
Tanaka, K
R. Tanaka, K. Nishida, and S. Yoshida. VisualMRC: Machine reading comprehension on document images. InAAAI Conference on Artificial Intelligence (AAAI), 2021. Cited on pages 53 and 54
2021
-
[275]
B. J. Tang, A. Boggust, and A. Satyanarayan. VisText: A benchmark for semantically rich chart captioning.Annual Meeting of the Association for Computational Linguistics (ACL), 2023. Cited on pages 53 and 54
2023
-
[276]
J. Tang, Q. Liu, Y . Ye, J. Lu, S. Wei, A.-L. Wang, C. Lin, H. Feng, Z. Zhao, Y . Wang, et al. Mtvqa: Benchmarking multilingual text-centric visual question answering. InFindings of the Association for Computational Linguistics: ACL 2025, pages 7748–7763, 2025. Cited on page 66
2025
-
[277]
C. Team, Z. Yue, Z. Lin, Y . Song, W. Wang, S. Ren, S. Gu, S. Li, P. Li, L. Zhao, L. Li, K. Bao, H. Tian, H. Zhang, G. Wang, D. Zhu, Cici, C. He, B. Ye, B. Shen, Z. Zhang, Z. Jiang, Z. Zheng, Z. Song, Z. Luo, Y . Yu, Y . Wang, Y . Tian, Y . Tu, Y . Yan, Y . Huang, X. Wang, X. ...
2025
-
[278]
G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. bastien Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. T...
2025 arXiv
-
[279]
K. Team, A. Du, B. Yin, B. Xing, B. Qu, B. Wang, C. Chen, C. Zhang, C. Du, C. Wei, et al. Kimi-vl technical report.arXiv preprint arXiv:2504.07491, 2025. Cited on page 2
2025 arXiv
-
[280]
T. M. A. Team. Mai-thinking-1: Building a hill-climbing machine. Technical report, Microsoft AI, 2026. URLhttps://microsoft.ai/pdf/mai-thinking-1.pdf. Cited on pages 4 and 82
2026
-
[281]
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. Advances in Neural Information Processing Systems (NeurIPS), 37:87310–87356, 2024. Cited on pages 2, ...
2024
-
[282]
S. Tong, Z. Liu, Y . Zhai, Y . Ma, Y . LeCun, and S. Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9568–9578, 2024. Cited on page 66
2024
-
[283]
T. H. Trinh and Q. V . Le. A simple method for commonsense reasoning.arXiv preprint arXiv:1806.02847, 2018. Cited on page 46
2018
-
[284]
Tschannen, A
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y . Xia, B. Mustafa, et al. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:250...
2025 arXiv
-
[285]
Y . Tuo, W. Xiang, J.-Y . He, Y . Geng, and X. Xie. Anytext: Multilingual visual text generation and editing. InInternational Conference on Learning Representations (ICLR), 2024. URL https://openreview.net/forum?id=ezBH9WE9s2. Cited on pages 53 and 54
2024
-
[286]
zero-shot
V . Udandarao, A. Prabhu, A. Ghosh, Y . Sharma, P. H. Torr, A. Bibi, S. Albanie, and M. Bethge. No" zero-shot" without exponential data: Pretraining concept frequency determines multimodal model performance.Advances in Neural Information Processing Systems (NeurIPS), 37:61735–...
2024
-
[287]
Udandarao, Z
V . Udandarao, Z. Lu, X. Chang, Y . Wang, V . Z. Yao, A. M. Jose, F. Faghri, J. Gardner, and C.-C. Chiu. Data-centric lessons to improve speech-language pretraining.arXiv preprint arXiv:2510.20860, 2025. Cited on page 1
2025
-
[288]
Udandarao, N
V . Udandarao, N. Parthasarathy, M. F. Naeem, T. Evans, S. Albanie, F. Tombari, Y . Xian, A. Tonioni, and O. J. Hénaff. Active data curation effectively distills large-scale multimodal models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14422...
2025
-
[289]
Ustalov, N
D. Ustalov, N. Pavlichenko, S. Koshelev, D. Likhobaba, and A. Smirnova. Toloka visual question answering benchmark.arXiv preprint arXiv:2309.16511, 2023. Cited on pages 53, 54, and 66
2023
-
[290]
Van Horn, O
G. Van Horn, O. Mac Aodha, Y . Song, Y . Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie. The iNaturalist species classification and detection dataset. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. Cited on pages 53 and 54. 34
2018
-
[291]
A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie. COCO-Text: Dataset and benchmark for text detection and recognition in natural images. InarXiv preprint arXiv:1601.07140, 2016. Cited on pages 53 and 54
2016 arXiv
-
[292]
H. V . V o, V . Khalidov, T. Darcet, T. Moutakanni, N. Smetanin, M. Szafraniec, H. Touvron, C. Couprie, M. Oquab, A. Joulin, et al. Automatic data curation for self-supervised learning: A clustering-based approach.arXiv preprint arXiv:2405.15613, 2024. Cited on page 45
2024
-
[293]
B. Wang, G. Li, X. Zhou, Z. Chen, T. Grossman, and Y . Li. Screen2words: Automatic mobile ui summarization with multimodal learning. InThe 34th Annual ACM Symposium on User Interface Software and Technology, pages 498–510, 2021. Cited on pages 53 and 54
2021
-
[294]
F. Wang, X. Fu, J. Y . Huang, Z. Li, Q. Liu, X. Liu, M. D. Ma, N. Xu, W. Zhou, K. Zhang, et al. Muirbench: A comprehensive benchmark for robust multi-image understanding.arXiv preprint arXiv:2406.09411, 2024. Cited on page 66
2024 arXiv
-
[295]
J. Wang, L. Meng, Z. Weng, B. He, Z. Wu, and Y .-G. Jiang. To see is to believe: Prompting gpt- 4v for better visual instruction tuning, 2023. URL https://arxiv.org/abs/2311.07574. Cited on pages 53 and 54
2023
-
[296]
J. Wang, Y . Wang, G. Xu, J. Zhang, Y . Gu, H. Jia, J. Wang, H. Xu, M. Yan, J. Zhang, et al. Amber: An llm-free multi-dimensional benchmark for mllms hallucination evaluation.arXiv preprint arXiv:2311.07397, 2023. Cited on page 66
2023 arXiv
-
[297]
J. Wang, P. Zhang, T. Chu, Y . Cao, Y . Zhou, T. Wu, B. Wang, C. He, and D. Lin. V3Det: Vast vocabulary visual detection dataset. InIEEE/CVF International Conference on Computer Vision (ICCV), 2023. Cited on pages 53 and 54
2023
-
[298]
K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, and H. Li. Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024. Cited on page 66
2024
-
[299]
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. Cited on pages 5, 62, 78, and 79
2024 arXiv
-
[300]
W. Wang, Y . Ren, H. Luo, T. Li, C. Yan, Z. Chen, W. Wang, Q. Li, L. Lu, X. Zhu, et al. The all-seeing project V2: Towards general relation comprehension of the open world. InEuropean Conference on Computer Vision (ECCV), 2024. Cited on pages 53, 54, and 66
2024
Reviewed July 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.