Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

DataComp-VLM: Improved Open Datasets for Vision-Language Models

T0 review · 2 major / 1 minor · reviewed 2026-07-02 · grok-4.3

Pith's one-line read Instruction-heavy data mixtures scale better than caption-heavy ones for vision-language model training.

desk verdict DCVLM gives a controlled benchmark and 6T-token corpus for VLM data curation, with evidence that instruction-heavy mixing beats filtering at scale, but the 33-task suite may over-weight the very capabilities that mixing targets. read the letter →

arxiv 2606.28551 v2 pith:HLF232VG submitted 2026-06-26 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords vision-languagemodelsdatasetcurationdatamixinginstructiontuningbenchmarkmultimodaltrainingopendatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces DataComp for VLMs, a benchmark that collects 160 datasets totaling 6T tokens across image captions, interleaved documents, text, and instructions. It runs controlled experiments varying filtering, mixing, formatting, and sampling on models from 1B to 8B parameters and token budgets up to 200B. Results show that the choice of mixture composition matters more than aggressive filtering, with instruction-heavy blends producing larger gains as scale increases. The resulting DCVLM-Baseline dataset trains an 8B VLM to 63.6 percent accuracy on a 33-task core evaluation suite, exceeding the prior open dataset FineVision by 5.4 points.

What carries the argument

The DCVLM benchmark, which standardizes curation operations (filtering, mixing, formatting, sampling) across fixed model sizes and token budgets and evaluates on a fixed suite of up to 52 downstream tasks in nine domains.

What would settle it

A caption-heavy mixture or a purely filtering-based curation strategy that achieves higher average accuracy than DCVLM-Baseline on the same 33-task core suite when trained at the 8B scale with 200B tokens.

Watch

Extended reading notes

Core claim

Data mixing, not filtering, is the dominant factor in building high-quality VLM training sets; instruction-heavy mixtures outperform caption-heavy ones, with the performance gap widening at larger model and data scales. The DCVLM-Baseline mixture derived from these experiments enables an 8B-parameter VLM trained on 200B tokens to reach 63.6 percent average accuracy across the 33-task core suite, a 5.4-point improvement over the previous state-of-the-art open VLM dataset FineVision.

Load-bearing premise

The selected downstream benchmarks adequately represent the full range of general VLM capabilities.

Editorial extensions

If this is right

  • Instruction-heavy mixtures deliver increasing returns as model size and token count grow.
  • Filtering alone yields smaller gains than careful composition of data types.
  • The DCVLM-Baseline dataset can be used directly to train stronger open VLMs without proprietary data.
  • Curation effort should prioritize mixing ratios over removal of individual low-quality examples.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Future work could test whether the same mixing preference holds when the evaluation suite is expanded to include more reasoning-heavy or long-context tasks.
  • The benchmark design makes it straightforward to measure whether new data sources improve performance mainly by changing the overall mixture balance.
  • Practitioners building custom VLM datasets may achieve comparable gains by reweighting existing public collections toward instruction data rather than collecting new filtered captions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces DataComp-VLM (DCVLM), a benchmark and corpus of 6T multimodal tokens from 160 datasets across four data types. It enables controlled experiments on curation operations (filtering, mixing, formatting, sampling) for 1B-8B VLMs trained on 6.25B-200B tokens, evaluated on up to 52 downstream benchmarks across 9 domains. The central empirical finding is that data mixing—not filtering—is the dominant factor, with instruction-heavy mixtures outperforming caption-heavy ones (gains widening at larger scales); the resulting DCVLM-Baseline yields an 8B VLM at 63.6% on the 33-task core suite (+5.4pp over FineVision).

Significance. If the results hold, the work supplies the first large-scale, controlled benchmark for VLM data curation and demonstrates that mixing strategies can be systematically optimized, with public release of the corpus, baseline, and evaluation suite as a concrete community resource. The scale of the experiments (multiple model sizes and token budgets) and the quantitative improvement over an existing SOTA open dataset are notable strengths.

major comments (2)
  1. [Evaluation] Evaluation section (and abstract): the claim that 'instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales' is supported only by relative performance on the 33-task core suite. The manuscript provides no evidence that this suite was constructed with held-out domains, that deltas were tested for sensitivity to task re-weighting, or that performance was measured on underrepresented categories such as pure captioning, OCR, or retrieval; this is load-bearing for the generalization that mixing is the key curation operation.
  2. [Experiments] Methods / Experiments: the abstract and results report specific accuracy gains (e.g., 63.6% and +5.4pp) without mention of error bars, multiple random seeds, or data-exclusion criteria for the downstream benchmarks. Because the central mixing-vs-filtering conclusion rests on these measured deltas, the absence of statistical characterization weakens verification of the reported improvements.
minor comments (1)
  1. [Abstract] The abstract states 'up to 52 downstream benchmarks' while the core suite is described as 33 tasks; clarify the exact overlap and selection criteria in the main text.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive feedback. We address the two major comments point by point below, indicating where revisions will be made.

read point-by-point responses
  1. Referee: [Evaluation] Evaluation section (and abstract): the claim that 'instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales' is supported only by relative performance on the 33-task core suite. The manuscript provides no evidence that this suite was constructed with held-out domains, that deltas were tested for sensitivity to task re-weighting, or that performance was measured on underrepresented categories such as pure captioning, OCR, or retrieval; this is load-bearing for the generalization that mixing is the key curation operation.

    Authors: The 33-task core suite was explicitly selected to ensure coverage across all 9 evaluation domains (including captioning, OCR, retrieval, and VQA), with the full 52-task results provided in the appendix showing consistent trends. We will add an explicit subsection on task selection methodology, domain balance, and sensitivity analysis to re-weighting in the revised manuscript to further support the generalization. revision: partial

  2. Referee: [Experiments] Methods / Experiments: the abstract and results report specific accuracy gains (e.g., 63.6% and +5.4pp) without mention of error bars, multiple random seeds, or data-exclusion criteria for the downstream benchmarks. Because the central mixing-vs-filtering conclusion rests on these measured deltas, the absence of statistical characterization weakens verification of the reported improvements.

    Authors: We agree that additional statistical detail would strengthen the presentation. All experiments use fixed seeds for reproducibility; the scale of the 8B/200B-token runs made multiple independent seeds computationally prohibitive. We will add a limitations paragraph on this point, report error bars for all smaller-scale ablations, and expand the existing description of benchmark data-exclusion criteria in Section 4. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical results on external benchmarks

full rationale

The paper's central claims rest on empirical measurements: models trained on curated mixtures are evaluated on a held-out suite of up to 52 downstream benchmarks across 9 domains. No equations, fitted parameters, or self-referential definitions appear in the derivation; the superiority of instruction-heavy mixing is reported as observed performance deltas (e.g., +5.4pp over FineVision), not as a quantity forced by the curation process itself. Self-citations, if present, are not load-bearing for the mixing-vs-filtering conclusion. The evaluation distribution is external to the training data construction, satisfying the condition for a self-contained empirical result.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Empirical benchmark paper with no mathematical derivations. No free parameters are fitted to the reported accuracy numbers. Relies only on standard machine-learning training assumptions.

assumptions (1)
  • standard math Standard i.i.d. sampling and gradient-based optimization assumptions used in large-scale model training.
    Implicit background for all reported training runs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DataComp-VLM: Improved Open Datasets for Vision-Language Models." pith.science (2026). https://pith.science/paper/HLF232VG

@misc{pith2026260628551,
  author       = {Pith},
  title        = {Pith review of: DataComp-VLM: Improved Open Datasets for Vision-Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HLF232VG}},
  note         = {Machine review of arXiv:2606.28551}
}
read the original abstract

Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we collect 160 datasets spanning four data types -- image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data -- into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B-8B models and 6.25B-200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-Baseline, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FineVision, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available at https://www.datacomp.ai/dcvlm/.

Figures

Figures reproduced from arXiv: 2606.28551 by the authors.

Figure 1
Figure 1. DCVLM-BASELINE outperforms open VLM training datasets. DCVLM-BASELINE (left) combines 160 sources as 10% image-caption pairs, 5% multimodal documents, 15% text-only, and 70% multimodal instruction-tuning data. On our 33-evaluation Core set (right), it outperforms existing datasets [12, 58, 310] across all scales. Notably, a 4B model trained on DCVLM-BASELINE for 100B tokens beats an 8B model trained on FINEVISION fo… view at source ↗
Figure 2
Figure 2. DCVLM allows researchers to construct effective multimodal datasets. Participants can choose one of four scales (small, medium, large, and x-large) according to their compute availability. We provide tools to format, filter, and mix the data pool so that participants can create their own datasets. The resulting datasets are then used to train an autoregressive VLM using a fixed training recipe. Models are comprehens… view at source ↗
Figure 3
Figure 3. Filtering rarely helps, but changing the data composition does move performance substantially. Established data filtering techniques do not significantly outperform a no-filter baseline. This observation holds consistently at both the small ( [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (35 more)
Figure 4
Figure 4. Figure 4: Upstream filtering leads to diminishing returns from additional (i.e., “downstream”) filtering. This failure is surprising, especially in light of strong results from prior works. We hypothesize this is because there is no significant noise to remove from our base pool…
Figure 5
Figure 5. Figure 5: Instruction-heavy mixtures scale better with compute. For the 1B model (left), the Instruction-heavy mix (red) starts as the worst mixture with 6.25B training tokens, but recovers quickly up to becoming the second-best with 25B training tokens. For the 2B model (middle…
Figure 6
Figure 6. Figure 6: Control experiments. (Left) Pretraining performance predicts post-SFT performance with near-perfect fidelity. (Right) Data mixture rankings are preserved when switching the LM backbone from Qwen2.5-Base to Qwen2.5-Instruct, verifying robustness of our results to choice…
Figure 7
Figure 7. Figure 7: DCVLM pool composition by data type. Share of total samples (top) vs. share of total multimodal tokens (bottom) for each of the four data types. The pool is dominated by image-caption pairs on both axes (83% tokens vs 74% samples). Text-only data exhibits the opposite …
Figure 7
Figure 7. Figure 7: DCVLM pool composition by data type. Share of total samples (top) vs. share of total multimodal tokens (bottom) for each of the four data types. The pool is dominated by image-caption pairs on both axes (83% tokens vs 74% samples). Text-only data exhibits the opposite …
Figure 8
Figure 8. Figure 8: Samples vs. multimodal tokens (log–log). Each marker is one of 160 datasets in our pool, colored by data type. Diagonal reference lines mark constant tokens-per-sample regimes (100, 1K, 10K). Image-caption datasets cluster tightly along the 1–2K tok/sample diagonal dri…
Figure 8
Figure 8. Figure 8: Samples vs. multimodal tokens (log–log). Each marker is one of 160 datasets in our pool, colored by data type. Diagonal reference lines mark constant tokens-per-sample regimes (100, 1K, 10K). Image-caption datasets cluster tightly along the 1–2K tok/sample diagonal dri…
Figure 9
Figure 9. Figure 9: Language distribution of DCVLM pool. Top-15 languages by per-sample top-1 prediction from Lingua (left, blue) and NLLB (right, red), shown on a log-scaled x-axis. The pool is dominated by English (91.1% Lingua / 92.8% NLLB), followed by Chinese (3.9% / 3.5%); the remai…
Figure 9
Figure 9. Figure 9: Language distribution of DCVLM pool. Top-15 languages by per-sample top-1 prediction from Lingua (left, blue) and NLLB (right, red), shown on a log-scaled x-axis. The pool is dominated by English (91.1% Lingua / 92.8% NLLB), followed by Chinese (3.9% / 3.5%); the remai…
Figure 10
Figure 10. Figure 10: Not all benchmarks monotonically increase with scale. We report the performance of 65 benchmarks (averaged over three runs at both the small and the medium scales) with a focus on their scaling behaviour. Average performance at the small scale is arranged on the x-axi…
Figure 10
Figure 10. Figure 10: Not all benchmarks monotonically increase with scale. We report the performance of 65 benchmarks (averaged over three runs at both the small and the medium scales) with a focus on their scaling behaviour. Average performance at the small scale is arranged on the x-axi…
Figure 11
Figure 11. Figure 11: Seed variance of the initial 65-benchmark pool. We conduct runs at small and medium scales, to measure the seed variance of benchmarks in our initial 65 candidates. We remove POPE [164], as it leads to a standard deviation of up to 16% over three runs at the small sca…
Figure 11
Figure 11. Figure 11: Seed variance of the initial 65-benchmark pool. We conduct runs at small and medium scales, to measure the seed variance of benchmarks in our initial 65 candidates. We remove POPE [164], as it leads to a standard deviation of up to 16% over three runs at the small sca…
Figure 12
Figure 12. Figure 12: Image-decontamination removal rate vs. SSCD similarity threshold. Fraction of pool images removed as a function of the similarity threshold for four sources. The shaded band between our threshold (0.75, green) and FineVision’s (0.95, red) is detected train/test overla…
Figure 12
Figure 12. Figure 12: Image-decontamination removal rate vs. SSCD similarity threshold. Fraction of pool images removed as a function of the similarity threshold for four sources. The shaded band between our threshold (0.75, green) and FineVision’s (0.95, red) is detected train/test overla…
Figure 13
Figure 13. Figure 13: Qualitative train/test image matches by SSCD similarity band. Each pair shows a training-pool image (left) and its top-1 match in the DCVLM-Extended evaluation suite (right). Rows are grouped by the maximum SSCD cosine similarity of the pair, and labels above each pai…
Figure 13
Figure 13. Figure 13: Qualitative train/test image matches by SSCD similarity band. Each pair shows a training-pool image (left) and its top-1 match in the DCVLM-Extended evaluation suite (right). Rows are grouped by the maximum SSCD cosine similarity of the pair, and labels above each pai…
Figure 14
Figure 14. Figure 14: Text-decontamination hyperparameter sweep. Per-dataset match rate (% of training queries flagged) across the three MinHash design choices. Signature size is immaterial (a), so we use 128 permutations. Longer n-grams miss short evaluation prompts and detect almost no o…
Figure 14
Figure 14. Figure 14: Text-decontamination hyperparameter sweep. Per-dataset match rate (% of training queries flagged) across the three MinHash design choices. Signature size is immaterial (a), so we use 128 permutations. Longer n-grams miss short evaluation prompts and detect almost no o…
Figure 15
Figure 15. Figure 15: Threshold selection via human annotation. True-positive (TPR, green) and false￾positive (FPR, red) rates of MinHash matches at each similarity bin, aggregated over seven annotators. Below ∼0.55 most matches are spurious (high FPR). The TPR climbs steeply just above it…
Figure 15
Figure 15. Figure 15: Threshold selection via human annotation. True-positive (TPR, green) and false￾positive (FPR, red) rates of MinHash matches at each similarity bin, aggregated over seven annotators. Below ∼0.55 most matches are spurious (high FPR). The TPR climbs steeply just above it…
Figure 16
Figure 16. Figure 16: Per-dataset removal rates of the full decontamination protocol. The 30 pool sources with the highest fraction of samples removed by the combined image-based (SSCD s ≥ 0.75) and text-based (MinHash Jaccard ≥ 0.55 + substring check) decontamination against the DCVLM￾Ext…
Figure 16
Figure 16. Figure 16: Per-dataset removal rates of the full decontamination protocol. The 30 pool sources with the highest fraction of samples removed by the combined image-based (SSCD s ≥ 0.75) and text-based (MinHash Jaccard ≥ 0.55 + substring check) decontamination against the DCVLM￾Ext…
Figure 17
Figure 17. Figure 17: Filtering rarely helps, but changing the data composition does move performance substantially (cont). Complementary experiments at the small scale of our benchmark confirm the outcome of Sec. 4.1: established quality filtering rarely provides significant gains over a …
Figure 17
Figure 17. Figure 17: Filtering rarely helps, but changing the data composition does move performance substantially (cont). Complementary experiments at the small scale of our benchmark confirm the outcome of Sec. 4.1: established quality filtering rarely provides significant gains over a …
Figure 18
Figure 18. Figure 18: Length-proportional sampling (T = 1) is near-optimal. Validation average as a function of sampling temperature T in p(d) ∝ len(d) 1/T . Both sharpening (T < 1) and flattening (T > 1) degrade performance relative, with the near-uniform T = 4 setting losing 4.1pp. We st…
Figure 18
Figure 18. Figure 18: Length-proportional sampling (T = 1) is near-optimal. Validation average as a function of sampling temperature T in p(d) ∝ len(d) 1/T . Both sharpening (T < 1) and flattening (T > 1) degrade performance relative, with the near-uniform T = 4 setting losing 4.1pp. We st…
Figure 19
Figure 19. Figure 19: Online filtering adds negligible overhead to training. Left: Runtime for 50M tokens (2 nodes, 8 GPUs, 1 filter) across rejection percentiles—total spread is 3.5%, within normal run￾to-run variance. Center: Same measurement at 500M tokens, where the spread shrinks to 0…
Figure 19
Figure 19. Figure 19: Online filtering adds negligible overhead to training. Left: Runtime for 50M tokens (2 nodes, 8 GPUs, 1 filter) across rejection percentiles—total spread is 3.5%, within normal run￾to-run variance. Center: Same measurement at 500M tokens, where the spread shrinks to 0…
Figure 20
Figure 20. Figure 20: The optimal pretraining data mixture is scale-dependent. At the small scale, performance is flat or slightly degrades as the instruction-tuning share grows; at the medium scale, the trend reverses, with performance improving monotonically and peaking at 65% instructio…
Figure 20
Figure 20. Figure 20: The optimal pretraining data mixture is scale-dependent. At the small scale, performance is flat or slightly degrades as the instruction-tuning share grows; at the medium scale, the trend reverses, with performance improving monotonically and peaking at 65% instructio…
Figure 21
Figure 21. Figure 21: Streaming best-fit sequence packing. (a) Incoming samples are assigned online to a pool of open buffers kept sorted from fullest to emptiest. A sample Si is placed in the first (hence fullest) buffer whose token budget L and image-tile budget M both still accommodate …
Figure 21
Figure 21. Figure 21: Streaming best-fit sequence packing. (a) Incoming samples are assigned online to a pool of open buffers kept sorted from fullest to emptiest. A sample Si is placed in the first (hence fullest) buffer whose token budget L and image-tile budget M both still accommodate …
Figure 22
Figure 22. Figure 22: replicates the analysis of Section 4.3 using Mammoth-VL-12M [89] as the SFT dataset in place of LLaVA-665K [151]. Results are fully consistent: pretraining and post-SFT scores remain near-perfectly correlated (Pearson r=0.99; Spearman ρ=0.99), and the pretraining rank…
Figure 22
Figure 22. Figure 22: replicates the analysis of Section 4.3 using Mammoth-VL-12M [89] as the SFT dataset in place of LLaVA-665K [151]. Results are fully consistent: pretraining and post-SFT scores remain near-perfectly correlated (Pearson r=0.99; Spearman ρ=0.99), and the pretraining rank…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Caption quality can be split into Coverage and Precision; in controlled training runs, Coverage predicts VLM understanding while Precision predicts T2I generation.

Reference graph

Works this paper leans on

299 extracted references · 299 canonical work pages · cited by 1 Pith paper

  1. [1]

    SemDeDup: Data-efficient learning at web-scale through semantic deduplication

    A. Abbas, K. Tirumala, D. Simig, S. Ganguli, and A. S. Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication.arXiv preprint arXiv:2303.09540, 2023. Cited on page 45

  2. [2]

    Abbas, E

    A. Abbas, E. Rusak, K. Tirumala, W. Brendel, K. Chaudhuri, and A. S. Morcos. Effective pruning of web-scale datasets based on complexity of concept clusters.arXiv preprint arXiv:2401.04578, 2024. Cited on page 45

  3. [3]

    Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

    A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, J. Bao, A. Benhaim, M. Cai, V . Chaudhary, C. Chen, et al. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras.arXiv preprint arXiv:2503.01743, 2025. Cited on page 2

  4. [4]

    Acharya, K

    M. Acharya, K. Kafle, and C. Kanan. TallyQA: Answering complex counting questions. In AAAI Conference on Artificial Intelligence (AAAI), 2019. Cited on pages 53 and 54

  5. [5]

    Agnolucci, L

    L. Agnolucci, L. Galteri, M. Bertini, and A. Del Bimbo. Arniqa: Learning distortion manifold for image quality assessment. InIEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 189–198, 2024. Cited on pages 45 and 75

  6. [6]

    Ainslie, J

    J. Ainslie, J. Lee-Thorp, M. De Jong, Y . Zemlyanskiy, F. Lebrón, and S. Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. InConference on Empirical Methods in Natural Language Processing (EMNLP), pages 4895–4901, 2023. Cited on page 48

  7. [7]

    S. N. Akter, S. Prabhumoye, E. Nyberg, M. Patwary, M. Shoeybi, Y . Choi, and B. Catanzaro. Front-loading reasoning: The synergy between pretraining and post-training data.arXiv preprint arXiv:2510.03264, 2025. Cited on page 8

  8. [8]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS), 35:23716–23736, 2022. Cited on pages 2 and 45

Show all 299 references
  1. [9]

    L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíˇcek, A. P. Lajarín, V . Srivastav, et al. Smollm2: When smol goes big–data-centric training of a small language model.arXiv preprint arXiv:2502.02737, 2025. Cited on pages 6 and 46

  2. [10]

    Allen-Zhu and Y

    Z. Allen-Zhu and Y . Li. Physics of language models: Part 3.1, knowledge storage and extraction.arXiv preprint arXiv:2309.14316, 2023. Cited on page 8

  3. [11]

    Amini, S

    A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y . Choi, and H. Hajishirzi. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In J. Burstein, C. Doran, and T. Solorio, editors,Proceedings of the 2019 Conference of the North American ...

  4. [12]

    X. An, Y . Xie, K. Yang, W. Zhang, X. Zhao, Z. Cheng, Y . Wang, S. Xu, C. Chen, D. Zhu, et al. Llava-onevision-1.5: Fully open framework for democratized multimodal training.arXiv preprint arXiv:2509.23661, 2025. Cited on pages 2, 3, 9, and 45. 12

  5. [13]

    Ankner, C

    Z. Ankner, C. Blakeney, K. Sreenivasan, M. Marion, M. L. Leavitt, and M. Paul. Perplexed by perplexity: Perplexity-based data pruning with small reference models.arXiv preprint arXiv:2405.20541, 2024. Cited on page 5

  6. [14]

    Awadalla, L

    A. Awadalla, L. Xue, O. Lo, M. Shu, H. Lee, E. Guha, M. Jordan, S. Shen, M. Awadalla, S. Savarese, et al. Mint-1t: Scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens.Advances in Neural Information Processing Systems (NeurIPS), 37: 36805–3...

  7. [15]

    C. Baek, R. P. Monti, D. Schwab, A. Abbas, R. Adiga, C. Blakeney, M. Böther, P. Burstein, A. G. Carranza, A. Deng, et al. The finetuner’s fallacy: When to pretrain with your finetuning data.arXiv preprint arXiv:2603.16177, 2026. Cited on page 8

  8. [16]

    S. Bai, Y . Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025. Cited on pages 2, 3, and 4

  9. [17]

    S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin. Qwen2.5-vl technical report, 2025. URLhttp...

  10. [18]

    Berasi, M

    D. Berasi, M. Farina, M. Mancini, and E. Ricci. Linear model merging unlocks simple and scalable multimodal data mixture optimization.arXiv preprint arXiv:2602.04937, 2026. Cited on pages 3 and 45

  11. [19]

    Bevli, S

    A. Bevli, S. Chaybouti, Y . Dahou, H. Hacid, N. D. Huynh, P. H. L. Khac, S. Narayan, W. R. Para, and A. Singh. Falcon perception.arXiv preprint arXiv:2603.27365, 2026. Cited on page 2

  12. [20]

    L. Beyer. On the speed of ViTs and CNNs.http://lb.eyer.be/a/vit-cnn-speed.html, 2024. Cited on page 68

  13. [21]

    Beyer, A

    L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, et al. Paligemma: A versatile 3b vlm for transfer.arXiv preprint arXiv:2407.07726, 2024. Cited on pages 2, 45, and 46

  14. [22]

    A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas. Scene text visual question answering. InIEEE/CVF International Conference on Computer Vision (ICCV), pages 4291–4301, 2019. Cited on pages 53 and 54

  15. [23]

    Bordt, S

    S. Bordt, S. Srinivas, V . Boreiko, and U. V on Luxburg. How much can we forget about data contamination?arXiv preprint arXiv:2410.03249, 2024. Cited on page 46

  16. [24]

    Breuel and WebDataset Contributors

    T. Breuel and WebDataset Contributors. WebDataset: A high-performance Python-based I/O system for large (and small) deep learning problems, with strong support for PyTorch. https://github.com/webdataset/webdataset, 2020. Cited on page 85

  17. [25]

    A. Z. Broder. On the resemblance and containment of documents. InProceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), pages 21–29. IEEE, 1997. Cited on pages 4, 46, and 69. 13

  18. [26]

    Cahyawijaya, H

    S. Cahyawijaya, H. Lovenia, J. R. A. Moniz, T. H. Wong, M. R. Farhansyah, T. T. Maung, F. Hudi, D. Anugraha, M. R. S. Habibi, M. R. Qorib, et al. Crowdsource, crawl, or generate? creating sea-vl, a multicultural vision-language dataset for southeast asia. InProceedings of the ...

  19. [27]

    Cao and J

    J. Cao and J. Xiao. An augmented benchmark dataset for geometric question answering through dual parallel text encoding. InInternational Conference on Computational Linguistics (COLING), 2022. Cited on pages 53 and 54

  20. [28]

    Carlini, D

    N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang. Quantifying memorization across neural language models. InInternational Conference on Learning Representations (ICLR), 2022. Cited on page 8

  21. [29]

    J. Carter. TextOCR-GPT4V: A re-captioning of TextOCR with GPT-4V, 2024. Hugging Face dataset card,https://huggingface.co/datasets/jimmycarter/textocr-gpt4v. Cited on pages 53 and 54

  22. [30]

    Chang, D

    S. Chang, D. Palzer, J. Li, E. Fosler-Lussier, and N. Xiao. MapQA: A dataset for question answering on choropleth maps.arXiv preprint arXiv:2211.08545, 2022. Cited on pages 53 and 54

  23. [31]

    Changpinyo, P

    S. Changpinyo, P. Sharma, N. Ding, and R. Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3558–3568, 2021. Cited on page 45

  24. [32]

    G. H. Chen, S. Chen, R. Zhang, J. Chen, X. Wu, Z. Zhang, Z. Chen, J. Li, X. Wan, and B. Wang. ALLaV A: Harnessing GPT4V-synthesized data for a lite vision-language model. arXiv preprint arXiv:2402.11684, 2024. Cited on pages 53 and 54

  25. [33]

    J. Chen, T. Li, J. Qin, P. Lu, L. Lin, C. Chen, and X. Liang. UniGeo: Unifying geometry logical reasoning via reformulating mathematical expression. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2022. Cited on pages 53 and 54

  26. [34]

    L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin. ShareGPT4V: Improving large multi-modal models with better captions. InEuropean Conference on Computer Vision (ECCV), pages 370–387. Springer, 2024. Cited on pages 53, 54, and 55

  27. [35]

    L. Chen, J. Li, X. Dong, P. Zhang, Y . Zang, Z. Chen, H. Duan, J. Wang, Y . Qiao, D. Lin, et al. Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. Cited on page 66

  28. [36]

    M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y . Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021. Cited on page 66

  29. [37]

    M. F. Chen, T. Murray, D. Heineman, M. Jordan, H. Hajishirzi, C. Ré, L. Soldaini, and K. Lo. Olmix: A framework for data mixing throughout lm development.arXiv preprint arXiv:2602.12237, 2026. Cited on pages 3, 45, and 93

  30. [38]

    W. Chen, M. Yin, M. Ku, P. Lu, Y . Wan, X. Ma, J. Xu, X. Wang, and T. Xia. Theoremqa: A theorem-driven question answering dataset. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7889–7901, 2023. Cited on page 66. 14

  31. [39]

    X. Chen, H. Fang, T.-Y . Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft COCO captions: Data collection and evaluation server.arXiv preprint arXiv:1504.00325, 2015. Cited on pages 53, 54, and 66

  32. [40]

    Y . Chen, S. Qian, H. Tang, X. Lai, Z. Liu, S. Han, and J. Jia. LongloRA: Efficient fine-tuning of long-context large language models. InInternational Conference on Learning Representations (ICLR), 2024. URLhttps://openreview.net/forum?id=6PmJoRfdaK. Cited on pages 53 and 54

  33. [41]

    Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling.arXiv preprint arXiv:2412.05271, 2024. Cited on pages 1, 4, 5, 47, 50, 52, 54, 6...

  34. [42]

    Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. Cited on p...

  35. [43]

    C. K. Chng, Y . Liu, Y . Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding, et al. ICDAR2019 robust reading challenge on arbitrary-shaped text - RRC-ArT. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53 and 54

  36. [44]

    J. H. Cho, A. Madotto, E. Mavroudi, T. Afouras, T. Nagarajan, M. Maaz, Y . Song, T. Ma, S. Hu, S. Jain, et al. Perceptionlm: Open-access data and models for detailed visual understanding. arXiv preprint arXiv:2504.13180, 2025. Cited on pages 2 and 86

  37. [45]

    Clark, J

    C. Clark, J. Zhang, Z. Ma, J. S. Park, M. Salehi, R. Tripathi, S. Lee, Z. Ren, C. D. Kim, Y . Yang, et al. Molmo2: Open weights and data for vision-language models with video understanding and grounding.arXiv preprint arXiv:2601.10611, 2026. Cited on pages 3, 45, and 77

  38. [47]

    Cobbe, V

    K. Cobbe, V . Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021. Cited on page 66

  39. [48]

    Conover, M

    M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin. Free Dolly: Introducing the world’s first truly open instruction- tuned LLM, 2023. Databricks Blog https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially...

  40. [49]

    Contributors

    O. Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023. Cited on page 62

  41. [50]

    M. R. Costa-Jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, et al. No language left behind: Scaling human-centered machine translation.arXiv preprint arXiv:2207.04672, 2022. Cited on page 59

  42. [51]

    E. Cui, Y . He, Z. Ma, Z. Chen, H. Tian, W. Wang, K. Li, Y . Wang, W. Wang, X. Zhu, L. Lu, T. Lu, Y . Wang, L. Wang, Y . Qiao, and J. Dai. Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o, 2024. URLhttps://sharegpt4o.github.io/. Cited on page 3. 15

  43. [52]

    G. Cui, L. Yuan, N. Ding, G. Yao, B. He, W. Zhu, Y . Ni, G. Xie, R. Xie, Y . Lin, Z. Liu, and M. Sun. UltraFeedback: Boosting language models with scaled AI feedback.International Conference on Machine Learning (ICML), 2024. Cited on pages 53 and 54

  44. [53]

    D. Dai, Y . Li, Y . Liu, M. Jia, Z. YuanHui, and G. Wang. 15M multimodal facial image-text dataset.arXiv preprint arXiv:2407.08515, 2024. Cited on pages 53 and 54

  45. [54]

    T. Dao. Flashattention-2: Faster attention with better parallelism and work partitioning.arXiv preprint arXiv:2307.08691, 2023. Cited on page 47

  46. [55]

    A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra. Visual dialog. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. Cited on pages 53 and 54

  47. [56]

    J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, et al. Large scale distributed deep networks.Advances in Neural Information Processing Systems (NeurIPS), 25, 2012. Cited on page 50

  48. [57]

    Deitke, C

    M. Deitke, C. Clark, S. Lee, R. Tripathi, Y . Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al. Molmo and pixmo: Open weights and open data for state-of-the-art vision- language models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...

  49. [58]

    A. S. Deshmukh, K. Chumachenko, T. Rintamaki, M. Le, T. Poon, D. M. Taheri, I. Karmanov, G. Liu, J. Seppanen, G. Chen, et al. Nvidia nemotron nano v2 vl.arXiv preprint arXiv:2511.03929, 2025. Cited on pages 2, 3, 9, 45, and 86

  50. [59]

    S. Diao, Y . Yang, Y . Fu, X. Dong, D. Su, M. Kliegl, Z. Chen, P. Belcak, Y . Suhara, H. Yin, et al. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training.arXiv preprint arXiv:2504.13161, 2025. Cited on pages 3, 45, and 77

  51. [60]

    N. Ding, Y . Chen, B. Xu, Y . Qin, S. Hu, Z. Liu, M. Sun, and B. Zhou. Enhancing chat language models by scaling high-quality instructional conversations. In H. Bouamor, J. Pino, and K. Bali, editors,Conference on Empirical Methods in Natural Language Processing (EMNLP), pages...

  52. [61]

    Dodge, M

    J. Dodge, M. Sap, A. Marasovi ´c, W. Agnew, G. Ilharco, D. Groeneveld, M. Mitchell, and M. Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. InProceedings of the 2021 conference on empirical methods in natural language processing, p...

  53. [62]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020. Cited on page 47

  54. [63]

    Douze, A

    M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou. The faiss library.IEEE Transactions on Big Data, 2025. Cited on page 68

  55. [64]

    H. Duan, J. Yang, Y . Qiao, X. Fang, L. Chen, Y . Liu, X. Dong, Y . Zang, P. Zhang, J. Wang, et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. InACM International Conference on Multimedia, pages 11198–11201, 2024. Cited on pages 2 and 62. 16

  56. [65]

    Evans, N

    T. Evans, N. Parthasarathy, H. Merzi ´c, and O. J. Henaff. Data curation via joint example selection further accelerates multimodal learning.Advances in Neural Information Processing Systems (NeurIPS), 37:141240–141260, 2024. Cited on pages 5 and 6

  57. [66]

    L. Fan, D. Krishnan, P. Isola, D. Katabi, and Y . Tian. Improving clip training with language rewrites.Advances in Neural Information Processing Systems (NeurIPS), 36:35544–35575, 2023. Cited on page 45

  58. [67]

    A. Fang, G. Ilharco, M. Wortsman, Y . Wan, V . Shankar, A. Dave, and L. Schmidt. Data determines distributional robustness in contrastive language image pre-training (clip). In International Conference on Machine Learning (ICML), pages 6216–6234. PMLR, 2022. Cited on page 1

  59. [68]

    A. Fang, A. M. Jose, A. Jain, L. Schmidt, A. Toshev, and V . Shankar. Data filtering networks. arXiv preprint arXiv:2309.17425, 2023. Cited on pages 5 and 45

  60. [69]

    A. Fang, H. Pouransari, M. Jordan, A. Toshev, V . Shankar, L. Schmidt, and T. Gunter. Datasets, documents, and repetitions: The practicalities of unequal data quality.arXiv preprint arXiv:2503.07879, 2025. Cited on page 8

  61. [70]

    L. Feng, G. R. Ghosal, J. M. Springer, Z. Zhong, and A. Raghunathan. Early data exposure improves robustness to subsequent fine-tuning.arXiv preprint arXiv:2605.12705, 2026. Cited on page 8

  62. [71]

    C. Fu, P. Chen, Y . Shen, Y . Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models.arXiv preprint arXiv:2306.13394, 2023. Cited on pages 65 and 66

  63. [72]

    X. Fu, Y . Hu, B. Li, Y . Feng, H. Wang, X. Lin, D. Roth, N. A. Smith, W.-C. Ma, and R. Krishna. Blink: Multimodal large language models can see but not perceive. InEuropean Conference on Computer Vision, pages 148–166. Springer, 2024. Cited on page 66

  64. [73]

    S. Y . Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al. Datacomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Systems (NeurIPS), 36:27092–27112, 2023. Cited o...

  65. [74]

    L. Gao. An empirical exploration in quality filtering of text data.arXiv preprint arXiv:2109.00698, 2021. Cited on page 7

  66. [75]

    L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027, 2020. Cited on page 46

  67. [76]

    Gervais, A

    P. Gervais, A. Fadeeva, and A. Maksai. Mathwriting: A dataset for handwritten mathematical expression recognition. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V .2, KDD ’25, page 5459–5469, New York, NY , USA, 2025. Association for Co...

  68. [77]

    Ghosal, V

    D. Ghosal, V . T. Y . Han, C. Y . Ken, and S. Poria. Are language models puzzle prodigies? Algorithmic puzzles unveil serious challenges in multimodal reasoning.arXiv preprint arXiv:2403.03864, 2024. Cited on pages 53 and 54. 17

  69. [78]

    Ghosh, S

    A. Ghosh, S. Dziadzio, A. Prabhu, V . Udandarao, S. Albanie, and M. Bethge. Onebench to test them all: Sample-level benchmarking over open-ended capabilities. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pag...

  70. [79]

    Ghosh, V

    A. Ghosh, V . Udandarao, T. Nguyen, M. Farina, M. Cherti, J. Jitsev, S. Oh, E. Ricci, L. Schmidt, and M. Bethge. Concept-aware batch sampling improves language-image pretraining.arXiv preprint arXiv:2511.20643, 2025. Cited on pages 1, 45, 78, and 79

  71. [80]

    Glaive-Code-Assistant, 2023

    Glaive AI. Glaive-Code-Assistant, 2023. https://huggingface.co/datasets/ glaiveai/glaive-code-assistant. Cited on pages 53 and 54

  72. [81]

    Goyal, P

    S. Goyal, P. Maini, Z. C. Lipton, A. Raghunathan, and J. Z. Kolter. Scaling laws for data filtering–data curation cannot be compute agnostic. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22702–22711, 2024. Cited on pages 4, 7, 46, and 82

  73. [82]

    Goyal, T

    Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. Cited on pages 53 and 54

  74. [83]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024. Cited on page 1

  75. [84]

    T. Gu, Z. Zhou, K. Huang, D. Liang, Y . Wang, H. Zhao, Y . Yao, X. Qiao, K. Wang, Y . Yang, et al. Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models.Advances in Neural Information Processing Systems, 37:7256–7295, 2024. Cited on page 66

  76. [85]

    T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y . Yacoob, et al. Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. InProceedings of the IEEE/CVF conference on com...

  77. [86]

    E. Guha, R. Marten, S. Keh, N. Raoof, G. Smyrnis, H. Bansal, M. Nezhurina, J. Mercat, T. Vu, Z. Sprague, et al. Openthoughts: Data recipes for reasoning models.arXiv preprint arXiv:2506.04178, 2025. Cited on page 65

  78. [87]

    D. Guo, F. Wu, F. Zhu, F. Leng, G. Shi, H. Chen, H. Fan, J. Wang, J. Jiang, J. Wang, et al. Seed1. 5-vl technical report.arXiv preprint arXiv:2505.07062, 2025. Cited on page 2

  79. [88]

    H. Guo, X. Qin, J. Liu, J. Han, J. Liu, and E. Ding. EATEN: Entity-aware attention for single shot visual text extraction. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53 and 54

  80. [89]

    J. Guo, T. Zheng, Y . Li, Y . Bai, B. Li, Y . Wang, K. Zhu, G. Neubig, W. Chen, and X. Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa...

  81. [90]

    Gupta, A

    A. Gupta, A. Vedaldi, and A. Zisserman. Synthetic data for text localisation in natural images. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016. Cited on pages 53 and 54. 18

  82. [91]

    Gurari, Q

    D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham. Vizwiz grand challenge: Answering visual questions from blind people. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 3608–3617, 2018. Cited on page 66

  83. [92]

    Y . Han, C. Zhang, X. Chen, X. Yang, Z. Wang, G. Yu, B. Fu, and H. Zhang. ChartLlama: A multimodal LLM for chart understanding and generation.arXiv preprint arXiv:2311.16483, 2023. Cited on pages 53 and 54

  84. [93]

    Hanu and Unitary team

    L. Hanu and Unitary team. Detoxify. Github. https://github.com/unitaryai/detoxify, 2020. Cited on page 89

  85. [94]

    C. He, Z. Jin, C. Xu, J. Qiu, B. Wang, W. Li, H. Yan, J. Wang, and D. Lin. Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models.arXiv preprint arXiv:2308.10755, 2023. Cited on pages 4, 53, 54, and 55

  86. [95]

    M. He, Y . Liu, Z. Yang, S. Zhang, C. Luo, F. Gao, Q. Zheng, Y . Wang, X. Zhang, and L. Jin. ICPR 2018 contest on robust reading for multi-type web images (MTWI). InInternational Conference on Pattern Recognition (ICPR), 2018. Cited on pages 53 and 54

  87. [96]

    X. He, Y . Zhang, L. Mou, E. Xing, and P. Xie. PathVQA: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020. Cited on pages 53 and 54

  88. [97]

    Heineman, V

    D. Heineman, V . Hofmann, I. Magnusson, Y . Gu, N. A. Smith, H. Hajishirzi, K. Lo, and J. Dodge. Signal and noise: A framework for reducing uncertainty in language model evaluation.arXiv preprint arXiv:2508.13144, 2025. Cited on page 5

  89. [98]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020. Cited on page 66

  90. [99]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset.arXiv preprint arXiv:2103.03874, 2021. Cited on page 66

  91. [100]

    Hernandez, T

    D. Hernandez, T. Brown, T. Conerly, N. DasSarma, D. Drain, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, T. Henighan, T. Hume, et al. Scaling laws and interpretability of learning from repeated data.arXiv preprint arXiv:2205.10487, 2022. Cited on page 8

  92. [101]

    Hessel, A

    J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y . Choi. Clipscore: A reference-free evaluation metric for image captioning. InConference on Empirical Methods in Natural Language Processing (EMNLP), pages 7514–7528, 2021. Cited on pages 3 and 45

  93. [102]

    R. Hong, W. Agnew, T. Kohno, and J. Morgenstern. Who’s in and who’s out? a case study of multimodal clip-filtering in datacomp. InProceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–17, 2024. Cited on page 94

  94. [103]

    W. Hong, W. Yu, X. Gu, G. Wang, G. Gan, H. Tang, J. Cheng, J. Qi, J. Ji, L. Pan, et al. Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning.arXiv preprint arXiv:2507.01006, 2025. Cited on page 2. 19

  95. [104]

    Honovich, T

    O. Honovich, T. Scialom, O. Levy, and T. Schick. Unnatural instructions: Tuning language models with (almost) no human labor. In A. Rogers, J. Boyd-Graber, and N. Okazaki, editors,Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1...

  96. [105]

    A. Hu, H. Xu, J. Ye, M. Yan, L. Zhang, B. Zhang, J. Zhang, Q. Jin, F. Huang, and J. Zhou. mplug-docowl 1.5: Unified structure learning for ocr-free document understanding.Findings of the Association for Computational Linguistics: EMNLP 2024, pages 3096–3120, 2024. Cited on pag...

  97. [106]

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models.Iclr, 1(2):3, 2022. Cited on page 49

  98. [107]

    Huang, Y

    Y . Huang, Y . Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y . Zhang, Y . Fu, et al. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.Advances in neural information processing systems, 36:62991–63010, 2023. Cited on page 66

  99. [108]

    Huang, K

    Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. Jawahar. ICDAR2019 competition on scanned receipt OCR and information extraction. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53 and 54

  100. [109]

    D. A. Hudson and C. D. Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6700–6709, 2019. Cited on pages 53, 54, and 66

  101. [110]

    Ionescu, H

    B. Ionescu, H. Müller, et al. Overview of the ImageCLEF 2024: Multimedia retrieval in medical applications. InInternational Conference of the Cross-Language Evaluation Forum for European Languages, 2024. Cited on pages 53 and 54

  102. [111]

    Jhamtani and T

    H. Jhamtani and T. Berg-Kirkpatrick. Learning to describe differences between pairs of similar images. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2018. Cited on pages 53 and 54

  103. [112]

    C. Jia, Y . Yang, Y . Xia, Y .-T. Chen, Z. Parekh, H. Pham, Q. Le, Y .-H. Sung, Z. Li, and T. Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning (ICML), pages 4904–4916. PMLR, 2021....

  104. [113]

    Y . Jia, J. Li, X. Yue, B. Li, P. Nie, K. Zou, and W. Chen. VisualWebInstruct: Scaling up multimodal instruction data through web search. In C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, editors,Conference on Empirical Methods in Natural Language Processing (EM...

  105. [114]

    Jiang, X

    D. Jiang, X. He, H. Zeng, C. Wei, M. Ku, Q. Liu, and W. Chen. Mantis: Interleaved multi- image instruction tuning.arXiv preprint arXiv:2405.01483, 2024. Cited on page 66

  106. [115]

    Jiang, K

    M. Jiang, K. Z. Liu, M. Zhong, R. Schaeffer, S. Ouyang, J. Han, and S. Koyejo. Investigating data contamination for pre-training language models.arXiv preprint arXiv:2401.06059, 2024. Cited on page 46. 20

  107. [116]

    Joshi, E

    M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, 2017...

  108. [117]

    Joshi, H

    S. Joshi, H. Yin, R. Adiga, H. Mongstad, A. Deng, A. Carranza, A. Fang, A. Abbas, A. Suri, B. Larsen, et al. 20/20 vision language models: A prescription for better vlms through data curation alone.arXiv preprint arXiv:2605.11405, 2026. Cited on page 2

  109. [118]

    Joshi, H

    S. Joshi, H. Yin, R. Adiga, R. Monti, A. Carranza, A. Fang, A. Deng, A. Abbas, B. Larsen, C. Blakeney, et al. Datbench: Discriminative, faithful, and efficient vlm evaluations.arXiv preprint arXiv:2601.02316, 2026. Cited on page 2

  110. [119]

    Kafle, B

    K. Kafle, B. Price, S. Cohen, and C. Kanan. DVQA: Understanding data visualizations via question answering. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5648–5656, 2018. Cited on pages 53 and 54

  111. [120]

    S. E. Kahou, V . Michalski, A. Atkinson, Á. Kádár, A. Trischler, and Y . Bengio. FigureQA: An annotated figure dataset for visual reasoning.arXiv preprint arXiv:1710.07300, 2017. Cited on pages 53 and 54

  112. [121]

    F. Kang, Y . Sun, B. Wen, S. Chen, D. Song, R. Mahmood, and R. Jia. Autoscale: Scale-aware data mixing for pre-training llms.arXiv preprint arXiv:2407.20177, 2024. Cited on pages 3 and 45

  113. [122]

    Kantharaj, R

    S. Kantharaj, R. T. Leong, X. Lin, A. Masry, M. Thakkar, E. Hoque, and S. Joty. Chart-to-text: A large-scale benchmark for chart summarization. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4005–4023, 2...

  114. [123]

    Karamcheti, S

    S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh. Prismatic vlms: Investigating the design space of visually-conditioned language models. InForty- first International Conference on Machine Learning, 2024. Cited on page 7

  115. [124]

    Karpathy and L

    A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3128– 3137, 2015. Cited on page 66

  116. [125]

    Kazemi, H

    M. Kazemi, H. Alvari, A. Anand, J. Wu, X. Chen, and R. Soricut. GeomVerse: A systematic evaluation of large models for geometric reasoning.arXiv preprint arXiv:2312.12241, 2023. Cited on pages 53 and 54

  117. [126]

    Kazemzadeh, V

    S. Kazemzadeh, V . Ordonez, M. Matten, and T. Berg. ReferItGame: Referring to objects in photographs of natural scenes. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2014. Cited on pages 53, 54, and 66

  118. [127]

    Kembhavi, M

    A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi. A diagram is worth a dozen images. InEuropean Conference on Computer Vision (ECCV), pages 235–251. Springer, 2016. Cited on pages 53, 54, and 66

  119. [128]

    Kembhavi, M

    A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi. Are you smarter than a sixth grader? Textbook question answering for multimodal machine comprehension. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. Cited on pages 53 and 54. 21

  120. [129]

    Kiela, H

    D. Kiela, H. Firooz, A. Mohan, V . Goswami, A. Singh, P. Ringshia, and D. Testuggine. The hateful memes challenge: Detecting hate speech in multimodal memes.Advances in Neural Information Processing Systems (NeurIPS), 2020. Cited on pages 53 and 54

  121. [130]

    G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park. OCR-free document understanding transformer. InEuropean Conference on Computer Vision (ECCV), 2022. Cited on pages 53 and 54

  122. [131]

    W. Kim, S. Chun, T. Kim, D. Han, and S. Yun. Hype: Hyperbolic entailment filtering for underspecified images and texts. InEuropean Conference on Computer Vision (ECCV), pages 247–265. Springer, 2024. Cited on page 45

  123. [132]

    Know-Saraswati-CoT: Chain-of-thought Sanskrit/Hindi reasoning dataset

    knowrohit07 and Knowledge Tech Team. Know-Saraswati-CoT: Chain-of-thought Sanskrit/Hindi reasoning dataset. https://huggingface.co/datasets/knowrohit07/ know-saraswati-cot, 2024. Hugging Face dataset card. Cited on pages 53 and 54

  124. [133]

    Kuang, W

    J. Kuang, W. Hua, D. Liang, M. Yang, D. Jiang, B. Ren, and X. Bai. Visual information extraction in the wild: practical dataset and end-to-end solution.International Conference on Document Analysis and Recognition (ICDAR), 2023. Cited on pages 53 and 54

  125. [134]

    Kumar, A

    A. Kumar, A. Raghunathan, R. Jones, T. Ma, and P. Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution.arXiv preprint arXiv:2202.10054, 2022. Cited on page 8

  126. [135]

    Kuznetsova, H

    A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al. The Open Images dataset V4: Unified image classification, object detection, and visual relationship detection at scale.International Journal of Comp...

  127. [136]

    Kwiatkowski, J

    T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al. Natural questions: a benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019. Cite...

  128. [137]

    G. Lai, Q. Xie, H. Liu, Y . Yang, and E. Hovy. Race: Large-scale reading comprehension dataset from examinations. InProceedings of the 2017 conference on empirical methods in natural language processing, pages 785–794, 2017. Cited on page 66

  129. [138]

    Lambert, J

    N. Lambert, J. Morrison, V . Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V . Miranda, A. Liu, N. Dziri, S. Lyu, et al. Tulu 3: Pushing frontiers in open language model post-training.arXiv preprint arXiv:2411.15124, 2024. Cited on page 68

  130. [139]

    J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific Data, 5(1):1–10, 2018. Cited on pages 53 and 54

  131. [140]

    Laurençon, L

    H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. Rush, D. Kiela, et al. Obelics: An open web-scale filtered dataset of interleaved image-text documents.Advances in Neural Information Processing Systems (NeurIPS), 36:71683–7170...

  132. [141]

    Laurençon, L

    H. Laurençon, L. Tronchon, M. Cord, and V . Sanh. What matters when building vision- language models?Advances in Neural Information Processing Systems (NeurIPS), 37: 87874–87907, 2024. Cited on page 45. 22

  133. [142]

    Laurençon, L

    H. Laurençon, L. Tronchon, M. Cord, and V . Sanh. What matters when building vision- language models?Advances in Neural Information Processing Systems (NeurIPS), 37: 87874–87907, 2024. Cited on page 3

  134. [143]

    Laurençon, A

    H. Laurençon, A. Marafioti, V . Sanh, and L. Tronchon. Building and better understanding vision-language models: insights and future directions, 2024. URL https://arxiv.org/ abs/2408.12637. Cited on pages 53 and 54

  135. [144]

    Laurençon, L

    H. Laurençon, L. Tronchon, and V . Sanh. Unlocking the conversion of web screenshots into html code with the websight dataset, 2024. URLhttps://arxiv.org/abs/2403.09029. Cited on pages 53 and 54

  136. [145]

    K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini. Deduplicating training data makes language models better. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424–8445, 2...

  137. [146]

    Lee and S

    S. Lee and S. Hwang. Selective training for large vision language models via visual information gain.arXiv preprint arXiv:2602.17186, 2026. Cited on page 6

  138. [147]

    Lerner, O

    P. Lerner, O. Ferret, C. Guinaudeau, H. Le Borgne, R. Besançon, J. G. Moreno, and J. Lovon- Melgarejo. ViQuAE, a dataset for knowledge-based visual question answering about named entities. InACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 202...

  139. [148]

    B. Li, R. Wang, G. Wang, Y . Ge, Y . Ge, and Y . Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension.arXiv preprint arXiv:2307.16125, 2023. Cited on page 66

  140. [149]

    B. Li, Y . Ge, Y . Chen, Y . Ge, R. Zhang, and Y . Shan. Seed-bench-2-plus: Benchmarking multimodal large language models with text-rich visual comprehension.arXiv preprint arXiv:2404.16790, 2024. Cited on page 66

  141. [150]

    B. Li, Z. Lin, W. Peng, J. d. D. Nyandwi, D. Jiang, Z. Ma, S. Khanuja, R. Krishna, G. Neubig, and D. Ramanan. Naturalbench: Evaluating vision-language models on natural adversarial samples.Advances in Neural Information Processing Systems, 37:17044–17068, 2024. Cited on page 66

  142. [151]

    B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liu, et al. Llava-onevision: Easy visual task transfer.arXiv preprint arXiv:2408.03326, 2024. Cited on pages 3 and 87

  143. [152]

    C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao. LLaV A-Med: Training a large language-and-vision assistant for biomedicine in one day. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023. Ci...

  144. [153]

    H. Li, Y . Zhang, F. Koto, Y . Yang, H. Zhao, Y . Gong, N. Duan, and T. Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese. InFindings of the Association for Computational Linguistics: ACL 2024, pages 11260–11285, 2024. Cited on page 66

  145. [154]

    H. Li, Y . Chen, S. Miao, Q. Dong, J. Chen, Y . Hu, J. Chen, M. Qin, Y . Wu, Y . Zhou, et al. Legalone: a family of foundation models for reliable legal reasoning.arXiv preprint arXiv:2602.00642, 2026. Cited on page 8. 23

  146. [155]

    J. Li, D. Li, S. Savarese, and S. Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InInternational Conference on Machine Learning (ICML), pages 19730–19742. PMLR, 2023. Cited on page 45

  147. [156]

    J. LI, E. Beeching, L. Tunstall, B. Lipkin, R. Soletskyi, S. Huang, K. Rasul, L. Yu, A. Q. Jiang, Z. Shen, Z. Qin, B. Dong, L. Zhou, Y . Fleureau, G. Lample, and S. Polu. NuminaMath 1.5: Second iteration of NuminaMath, 2024. Hugging Face dataset card https://huggingface. co/da...

  148. [157]

    J. LI, E. Beeching, L. Tunstall, et al. NuminaMath-TIR: Tool-integrated reasoning math dataset, 2024. Hugging Face dataset card https://huggingface.co/datasets/AI-MO/ NuminaMath-TIR. Cited on pages 53 and 54

  149. [158]

    J. Li, A. Fang, G. Smyrnis, M. Ivgi, M. Jordan, S. Gadre, H. Bansal, E. Guha, S. Keh, K. Arora, et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems (NeurIPS), 37:14200–14282, 2024. Cited o...

  150. [159]

    J. Li, J. Chen, Y . Qu, S. Xu, Z. Lin, J. Zhu, B. Xu, W. Tan, P. Fu, J. Ju, et al. Xiaomi mimo-vl-miloco technical report.arXiv preprint arXiv:2512.17436, 2025. Cited on page 2

  151. [160]

    L. Li, Y . Wang, R. Xu, P. Wang, X. Feng, L. Kong, and Q. Liu. Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models. InAnnual Meeting of the Association for Computational Linguistics (ACL), 2024. Cited on pages 53 and 54

  152. [161]

    Q. Li, Z. Chen, W. Wang, W. Wang, S. Ye, Z. Jin, G. Chen, Y . He, Z. Gao, E. Cui, et al. Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text. arXiv preprint arXiv:2406.08418, 2024. Cited on pages 4, 53, 54, 55, and 76

  153. [162]

    Li and N

    S. Li and N. Tajbakhsh. SciGraphQA: A large-scale synthetic multi-turn question-answering dataset for scientific graphs.arXiv preprint arXiv:2308.03349, 2023. Cited on pages 53 and 54

  154. [163]

    X. Li, H. Tu, M. Hui, Z. Wang, B. Zhao, J. Xiao, S. Ren, J. Mei, Q. Liu, H. Zheng, et al. What if we recaption billions of web images with llama-3?arXiv preprint arXiv:2406.08478, 2024. Cited on pages 45 and 78

  155. [164]

    Y . Li, Y . Du, K. Zhou, J. Wang, X. Zhao, and J.-R. Wen. Evaluating object hallucination in large vision-language models. InProceedings of the 2023 conference on empirical methods in natural language processing, pages 292–305, 2023. Cited on pages 65 and 66

  156. [165]

    W. Lian, G. Wang, B. Goodson, E. Pentland, A. Cook, C. V ong, and "Teknium". Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023. URL https://https://huggingface.co/Open-Orca/SlimOrca. Cited on pages 4, 53, and 54

  157. [166]

    H. Lin, V . Hosu, and D. Saupe. Kadid-10k: A large-scale artificially distorted iqa database. In2019 Eleventh International Conference on Quality of Multimedia Experience (QoMEX), pages 1–3. IEEE, 2019. Cited on page 75

  158. [167]

    T.-Y . Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. InEuropean Conference on Computer Vision (ECCV), pages 740–755. Springer, 2014. Cited on page 66. 24

  159. [168]

    A. D. Lindström and S. S. Abraham. CLEVR-Math: A dataset for compositional language, visual and mathematical reasoning. InInternational Workshop on Neural-Symbolic Learning and Reasoning (NeSy), 2022. Cited on pages 53 and 54

  160. [169]

    Liu, L.-M

    B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y . Yang, and X.-M. Wu. SLAKE: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. InIEEE International Symposium on Biomedical Imaging (ISBI), 2021. Cited on pages 53 and 54

  161. [170]

    C.-L. Liu, F. Yin, D.-H. Wang, and Q.-F. Wang. CASIA online and offline Chinese handwriting databases.International Conference on Document Analysis and Recognition (ICDAR), 2011. Cited on pages 53 and 54

  162. [171]

    F. Liu, G. Emerson, and N. Collier. Visual spatial reasoning.Transactions of the Association for Computational Linguistics (TACL), 11:635–651, 2023. Cited on pages 53, 54, and 66

  163. [172]

    F. Liu, K. Lin, L. Li, J. Wang, Y . Yacoob, and L. Wang. Mitigating hallucination in large multi-modal models via robust instruction tuning.International Conference on Learning Representations (ICLR), 2024. Cited on pages 53 and 54

  164. [173]

    F. Liu, X. Wang, W. Yao, J. Chen, K. Song, S. Cho, Y . Yacoob, and D. Yu. Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...

  165. [174]

    F. Liu, W. Zhou, B. Liu, P. Guo, Z. Wang, B. Zhang, Y . Zhang, Y . Yu, X. Zhou, and T. Wang. Infolaw: Information scaling laws for large language models with quality-weighted mixture data and repetition.arXiv preprint arXiv:2605.02364, 2026. Cited on page 8

  166. [175]

    H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning.Advances in Neural Information Processing Systems (NeurIPS), 36:34892–34916, 2023. Cited on pages 1, 2, 3, 7, 8, 45, 47, and 86

  167. [176]

    H. Liu, C. Li, Y . Li, and Y . J. Lee. Improved baselines with visual instruction tuning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 26296– 26306, 2024. Cited on pages 53 and 54

  168. [177]

    H. Liu, C. Li, Y . Li, B. Li, Y . Zhang, S. Shen, and Y . J. Lee. Llavanext: Improved reasoning, ocr, and world knowledge, 2024. Cited on pages 2, 4, and 45

  169. [178]

    J. Liu, T. Ou, Y . Song, Y . Qu, W. Lam, C. Xiong, W. Chen, G. Neubig, and X. Yue. Harnessing webpage uis for text-rich visual understanding, 2024. URLhttps://arxiv.org/abs/2410. 13824. Cited on pages 53 and 54

  170. [179]

    Q. Liu, X. Zheng, N. Muennighoff, G. Zeng, L. Dou, T. Pang, J. Jiang, and M. Lin. Regmix: Data mixture as regression for language model pre-training, 2024. Cited on pages 3, 45, and 93

  171. [180]

    X. Liu, Y . Zhu, J. Gu, Y . Lan, C. Yang, and Y . Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. InEuropean Conference on Computer Vision, pages 386–403. Springer, 2024. Cited on page 66. 25

  172. [181]

    Y . Liu, Y . Cao, Z. Gao, W. Wang, Z. Chen, W. Wang, H. Tian, L. Lu, X. Zhu, T. Lu, et al. Mminstruct: A high-quality multi-modal instruction tuning dataset with extensive diversity. Science China Information Sciences, 67(12):220103, 2024. Cited on pages 53 and 54

  173. [182]

    Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu, et al. Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision (ECCV), pages 216–233. Springer, 2024. Cited on page 66

  174. [183]

    Y . Liu, Z. Li, M. Huang, B. Yang, W. Yu, C. Li, X.-C. Yin, C.-L. Liu, L. Jin, and X. Bai. Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 2024. Cited on page 66

  175. [184]

    Z. Liu, T. Chu, Y . Zang, X. Dong, P. Zhang, Z. Yang, Y . Duan, D. Lin, Y . Wang, and J. Wang. MMDU: A multi-turn multi-image dialog understanding benchmark and instruction-tuning dataset for LVLMs. InAdvances in Neural Information Processing Systems (NeurIPS), 2024. Cited on ...

  176. [185]

    Chinese OCR dataset

    longmaodata. Chinese OCR dataset. https://huggingface.co/datasets/ longmaodata/Chinese-OCR, 2024. Hugging Face dataset card. Cited on pages 53 and 54

  177. [186]

    Longpre, L

    S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y . Tay, D. Zhou, Q. V . Le, B. Zoph, J. Wei, and A. Roberts. The flan collection: Designing data and methods for effective instruction tuning. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, edi...

  178. [187]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Sgdr: Stochastic gradient descent with warm restarts.arXiv preprint arXiv:1608.03983, 2016. Cited on page 50

  179. [188]

    Loshchilov and F

    I. Loshchilov and F. Hutter. Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017. Cited on pages 4 and 50

  180. [189]

    H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding.arXiv preprint arXiv:2403.05525, 2024. Cited on page 2

  181. [190]

    P. Lu, R. Gong, S. Jiang, L. Qiu, S. Huang, X. Liang, and S.-C. Zhu. Inter-GPS: Interpretable geometry problem solving with formal language and symbolic reasoning. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2021. Cited on p...

  182. [191]

    P. Lu, L. Qiu, J. Chen, T. Xia, Y . Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu. IconQA: A new benchmark for abstract diagram understanding and visual language reasoning. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2021. Cit...

  183. [192]

    P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems (NeurIPS), 35:2507–2521, 2022. Cited on pa...

  184. [193]

    P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K.-W. Chang, M. Galley, and J. Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023. Cited on page 66

  185. [194]

    P. Lu, L. Qiu, K.-W. Chang, Y . N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. In International Conference on Learning Representations (ICLR), 2023. Cited on pages 53 and 54

  186. [195]

    Z. Luo, C. Xu, P. Zhao, Q. Sun, X. Geng, W. Hu, C. Tao, J. Ma, Q. Lin, and D. Jiang. WizardCoder: Empowering code large language models with Evol-Instruct.International Conference on Learning Representations (ICLR), 2024. Cited on pages 53 and 54

  187. [196]

    W. Ma, H. Chen, G. Zhang, Y .-C. Chou, J. Chen, C. de Melo, and A. Yuille. 3dsrbench: A comprehensive 3d spatial reasoning benchmark. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 6924–6934, 2025. Cited on page 66

  188. [197]

    Madaan, A

    L. Madaan, A. K. Singh, R. Schaeffer, A. Poulton, S. Koyejo, P. Stenetorp, S. Narang, and D. Hupkes. Quantifying variance in evaluation benchmarks.arXiv preprint arXiv:2406.10229, 2024. Cited on page 5

  189. [198]

    Magar and R

    I. Magar and R. Schwartz. Data contamination: From memorization to exploitation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 157–165, 2022. Cited on page 46

  190. [199]

    Mahmoud, M

    A. Mahmoud, M. Elhoushi, A. Abbas, Y . Yang, N. Ardalani, H. Leather, and A. S. Morcos. Sieve: Multimodal dataset pruning using image captioning models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22423–22432, 2024. Cited on page 45

  191. [200]

    Maini, S

    P. Maini, S. Goyal, Z. C. Lipton, J. Z. Kolter, and A. Raghunathan. T-mars: Improving visual representations by circumventing text feature learning.arXiv preprint arXiv:2307.03132, 2023. Cited on page 45

  192. [201]

    Mao, C.-W

    C. Mao, C.-W. Xie, C. Zhong, H. Deng, J. Zhao, J. Xiao, J. Xing, J. Zhang, J. Zhou, J. Zhang, et al. Wan-image: Pushing the boundaries of generative visual intelligence.arXiv preprint arXiv:2604.19858, 2026. Cited on page 3

  193. [202]

    H. Mao, M. Cheung, and J. She. Deepart: Learning joint representations of visual arts. In ACM International Conference on Multimedia, pages 1183–1191, 2017. Cited on pages 53 and 54

  194. [203]

    Marafioti, O

    A. Marafioti, O. Zohar, M. Farré, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Tazi, et al. Smolvlm: Redefining small and efficient multimodal models.arXiv preprint arXiv:2504.05299, 2025. Cited on page 2

  195. [204]

    Marino, M

    K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. Cited on pages 53, 54, and 66

  196. [205]

    Marti and H

    U.-V . Marti and H. Bunke. The IAM-database: an English sentence database for offline handwriting recognition.International Journal on Document Analysis and Recognition (IJDAR), 5:39–46, 2002. Cited on pages 53 and 54. 27

  197. [206]

    Masry, X

    A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguistics: ACL 2022, pages 2263–2279, 2022. Cited on pages 53, 54, and 66

  198. [207]

    Masry, P

    A. Masry, P. Kavehzadeh, D. X. Long, E. Hoque, and S. Joty. UniChart: A universal vision- language pretrained model for chart comprehension and reasoning. InConference on Empirical Methods in Natural Language Processing (EMNLP), 2023. Cited on pages 53 and 54

  199. [208]

    Masry, M

    A. Masry, M. Thakkar, A. Bajaj, A. Kartha, E. Hoque, and S. Joty. Chartgemma: Visual instruction-tuning for chart reasoning in the wild.Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, pages 625–643, 2025. Cited on pages 53 and 54

  200. [209]

    Mathew, D

    M. Mathew, D. Karatzas, and C. Jawahar. Docvqa: A dataset for vqa on document images. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021. Cited on pages 53, 54, and 66

  201. [210]

    Mathew, V

    M. Mathew, V . Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar. Infographicvqa. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 1697–1706, 2022. Cited on pages 53, 54, and 66

  202. [211]

    KOpen-Hermes-25: Korean translation of OpenHermes-2.5, 2024

    Maywell. KOpen-Hermes-25: Korean translation of OpenHermes-2.5, 2024. Hugging Face dataset cardhttps://huggingface.co/datasets/maywell/ko_Ultrafeedback_ binarized. Cited on pages 53 and 54

  203. [212]

    Mazumder, C

    M. Mazumder, C. Banbury, X. Yao, B. Karlaš, W. Gaviria Rojas, S. Diamos, G. Diamos, L. He, A. Parrish, H. R. Kirk, et al. Dataperf: Benchmarks for data-centric ai development.Advances in Neural Information Processing Systems (NeurIPS), 36:5320–5347, 2023. Cited on pages 3 and 45

  204. [213]

    McKinzie, Z

    B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, A. Belyi, et al. Mm1: methods, analysis and insights from multimodal llm pre-training. In European Conference on Computer Vision (ECCV), pages 304–323. Springer, 2024. Cited on pages...

  205. [214]

    Methani, P

    N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar. Plotqa: Reasoning over scientific plots. InIEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 1527–1536, 2020. Cited on pages 53 and 54

  206. [215]

    Mishra, S

    A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty. OCR-VQA: Visual question answering by reading text in images. InInternational Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on pages 53, 54, and 66

  207. [216]

    Mitra, H

    A. Mitra, H. Khanpour, C. Rosset, and A. Awadallah. Orca-math: Unlocking the potential of slms in grade school math, 2024. URLhttps://arxiv.org/abs/2402.14830. Cited on pages 53 and 54

  208. [217]

    Mizrahi, A

    D. Mizrahi, A. B. L. Larsen, J. Allardice, S. Petryk, Y . Gorokhov, J. Li, A. Fang, J. Gardner, T. Gunter, and A. Dehghan. Language models improve when pretraining data matches target tasks.arXiv preprint arXiv:2507.12466, 2025. Cited on pages 4, 6, 7, 46, and 82

  209. [218]

    Mohri, J

    C. Mohri, J. Duchi, and T. Hashimoto. A bitter lesson for data filtering.arXiv preprint arXiv:2605.19407, 2026. Cited on page 7. 28

  210. [219]

    Muennighoff, A

    N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel. Scaling data-constrained language models.Advances in Neural Information Processing Systems (NeurIPS), 36:50358–50376, 2023. Cited on page 8

  211. [220]

    V . K. Nagaraja, V . I. Morariu, and L. S. Davis. Modeling context between objects for referring expression understanding. InEuropean Conference on Computer Vision (ECCV), pages 792–

  212. [221]

    Cited on pages 66 and 67

    Springer, 2016. Cited on pages 66 and 67

  213. [222]

    Nezhurina, T

    M. Nezhurina, T. Porian, G. Pucceti, T. Kerssies, R. Beaumont, M. Cherti, and J. Jitsev. Scaling laws for robust comparison of open foundation language-vision models and datasets.arXiv preprint arXiv:2506.04598, 2025. Cited on pages 4 and 7

  214. [223]

    H. Ngo, M. Deitke, M. Bartelds, S. Pratt, J. Gardner, M. Jordan, and L. Schmidt. Olmoasr: Open models and data for training robust speech recognition models.arXiv preprint arXiv:2508.20869, 2025. Cited on page 1

  215. [224]

    Nguyen, V

    H. Nguyen, V . May, H. Raj, M. Nezhurina, Y . Wang, Y . Luo, M. C. Vu, T. Nakamura, K. Tsui, V . K. Nguyen, et al. Mixturevitae: Open web-scale pretraining dataset with high quality instruction and reasoning data built from permissive-first text sources.arXiv preprint arXiv:25...

  216. [225]

    Nguyen, G

    T. Nguyen, G. Ilharco, M. Wortsman, S. Oh, and L. Schmidt. Quality not quantity: On the interaction between dataset design and robustness of clip.Advances in Neural Information Processing Systems (NeurIPS), 35:21455–21469, 2022. Cited on page 1

  217. [226]

    Nguyen, S

    T. Nguyen, S. Y . Gadre, G. Ilharco, S. Oh, and L. Schmidt. Improving multimodal datasets with image captioning.Advances in Neural Information Processing Systems (NeurIPS), 36: 22047–22069, 2023. Cited on page 45

  218. [227]

    Nguyen, M

    T. Nguyen, M. Wallingford, S. Santy, W.-C. Ma, S. Oh, L. Schmidt, P. W. Koh, and R. Krishna. Multilingual diversity improves vision-language representations.Advances in Neural Information Processing Systems (NeurIPS), 37:91430–91459, 2024. Cited on pages 45 and 59

  219. [228]

    T. OLMo, P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y . Gu, S. Huang, M. Jordan, et al. 2 olmo 2 furious.arXiv preprint arXiv:2501.00656, 2024. Cited on pages 3, 5, 45, and 77

  220. [229]

    T. Olmo, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, et al. Olmo 3.arXiv preprint arXiv:2512.13961, 2025. Cited on page 5

  221. [230]

    Chat markup language (ChatML)

    OpenAI. Chat markup language (ChatML). https://github.com/openai/ openai-python/blob/main/chatml.md, 2022. Accessed: 29 April 2026. Cited on page 75

  222. [231]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023. Cited on pages 45 and 46

  223. [232]

    Parashar, Z

    S. Parashar, Z. Lin, T. Liu, X. Dong, Y . Li, D. Ramanan, J. Caverlee, and S. Kong. The neglected tails in vision-language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12988–12997, 2024. Cited on page 60. 29

  224. [233]

    Penedo, Q

    G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only.arXiv preprint arXiv:2306.01116, 2023. Cited on page 1

  225. [234]

    Penedo, H

    G. Penedo, H. Kydlí ˇcek, A. Lozhkov, M. Mitchell, C. Raffel, L. V on Werra, T. Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems (NeurIPS), 37:30811–30849, 2024. Cited on pages 1, 3, 5, 45,...

  226. [235]

    Z. Peng, W. Wang, L. Dong, Y . Hao, S. Huang, S. Ma, and F. Wei. Kosmos-2: Grounding multimodal large language models to the world.arXiv preprint arXiv:2306.14824, 2023. Cited on pages 53, 54, and 55

  227. [236]

    Pizzi, S

    E. Pizzi, S. D. Roy, S. N. Ravindra, P. Goyal, and M. Douze. A self-supervised descriptor for image copy detection. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14532–14542, 2022. Cited on pages 4 and 68

  228. [237]

    Ponomarenko, O

    N. Ponomarenko, O. Ieremeiev, V . Lukin, K. Egiazarian, L. Jin, J. Astola, B. V ozel, K. Chehdi, M. Carli, F. Battisti, et al. Color image database tid2013: Peculiarities and preliminary results. InEuropean workshop on visual information processing (EUVIP), pages 106–111. IEEE...

  229. [238]

    Pouget, L

    A. Pouget, L. Beyer, E. Bugliarello, X. Wang, A. P. Steiner, X. Zhai, and I. Alabdulmohsin. No filter: Cultural and socioeconomic diversity in contrastive vision-language models.Advances in Neural Information Processing Systems (NeurIPS), 37:106474–106496, 2024. Cited on page 59

  230. [239]

    Pramanick, R

    S. Pramanick, R. Chellappa, and S. Venugopalan. SPIQA: A dataset for multimodal question answering on scientific papers.Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2024. Cited on pages 53 and 54

  231. [240]

    R. Qiao, Q. Tan, G. Dong, M. MinhuiWu, C. Sun, X. Song, J. Wang, Z. Gongque, S. Lei, Y . Zhang, et al. We-math: Does your large multimodal model achieve human-like mathematical reasoning? InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics...

  232. [241]

    Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, ...

  233. [242]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. InInternational Conference on Machine Learning (ICML), pages 8748–8763. PmLR, 2021. Cited...

  234. [243]

    Rajani, L

    N. Rajani, L. Tunstall, E. Beeching, N. Lambert, A. M. Rush, and T. Wolf. No Robots, 2023. https://huggingface.co/datasets/HuggingFaceH4/no_robots. Cited on pages 53 and 54

  235. [244]

    Rajbhandari, J

    S. Rajbhandari, J. Rasley, O. Ruwase, and Y . He. Zero: Memory optimizations toward training trillion parameter models. InSC20: international conference for high performance computing, networking, storage and analysis, pages 1–16. IEEE, 2020. Cited on page 50. 30

  236. [245]

    leetcode

    RayBernard. leetcode. https://huggingface.co/datasets/RayBernard/leetcode,

  237. [246]

    Cited on pages 53 and 54

    Hugging Face dataset card. Cited on pages 53 and 54

  238. [247]

    Rodriguez, X

    J. Rodriguez, X. Jian, S. S. Panigrahi, T. Zhang, A. Feizi, A. Puri, A. Kalkunte Suresh, F. Savard, A. Masry, S. Nayak, R. Awal, M. Massoud, A. Abaskohi, Z. Li, S. Wang, P.-A. Noël, M. L. Richter, S. Vadacchino, S. Agarwal, S. Biswas, S. Shanian, Y . Zhang, N. Bolger, K. MacDo...

  239. [248]

    K. Roth, V . Udandarao, S. Dziadzio, A. Prabhu, M. Cherti, O. Vinyals, O. Hénaff, S. Albanie, M. Bethge, and Z. Akata. A practitioner’s guide to continual multimodal pretraining.arXiv preprint arXiv:2408.14471, 2024. Cited on page 50

  240. [249]

    Sainz, I

    O. Sainz, I. García-Ferrero, A. Jacovi, J. A. Campos, Y . Elazar, E. Agirre, Y . Goldberg, W.-L. Chen, J. Chim, L. Choshen, et al. Data contamination report from the 2024 conda shared task. InProceedings of the 1st Workshop on Data Contamination (CONDA), pages 41–56, 2024. Cit...

  241. [250]

    Sakaguchi, R

    K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y . Choi. Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021. Cited on page 66

  242. [251]

    Saxena, P

    R. Saxena, P. Minervini, and F. Keller. PosterSum: A multimodal benchmark for scientific poster summarization. In K. Inui, S. Sakti, H. Wang, D. F. Wong, P. Bhattacharyya, B. Banerjee, A. Ekbal, T. Chakraborty, and D. P. Singh, editors,Proceedings of the 14th International Joi...

  243. [252]

    Schaeffer, J

    R. Schaeffer, J. Kazdan, B. Abbasi, K. Z. Liu, B. Miranda, A. Ahmed, F. Berez, A. Puri, S. Biderman, N. Mireshghallah, et al. Quantifying the effect of test set contamination on generative evaluations.arXiv preprint arXiv:2601.04301, 2026. Cited on page 46

  244. [253]

    Schuhmann, R

    C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in Neural Information Processing Systems (NeurIPS), 35:252...

  245. [254]

    Schwenk, A

    D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowledge. InEuropean Conference on Computer Vision (ECCV), 2022. Cited on pages 53 and 54

  246. [255]

    S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar. KVQA: Knowledge-aware visual question answering. InAAAI Conference on Artificial Intelligence (AAAI), 2019. Cited on pages 53 and 54

  247. [256]

    S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun. Objects365: A large-scale, high-quality dataset for object detection. InIEEE/CVF International Conference on Computer Vision (ICCV), 2019. Cited on pages 53 and 54

  248. [257]

    Shapourian, K

    H. Shapourian, K. Hejazi, O. M. Sule, and B. Millidge. Zaya1-vl-8b technical report.arXiv preprint arXiv:2605.08560, 2026. Cited on page 3. 31

  249. [258]

    N. Shazeer. Glu variants improve transformer.arXiv preprint arXiv:2002.05202, 2020. Cited on page 48

  250. [259]

    W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pa...

  251. [260]

    Shinoda, K

    R. Shinoda, K. Saito, S. Tanaka, T. Hirasawa, and Y . Ushiku. SBS Figures: Pre-training figure QA from stage-by-stage synthesized images.arXiv preprint arXiv:2412.17606, 2024. Cited on pages 53 and 54

  252. [261]

    Shukor, L

    M. Shukor, L. Bethune, D. Busbridge, D. Grangier, E. Fini, A. El-Nouby, and P. Ablin. Scaling laws for optimal data mixtures.arXiv preprint arXiv:2507.09404, 2025. Cited on page 7

  253. [262]

    Sidorov, R

    O. Sidorov, R. Hu, M. Rohrbach, and A. Singh. TextCaps: A dataset for image captioning with reading comprehension. InEuropean Conference on Computer Vision (ECCV), pages 742–758. Springer, 2020. Cited on pages 53 and 54

  254. [263]

    Singh, V

    A. Singh, V . Natarajan, M. Shah, Y . Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach. Towards vqa models that can read. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8317–8326, 2019. Cited on pages 53, 54, and 66

  255. [264]

    Singh, G

    A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner. TextOCR: Towards large- scale end-to-end reasoning for arbitrary-shaped scene text. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Cited on pages 53 and 54

  256. [265]

    e. a. Singh. Persian synthetic OCR dataset (ParSynth-OCR-200K), 2021. Hugging Face dataset card. Cited on pages 53 and 54

  257. [266]

    Sorscher, R

    B. Sorscher, R. Geirhos, S. Shekhar, S. Ganguli, and A. Morcos. Beyond neural scaling laws: beating power law scaling via data pruning.Advances in Neural Information Processing Systems (NeurIPS), 35:19523–19536, 2022. Cited on page 45

  258. [267]

    Steiner, A

    A. Steiner, A. S. Pinto, M. Tschannen, D. Keysers, X. Wang, Y . Bitton, A. Gritsenko, M. Minderer, A. Sherbondy, S. Long, et al. Paligemma 2: A family of versatile vlms for transfer.arXiv preprint arXiv:2412.03555, 2024. Cited on page 2

  259. [268]

    Stiennon, L

    N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. V oss, A. Radford, D. Amodei, and P. F. Christiano. Learning to summarize with human feedback. InAdvances in Neural Information Processing Systems (NeurIPS), 2020. Cited on pages 53 and 54

  260. [269]

    D. Su, K. Kong, Y . Lin, J. Jennings, B. Norick, M. Kliegl, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguisti...

  261. [270]

    J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. Cited on page 48

  262. [271]

    Sun, D.-W

    H.-L. Sun, D.-W. Zhou, Y . Li, S. Lu, C. Yi, Q.-G. Chen, Z. Xu, W. Luo, K. Zhang, D.-C. Zhan, et al. Parrot: Multilingual visual instruction tuning.arXiv preprint arXiv:2406.02539, 2024. Cited on page 66. 32

  263. [272]

    T. Sun, X. Zhang, Z. He, P. Li, Q. Cheng, X. Liu, H. Yan, Y . Shao, Q. Tang, S. Zhang, et al. MOSS: An open conversational large language model.Machine Intelligence Research, 2024. Cited on pages 53 and 54

  264. [273]

    Y . Sun, Z. Ni, C.-K. Chng, Y . Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas, et al. ICDAR2019 competition on large-scale street view text with partial labeling - RRC-LSVT. In International Conference on Document Analysis and Recognition (ICDAR), 2019. Cited on ...

  265. [274]

    Tanaka, K

    R. Tanaka, K. Nishida, and S. Yoshida. VisualMRC: Machine reading comprehension on document images. InAAAI Conference on Artificial Intelligence (AAAI), 2021. Cited on pages 53 and 54

  266. [275]

    B. J. Tang, A. Boggust, and A. Satyanarayan. VisText: A benchmark for semantically rich chart captioning.Annual Meeting of the Association for Computational Linguistics (ACL), 2023. Cited on pages 53 and 54

  267. [276]

    J. Tang, Q. Liu, Y . Ye, J. Lu, S. Wei, A.-L. Wang, C. Lin, H. Feng, Z. Zhao, Y . Wang, et al. Mtvqa: Benchmarking multilingual text-centric visual question answering. InFindings of the Association for Computational Linguistics: ACL 2025, pages 7748–7763, 2025. Cited on page 66

  268. [277]

    C. Team, Z. Yue, Z. Lin, Y . Song, W. Wang, S. Ren, S. Gu, S. Li, P. Li, L. Zhao, L. Li, K. Bao, H. Tian, H. Zhang, G. Wang, D. Zhu, Cici, C. He, B. Ye, B. Shen, Z. Zhang, Z. Jiang, Z. Zheng, Z. Song, Z. Luo, Y . Yu, Y . Wang, Y . Tian, Y . Tu, Y . Yan, Y . Huang, X. Wang, X. ...

  269. [278]

    G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. bastien Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. T...

  270. [279]

    K. Team, A. Du, B. Yin, B. Xing, B. Qu, B. Wang, C. Chen, C. Zhang, C. Du, C. Wei, et al. Kimi-vl technical report.arXiv preprint arXiv:2504.07491, 2025. Cited on page 2

  271. [280]

    T. M. A. Team. Mai-thinking-1: Building a hill-climbing machine. Technical report, Microsoft AI, 2026. URLhttps://microsoft.ai/pdf/mai-thinking-1.pdf. Cited on pages 4 and 82

  272. [281]

    S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. Advances in Neural Information Processing Systems (NeurIPS), 37:87310–87356, 2024. Cited on pages 2, ...

  273. [282]

    S. Tong, Z. Liu, Y . Zhai, Y . Ma, Y . LeCun, and S. Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9568–9578, 2024. Cited on page 66

  274. [283]

    T. H. Trinh and Q. V . Le. A simple method for commonsense reasoning.arXiv preprint arXiv:1806.02847, 2018. Cited on page 46

  275. [284]

    Tschannen, A

    M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y . Xia, B. Mustafa, et al. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:250...

  276. [285]

    Y . Tuo, W. Xiang, J.-Y . He, Y . Geng, and X. Xie. Anytext: Multilingual visual text generation and editing. InInternational Conference on Learning Representations (ICLR), 2024. URL https://openreview.net/forum?id=ezBH9WE9s2. Cited on pages 53 and 54

  277. [286]

    zero-shot

    V . Udandarao, A. Prabhu, A. Ghosh, Y . Sharma, P. H. Torr, A. Bibi, S. Albanie, and M. Bethge. No" zero-shot" without exponential data: Pretraining concept frequency determines multimodal model performance.Advances in Neural Information Processing Systems (NeurIPS), 37:61735–...

  278. [287]

    Udandarao, Z

    V . Udandarao, Z. Lu, X. Chang, Y . Wang, V . Z. Yao, A. M. Jose, F. Faghri, J. Gardner, and C.-C. Chiu. Data-centric lessons to improve speech-language pretraining.arXiv preprint arXiv:2510.20860, 2025. Cited on page 1

  279. [288]

    Udandarao, N

    V . Udandarao, N. Parthasarathy, M. F. Naeem, T. Evans, S. Albanie, F. Tombari, Y . Xian, A. Tonioni, and O. J. Hénaff. Active data curation effectively distills large-scale multimodal models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14422...

  280. [289]

    Ustalov, N

    D. Ustalov, N. Pavlichenko, S. Koshelev, D. Likhobaba, and A. Smirnova. Toloka visual question answering benchmark.arXiv preprint arXiv:2309.16511, 2023. Cited on pages 53, 54, and 66

  281. [290]

    Van Horn, O

    G. Van Horn, O. Mac Aodha, Y . Song, Y . Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie. The iNaturalist species classification and detection dataset. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. Cited on pages 53 and 54. 34

  282. [291]

    A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie. COCO-Text: Dataset and benchmark for text detection and recognition in natural images. InarXiv preprint arXiv:1601.07140, 2016. Cited on pages 53 and 54

  283. [292]

    H. V . V o, V . Khalidov, T. Darcet, T. Moutakanni, N. Smetanin, M. Szafraniec, H. Touvron, C. Couprie, M. Oquab, A. Joulin, et al. Automatic data curation for self-supervised learning: A clustering-based approach.arXiv preprint arXiv:2405.15613, 2024. Cited on page 45

  284. [293]

    B. Wang, G. Li, X. Zhou, Z. Chen, T. Grossman, and Y . Li. Screen2words: Automatic mobile ui summarization with multimodal learning. InThe 34th Annual ACM Symposium on User Interface Software and Technology, pages 498–510, 2021. Cited on pages 53 and 54

  285. [294]

    F. Wang, X. Fu, J. Y . Huang, Z. Li, Q. Liu, X. Liu, M. D. Ma, N. Xu, W. Zhou, K. Zhang, et al. Muirbench: A comprehensive benchmark for robust multi-image understanding.arXiv preprint arXiv:2406.09411, 2024. Cited on page 66

  286. [295]

    J. Wang, L. Meng, Z. Weng, B. He, Z. Wu, and Y .-G. Jiang. To see is to believe: Prompting gpt- 4v for better visual instruction tuning, 2023. URL https://arxiv.org/abs/2311.07574. Cited on pages 53 and 54

  287. [296]

    J. Wang, Y . Wang, G. Xu, J. Zhang, Y . Gu, H. Jia, J. Wang, H. Xu, M. Yan, J. Zhang, et al. Amber: An llm-free multi-dimensional benchmark for mllms hallucination evaluation.arXiv preprint arXiv:2311.07397, 2023. Cited on page 66

  288. [297]

    J. Wang, P. Zhang, T. Chu, Y . Cao, Y . Zhou, T. Wu, B. Wang, C. He, and D. Lin. V3Det: Vast vocabulary visual detection dataset. InIEEE/CVF International Conference on Computer Vision (ICCV), 2023. Cited on pages 53 and 54

  289. [298]

    K. Wang, J. Pan, W. Shi, Z. Lu, H. Ren, A. Zhou, M. Zhan, and H. Li. Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024. Cited on page 66

  290. [299]

    P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. Cited on pages 5, 62, 78, and 79

  291. [300]

    W. Wang, Y . Ren, H. Luo, T. Li, C. Yan, Z. Chen, W. Wang, Q. Li, L. Lu, X. Zhu, et al. The all-seeing project V2: Towards general relation comprehension of the open world. InEuropean Conference on Computer Vision (ECCV), 2024. Cited on pages 53, 54, and 66

Pith tools

Reviewed July 2, 2026 · model on record in the stance chip above.