REVIEW 4 major objections 5 minor 1 cited by
Benchmarking Large and Small MLLMs
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A four-model benchmark of multimodal language models finds GPT-4o the strongest overall, with small models competitive only in narrow recognition and prompt-specific tasks.
desk verdict A broad but statistically fragile capability map; the qualitative pattern may hold, but a copied table row and tiny samples undermine the quantitative rankings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that carries the argument is a multi-part evaluation protocol rather than a single theorem. It organizes 48 tasks into four input formats—single image-text pairs, multi-image with text, multi-frame with text, and interleaved image-text—and maps each task to a capability such as understanding, reasoning, assessment, or interaction. Quantitative evaluations are anchored to existing datasets (MSCOCO, FSC-147, BLINK, ChartInsights, MVTec AD) plus a new TableInsights set with 166 table questions, while correctness is judged by three annotators; when a model's output was unsatisfactory, the authors refined the prompt themselves. This protocol is what lets the paper claim a systematic comparison instead of an anecdotal one.
What would settle it
Run the same comparison on a pre-registered set of 200 chart-reasoning and 200 spatial-reasoning questions with fixed prompts and independent blind scoring by three or more annotators; if LLaVA-NeXT or Phi-3-Vision match GPT-4o within a few percentage points on those tasks, the paper's central boundary claim would collapse.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the capability boundary between large and small MLLMs is real and task-dependent. In most understanding and reasoning tasks, GPT-4o ranks first, with GPT-4V close behind; in specific recognition tasks, LLaVA-NeXT and Phi-3-Vision can reach comparable accuracy, sometimes beating the large models, as with celebrity recognition for LLaVA-NeXT and several chart and prompt-format tasks for Phi-3-Vision. When the task asks for multi-step inference or fine-grained interpretation, the small models' accuracy falls sharply, in several reasoning tasks to zero. The paper also identifies common failure cases shared by all four models, especially in spatial localization, where F1 scores stay below 0.1, and in abstract visual reasoning such as Raven's progressive matrices.
Load-bearing premise
The load-bearing assumption is that small, hand-collected sample sets (sometimes 5-28 items per task) and the authors' own correctness judgments, made without reported inter-annotator agreement, are representative and fair enough to rank the four models.
Editorial extensions
If this is right
- Small MLLMs can serve in narrow recognition and inspection jobs—logo, emotion, and celebrity identification, and helmet counting when an external detector supplies crops—at lower cost and faster speed.
- Large models remain the default choice for reasoning-heavy multimodal tasks such as chart and table question answering, visual math, Raven's progressive matrices, and multi-frame temporal ordering.
- Model choice should be task-specific rather than size-based, since Phi-3-Vision beats GPT-4V on several chart-reasoning prompt formats and LLaVA-NeXT leads on celebrity recognition.
- Even the best model in the study, GPT-4o, shows a clear capability ceiling on precise object localization and abstract spatial reasoning, so these tasks remain open for future work.
- Adding a reference image of a defect-free product improves large models' anomaly detection but not small models' performance, suggesting a boundary in how small models exploit additional visual context.
Reading between the lines
- Not tested in the paper: whether the small-model gap on reasoning tasks comes from model size or from differences in training data and alignment, since the two small models differ on both axes; a size-matched, data-matched ablation would separate these causes.
- A practical extension the paper only implies: small models could be embedded in constrained recognition pipelines while large models handle chart/table QA, visual math, and temporal reasoning, creating a two-tier deployment strategy.
- Because the authors refined prompts when outputs were unsatisfactory, the absolute accuracy percentages should be read as prompt-dependent; a broader prompt search on the same tasks might narrow the measured gap on specific skills.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper benchmarks two large proprietary MLLMs (GPT-4V, GPT-4o) and two smaller open models (LLaVA-NeXT, Phi-3-Vision) across 48 tasks organized by input format: single image-text pairs, multiple images, multi-frame sequences, and interleaved image-text inputs. The main claims are that small models can match large models in specific recognition or narrow scenarios but lag substantially on complex reasoning, and that GPT-4o is the best overall model. The evaluation combines reused samples from external benchmarks with newly collected samples, uses three annotators to judge correctness, and includes both quantitative tables and qualitative failure-case analyses.
Significance. If the quantitative findings were reliable, this would be a useful practical guide: it would tell deployers where small models suffice and where only large models are adequate. The paper's strengths are its broad task taxonomy, the inclusion of multiple input formats, and the qualitative failure analysis, which highlights recurring weaknesses such as spatial localization and fine-grained counting. Credit is also due for grounding parts of the evaluation in existing datasets (MSCOCO, FSC-147, ChartInsights, BLINK) rather than relying solely on self-constructed examples. However, the quantitative evidence as reported is not strong enough to support the headline ranking claims, and at least one table contains an internal inconsistency that must be resolved before the numbers can be trusted.
major comments (4)
- [§2.1.1, §2.1.2, §6.2.2] The positive half of the central claim—that small MLLMs achieve comparable performance in specific scenarios—is not statistically supported. The evidence for this claim consists of post-hoc selections of tasks with very small samples, such as celebrity recognition (n=15), safety inspection (n=5), logo recognition in the wild (n=18), and guided prompting (n=15). No confidence intervals, significance tests, or multiple-comparison corrections are reported; scanning 48 tasks, a few nominal small-model wins are expected by chance. For example, in the safety-inspection task with n=5, a single response flip changes Phi-3-Vision's reported 100% accuracy to 80%, and a two-response flip changes it to 60%. The authors should either provide error bars and a pre-specified or corrected analysis of which tasks test the small-model claim, or explicitly reframe these observations as descriptive rather than as confirmatory evidence.
- [Tables 2.4 and 2.11] The GPT-4V row in Table 2.11 (TableInsights) lists the exact same ten per-task values as the GPT-4V row in Table 2.4 (ChartInsights)—39.55, 60.29, 65.87, 32.03, 73.96, 41.73, 53.69, 67.75, 88.91, 73.44—while the reported overall accuracy differs (50.47 versus 57.20). Identical per-task percentages across two different datasets are extremely unlikely and strongly suggest that one row was copied from the other. The authors must regenerate and verify all quantitative tables, and the discrepancy must be explained before any model-ranking conclusion can be accepted.
- [§2.1.2, Scene Understanding] The scene-understanding evaluation uses inconsistent denominators. The text states that 12 scene images were collected and three models are scored over 12, but Phi-3-Vision is reported as [2 + 2 × 0.5]/14. Either Phi-3-Vision was evaluated on a different sample size than the other models, or the printed denominator is a typo. In both cases the reported accuracy needs correction and the evaluation protocol should be stated uniformly.
- [§1.1 and throughout §2] The manuscript states that three annotators judged model outputs, but no inter-annotator agreement is reported. Many tasks use 0.5 partial-credit scores (e.g., scene text recognition, multilingual translation, multimodal commonsense), so a single annotator's disagreement can shift a reported accuracy by several percentage points and can flip a claimed small-model win. The authors should report agreement statistics, release the annotation rubric and raw outputs, or at least quantify the sensitivity of the rankings to individual annotator judgments.
minor comments (5)
- [§1.2, §6 intro] The model name 'LLaV A-NeXT' appears with an extra space throughout; the Section 6 introduction also contains 'LLaV A-Vision', which should be 'LLaVA-NeXT'.
- [§2.1.2] The word 'illustarted' should be 'illustrated'.
- [§2.1.4] The sentence 'We also nake an interesting finding' should read 'We also make an interesting finding'.
- [§2.2.1] The phrase 'the mmetric evaluation' should read 'the metric evaluation'.
- [General] Reproducibility details are missing: the authors should report API model versions and access dates, the exact random-sampling procedure for samples drawn from MSCOCO and FSC-147, and the annotator instructions and scoring rubric.
Circularity Check
No circularity: the paper is an empirical benchmark whose claims summarize external evaluations rather than deriving results from their own inputs.
full rationale
This manuscript is a benchmark study, not a derivation. Its central claims (‘small MLLMs can achieve comparable performance to large models in specific scenarios but lag significantly in complex tasks’) are empirical summaries of measured accuracies on external datasets (MSCOCO, FSC-147, ChartInsights, BLINK) and on newly collected samples. No parameter is fitted from one subset of data and then reported as a prediction of a closely related quantity; no quantity is defined in terms of another quantity in a way that makes a reported result true by construction; and no load-bearing conclusion is justified by a self-citation or by a uniqueness theorem imported from the authors’ prior work. The only experimenter involvement described is prompt refinement for tasks with unsatisfactory outputs (Section 1.2), which is a validity concern about fairness and reproducibility, not circularity. Similarly, the small sample sizes (e.g., 5 safety-inspection cases, 15 celebrity cases) and post-hoc highlighting of small-model wins are statistical-evidence concerns, not circular-reasoning concerns. The apparent duplication of the GPT-4V row in Tables 2.4 and 2.11 with different overall accuracies is a data-integrity problem, not a circularity problem. Consequently, there are no circular steps to report, and the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Human annotators' binary judgments are treated as ground truth for correctness, with three annotators but no agreement metric.
- domain assumption Small convenience samples (often 5 to 28 items per task) are sufficient to rank model capabilities.
- domain assumption Proprietary API model outputs are treated as stable point estimates.
- domain assumption The scoring rubric with 0.5 partial credit is a valid ordinal measure.
Cite this review
Pith. "Pith review of Benchmarking Large and Small MLLMs." pith.science (2026). https://pith.science/paper/ULMO4LPU
@misc{pith2026250104150,
author = {Pith},
title = {Pith review of: Benchmarking Large and Small MLLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/ULMO4LPU}},
note = {Machine review of arXiv:2501.04150}
}
read the original abstract
Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior quality and capabilities across diverse tasks. However, their deployment faces significant challenges, including slow inference, high computational cost, and impracticality for on-device applications. In contrast, the emergence of small MLLMs, exemplified by the LLava-series models and Phi-3-Vision, offers promising alternatives with faster inference, reduced deployment costs, and the ability to handle domain-specific scenarios. Despite their growing presence, the capability boundaries between large and small MLLMs remain underexplored. In this work, we conduct a systematic and comprehensive evaluation to benchmark both small and large MLLMs, spanning general capabilities such as object recognition, temporal reasoning, and multimodal comprehension, as well as real-world applications in domains like industry and automotive. Our evaluation reveals that small MLLMs can achieve comparable performance to large models in specific scenarios but lag significantly in complex tasks requiring deeper reasoning or nuanced understanding. Furthermore, we identify common failure cases in both small and large MLLMs, highlighting domains where even state-of-the-art models struggle. We hope our findings will guide the research community in pushing the quality boundaries of MLLMs, advancing their usability and effectiveness across diverse applications.
Figures
Figures from the paper (69 more)
Forward citations
Cited by 1 Pith paper
-
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
CompSlider learns to synthesize image-conditioning latents from multiple attribute sliders at once, aiming for more disentangled and structure-preserving multi-attribute control in text-to-image generation.
Reference graph
Works this paper leans on
-
[3]
arXiv preprint arXiv:2307.01952
Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952. A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. V oss, A. Radford, M. Chen, and I. Sutskever. 2021. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092. V . Ranjan, U. Sharma, T. Nguyen, and M. Hoai. 2021. Learning to count everything. I...
arXiv 2021
-
[4]
arXiv preprint arXiv:2309.15112
Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112. Y .-T. Zheng, M. Zhao, Y . Song, H. Adam, U. Buddemeier, A. Bissacco, F. Brucher, T.-S. Chua, and H. Neven. 2009. Tour the world: building a web-scale landmark recognition engine. In 2009 IEEE Conference on Computer Vis...
arXiv 2009
-
[2017]
In Proceedings of the IEEE conference on computer vision and pattern recognition, pp
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2901–2910. D. Kang, Z. Ma, and A. B. Chan. 2018. Beyond counting: comparisons of density maps for crowd analysis tasks—counting, detection, and tracking. IEEE Transactions on Circuits...
arXiv 2018
-
[2023]
Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3): 8. A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, E. Valveny, C. Jawahar, and D. Karatzas. 2019. Scene text visual question answering. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4291–4301. L. B...
arXiv 2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.