Pith. sign in

REVIEW 3 major objections 4 minor 62 references

Affordance Benchmark for MLLMs

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that even the best multimodal large language models perceive object affordances far below human level, with the top model scoring 18.05% exact match against 81.25–85.34% for humans on its 2,000-question A4Bench.

desk verdict A worthwhile affordance benchmark, but the headline 18-vs-85 gap is partly an artifact of adversarial filtering; still deserves review. read the letter →

arxiv 2506.00893 v2 pith:DOQPUIIX submitted 2025-06-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords affordanceperceptionmultimodallargelanguagemodelsbenchmarkexactmatchaccuracyconstitutivetransformativehumanperformancecomparisonGibsontheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

A4Bench is a new benchmark of 2,000 multimodal question-answer pairs designed to test whether multimodal large language models perceive what objects afford—what actions an object or environment makes possible. The paper evaluates 17 proprietary and open-source models and finds that all of them score far below human performance, with the best model, Gemini-2.0-Pro, at 18.05% overall exact match accuracy versus humans at 81.25–85.34%. The benchmark splits questions into constitutive affordance (static, inherent object properties, 1,282 pairs) and transformative affordance (dynamic, contextual, cultural, or individual-dependent affordances, 718 pairs), and the models are especially weak on transformative affordance. The authors conclude that current MLLMs are still poor at affordance perception, which matters because robust affordance understanding is needed for safe and intuitive interaction in robotics, autonomous systems, and other embodied AI applications.

What carries the argument

The central object is A4Bench, a benchmark constructed from 2,000 image-question-answer pairs organized along two dimensions: Constitutive Affordance (1,282 pairs across daily life, natural and rural, urban and industrial, and medical and scientific environments) and Transformative Affordance (718 pairs covering misleading, time-dependent, cross-cultural, and individual-variant affordances). The benchmark uses a multi-select question format where the number of correct answers is hidden, explicit object names are replaced with 'this object' in some items, and answers are randomized and evaluated by exact match. The key mechanism is the pairing of diverse, context-rich images with questions that require reasoning about action possibilities, combined with a human-MLLM iterative revision process intended to ensure that the questions are genuinely difficult for models.

What would settle it

Have a new, independent panel of human annotators re-label a random sample of the transformative-affordance questions, especially the cross-cultural and individual-variant subsets, without seeing the original answers. If inter-annotator agreement on those subsets is low or near chance, then the fixed ground-truth labels are not stable across people, and the reported human-model performance gap would not be a well-defined measurement.

Watch

Extended reading notes

Core claim

The central discovery is that current multimodal large language models, despite strong performance on many vision-language tasks, perceive affordance at a level dramatically below human performance. On A4Bench, the top-performing model (Gemini-2.0-Pro) achieves only 18.05% exact match accuracy, while human best and worst scores are 85.34% and 81.25%. Proprietary models generally outperform open-source models, yet all models fall far short of humans, particularly on transformative affordance categories such as misleading affordance, time-dependent affordance, cross-cultural affordance, and individual-variant affordance. The paper argues that this gap reveals a fundamental deficiency in models' understanding of object-human-environment interactions, and it positions A4Bench as a diagnostic tool to measure and guide progress toward more context-aware, safe multimodal AI.

Load-bearing premise

The benchmark's labeled correct answers are treated as objective and unambiguous even for affordances that the paper itself says vary by culture or by individual, and the paper does not explain how annotator disagreements were resolved.

Editorial extensions

If this is right

  • If the benchmark is accurate, current MLLMs cannot be relied upon for tasks that require understanding what objects afford, such as robotic manipulation or navigation, without substantial further development.
  • The proprietary-versus-open-source performance gap suggests that larger or more diverse training data helps affordance perception but is not sufficient to approach human-level ability.
  • The very low performance on transformative affordance implies that models lack robust reasoning about context, culture, time, and individual differences, not just object recognition.
  • A4Bench can serve as a standardized diagnostic for measuring progress in affordance perception across future multimodal models.
  • The finding that all models outperform random guessing but remain far below humans indicates that models have partial, but incomplete, affordance knowledge.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the exact-match scoring on multi-select questions with a hidden number of correct answers may understate partial knowledge; the paper reports moderate precision but low recall, so models may identify some correct affordances while missing others.
  • My inference: the fixed 'correct answer' labels for cross-cultural and individual-variant affordance are inherently contestable, since the paper itself defines these affordances as varying by culture or person; a different panel of annotators might produce different labels and a different human-model gap.
  • My inference: a testable extension would be to ask models to provide explanations or confidence scores alongside their answers, which could separate genuine affordance understanding from pattern matching and could inform training strategies.
  • My inference: because transformative affordance failures are most severe, embodied AI systems should be evaluated on interactive affordance tasks (e.g., actually picking up or using objects) rather than only on image-based question answering.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper introduces A4Bench, a benchmark of 2,000 multimodal question-answer pairs for evaluating affordance perception in multimodal large language models (MLLMs). The benchmark is split into Constitutive Affordance (1,282 pairs across four disciplines and nine sub-disciplines) and Transformative Affordance (718 pairs covering misleading, time-dependent, cross-cultural, and individual-variant affordances). The authors evaluate 17 MLLMs (nine proprietary and eight open-source) plus human participants, reporting exact-match accuracy and several auxiliary metrics. Their headline result is a large human-model gap: the best model, Gemini-2.0-Pro, reaches 18.05% overall exact match accuracy while human best and worst are 85.34% and 81.25%, respectively. The paper concludes that MLLMs are still poor at affordance perception, especially in transformative affordance.

Significance. If the reported results are taken at face value, A4Bench would be a useful diagnostic resource for a relatively under-explored capability of multimodal models, and the paper's breadth of evaluation—nine disciplines, four transformative dimensions, 17 models, multiple metrics, and a human study—is a strength. The benchmark construction is also thoughtful in several respects: it uses multiple choice formats with concealed answer counts, removes explicit object names to force visual grounding, and applies expert review to questions. However, the central claim that MLLMs are generally poor at affordance perception is undermined by the model-in-the-loop adversarial filtering used during benchmark construction, which selects items specifically until current models fail. Because the reported gap is therefore partly by construction, the paper's main conclusion overgeneralizes unless the authors provide unfiltered or control evidence. The absence of variance measures and significance testing further weakens the quantitative comparisons. With additional analysis and reframing, the contribution could be valuable for the MLLM benchmarking community.

major comments (3)
  1. [Constructing the A4Bench, Key Principles] The quality control mechanism described in the Key Principles section is an adversarial, model-in-the-loop selection procedure: items are revised by alternating model and expert input 'until the model fails to respond correctly.' This means the final 2,000 items are chosen specifically to make current MLLMs fail, so the reported 18.05% versus ~85% human gap is not an unbiased estimate of general affordance perception ability. The paper should report results on the pre-filtering or unfiltered question set, name the MLLMs used in the loop, and provide a random-sample control subset, or alternatively reframe the central claim to say that MLLMs perform poorly on an adversarially filtered benchmark rather than that they are poor at affordance perception generally.
  2. [Table 1 and Findings of Transformative Affordance] The metric aggregation is internally inconsistent. Table 1 computes the Transformative Overall as a weighted average of Time, Culture, and Individual dimensions, deliberately excluding Misleading, while the final Overall includes Misleading. The text in Findings of Transformative Affordance explains the exclusion by the high random-guess rate for Misleading, but the headline 18.05% overall and the per-dimension comparisons are affected by this choice. The paper should report results both with and without the Misleading dimension consistently, or use a baseline-subtracted score for Misleading, so that readers can compare subdimensions and the overall score on the same basis.
  3. [Human Performance and Human Expert Annotation] No variance or significance information is reported for either human or model scores, and the human study lacks key details: the number of participants, whether the same participants answered all questions, and how disagreement among the five expert annotators was resolved. This matters most for Cross-Cultural and Individual-Variant Affordance, where the correct answer is defined as observer- or culture-dependent; if the ground-truth labels are a single fixed answer, the paper needs to show inter-annotator agreement and explain whether the label represents majority vote, expert consensus, or something else. Without this, the claimed human-model gap cannot be distinguished from label instability.
minor comments (4)
  1. [Introduction] There are several typographical errors, including 'cross-cluture,' 'discplines,' 'Indenpendent,' 'quzlity,' 'choosen,' 'percception,' and duplicated words such as 'category category'; a careful proofreading pass is needed.
  2. [Figure 4] Figure 4(b) is described as excluding the Misleading dimension, which is consistent with the Transformative Overall definition in Table 1, but the text in Findings of Transformative Affordance and the Table 1 footnote should state this exclusion in the same place so readers are not confused about why Misleading appears in the final Overall but not in the Transformative Overall.
  3. [References] Several references are incomplete or inconsistently formatted, such as 'step-1o' with a non-descriptive URL and 'BailingMM-Pro-0120' as an entry without full author or venue information; these should be standardized before publication.
  4. [Experiment Results, Human Performance] The sentence 'This rigorous methodology ensures reliable and results' appears to be missing a word, and the paragraph would benefit from stating explicitly whether the human participants received the same prompts as the models and whether they were allowed to ask clarifying questions.

Circularity Check

1 steps flagged · score 6.0 of 10

The headline claim 'MLLMs are still poor at affordance perception' is partly by construction: A4Bench's quality-control loop keeps questions only after 'the model fails to respond correctly,' so low MLLM scores on the final set are an inclusion-criterion output rather than an independent estimate.

  1. fitted input called prediction [Constructing the A4Bench, Key Principles, 'Guaranteeing Benchmarking Difficulty']
    "Quality control mechanism employs a human-MLLM mixed-obfuscation adaptation process. This iterative method starts with human-generated problems, followed by alternating revisions between models and experts until the model fails to respond correctly, ensuring a challenging benchmark."

    The rule for keeping a revised question is that an unnamed MLLM fails it. The paper then presents exact-match accuracy on this same filtered set (Gemini-2.0-Pro 18.05% overall) as evidence that 'MLLMs are still poor at affordance perception.' Because the benchmark's difficulty was tuned to produce model failure, low scores are partly a restatement of the selection rule rather than an unbiased measurement of general affordance ability. The paper reports no pre-revision control subset, no random-sample control, and does not identify which MLLMs were in the revision loop, so the headline human-model gap cannot be separated from the filter. Human expert labels are independent, so the result is partially but not fully circular.

full rationale

The benchmark's ground truth is anchored by expert human annotation, so the human-model comparison is not definitionally forced: humans reach 85.34% best accuracy on the same items, showing the tasks are answerable. However, the central conclusion generalizes from a set whose construction criterion is 'the model fails to respond correctly' after alternating model-expert revisions. That makes the low MLLM scores partly an artifact of model-in-the-loop selection rather than a free empirical discovery. No self-citation chain or imported uniqueness theorem is load-bearing; the paper's many author-overlapping related-work citations (e.g., AIBench, The Ever-Evolving Science Exam) are contextual, not the basis of the derivation. The remaining concerns—unresolved label-subjectivity for Cross-Cultural and Individual-Variant affordance, and the exclusion of Misleading from the Transformative Overall score—are validity and metric-consistency issues, not circularity. Overall score 6 reflects one central reduction: the benchmark's difficulty is fitted to failure and the 'poor at affordance perception' claim is reported from that fitted set, giving partial circularity rather than full equivalence.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper makes no mathematical derivation and introduces no physical entities or fitted parameters. Its load-bearing assumptions are about construct validity, label objectivity, and evaluation comparability.

assumptions (4)
  • domain assumption Affordance perception can be validly measured by multi-select QA pairs where object names are replaced with 'this object'.
    The 'Vision-language prompt removal approach' in 'Guaranteeing Benchmarking Difficulty' assumes the image alone carries enough information to answer, isolating affordance perception from textual object recognition.
  • domain assumption Expert annotator consensus provides objective ground truth, including for cross-cultural and individual-variant affordance.
    The 'Human Expert Annotation' paragraph says each question is reviewed by at least five experts, but for affordances that vary by culture or individual, a single fixed label set needs justification that the paper does not provide.
  • domain assumption The quality-control process, which revises questions until MLLMs fail, does not bias the benchmark against the evaluated models.
    In 'Guaranteeing Benchmarking Difficulty', questions are iteratively revised until models fail; this intentionally selects hard items and may inflate the reported human-model gap.
  • domain assumption Zero-shot API evaluations with tailored prompts are comparable across models.
    The 'Benchmark Candidates' section states prompts vary slightly across MLLMs; this assumes prompt differences do not drive score differences.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Affordance Benchmark for MLLMs." pith.science (2026). https://pith.science/paper/DOQPUIIX

@misc{pith2026250600893,
  author       = {Pith},
  title        = {Pith review of: Affordance Benchmark for MLLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DOQPUIIX}},
  note         = {Machine review of arXiv:2506.00893}
}
read the original abstract

Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong performance in vision-language tasks, their ability to perceive affordance, which is crucial for intuitive and safe interactions, remains underexplored. To address this, we introduce **A4Bench**, a novel benchmark designed to evaluate the affordance perception abilities of MLLMs across two dimensions: 1) Constitutive Affordance, assessing understanding of inherent object properties through 1,282 questionanswer pairs spanning nine sub-disciplines, and 2) Transformative Affordance, probing dynamic and contextual nuances (e.g., misleading, time-dependent, cultural, or individual-specific affordance) with 718 challenging question-answer pairs. We evaluate 17 MLLMs (nine proprietary and eight open-source) and compare them to human performance. Results show that proprietary models generally outperform open-source ones, yet all models perform far below humans, especially in transformative affordance. Furthermore, even top-performing models, such as Gemini-2.0-Pro (18.05% overall exact match accuracy), significantly lag behind human performance (best: 85.34%, worst: 81.25%). These findings highlight critical gaps in environmental understanding of MLLMs and provide a foundation for advancing AI systems toward more robust, context-aware interactions.

Figures

Figures reproduced from arXiv: 2506.00893 by the authors.

Figure 1
Figure 1. The motivation of the A4Bench. The affordance theory proposed by James J. Gibson(Gibson 2014) defines the action [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Structure and quantitative overview of A4Bench. The left panel presents the focused aspects of the benchmark, de￾tailing the primary dimensions (Constitutive and Transformative Affordance) and their respective sub-dimensions. The middle panel depicts the distributions of answer counts and option counts. The right panel reports the number of multimodal ques￾tion–answer pairs across each sub-dimension, providing a com… view at source ↗
Figure 3
Figure 3. Typical samples from the A4Bench. Each sample is accompanied by a image-question-answer pair. A4Bench eval￾uates models across diverse discplines (Constitutive Affordance) and challenging dimensions (Transformative Affordance), ensuring a comprehensive evaluation of the affordance perception capabilities. Constitutive Affordance Perception We propose A4Bench, a robust benchmark designed to assess affor￾dance percept… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: A Quick Look of the A4Bench Outcomes. (a) showcases a comparative analysis of the overall match score between [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 31 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al

    Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  4. [4]

    B.; Bout, B.; Chaplot, D.; Chudnovsky, J.; Costa, D.; De Monicault, B.; Garg, S.; Gervet, T.; et al

    Agrawal, P.; Antoniak, S.; Hanna, E. B.; Bout, B.; Chaplot, D.; Chudnovsky, J.; Costa, D.; De Monicault, B.; Garg, S.; Gervet, T.; et al. 2024. Pixtral 12B. arXiv preprint arXiv:2410.07073

  5. [5]

    Anthropic . 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku. https://www.anthropic.com/news/claude-3-family

  6. [6]

    Anthropic . 2025. Claude 3.7 Sonnet. https://www.anthropic.com/news/claude-3-7-sonnet

  7. [7]

    Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; et al. 2025. Qwen2.5-VL Technical Report. arXiv:2502.13923

  8. [8]

    BailingMM-Pro-0120. 2025. Https://github.com/wwbin2017/bailing/

Show all 62 references
  1. [9]

    Blake, R.; and Shiffrar, M. 2007. Perception of human motion. Annu. Rev. Psychol., 58(1): 47--73

  2. [10]

    Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; and Zhao, F. 2024 a . Are We on the Right Way for Evaluating Large Vision-Language Models? In Globerson, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J.; and Zhang,...

  3. [11]

    Chen, L.; Zhang, Y.; Ren, S.; Zhao, H.; Cai, Z.; Wang, Y.; Wang, P.; Meng, X.; Liu, T.; and Chang, B. 2024 b . PCA -Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Assoc...

  4. [12]

    Chen, X.; Wu, Z.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; and Ruan, C. 2025. Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. arXiv preprint arXiv:2501.17811

  5. [13]

    Dai, W.; Lee, N.; Wang, B.; Yang, Z.; Liu, Z.; Barker, J.; Rintamaki, T.; Shoeybi, M.; Catanzaro, B.; and Ping, W. 2024. Nvlm: Open frontier-class multimodal llms. arXiv preprint arXiv:2409.11402

  6. [14]

    Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023. InstructBLIP : Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv preprint arXiv:2305.06500. Originally announced May 2023

  7. [15]

    Davis, J.; and Goadrich, M. 2006. The relationship between Precision-Recall and ROC curves. ICML '06. New York, NY, USA: Association for Computing Machinery. ISBN 1595933832

  8. [16]

    Dong, H.; Kang, Z.; Yin, W.; Liang, X.; Feng, C.; and Ran, J. 2025. Scalable vision language model training via high quality data curation. arXiv preprint arXiv:2501.05952

  9. [17]

    I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al

    Georgiev, P.; Lei, V. I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al. 2024. Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context. arXiv preprint arXiv:2403.05530

  10. [18]

    Gibson, J. J. 2014. The ecological approach to visual perception: classic edition. Psychology press

  11. [19]

    Google . 2024 a . Gemini . https://gemini.google.com. Large language model

  12. [20]

    Google . 2024 b . Introducing Gemini 2.0: Our New AI Model for the Agentic Era. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/

  13. [21]

    Guo, D.; Wu, F.; Zhu, F.; Leng, F.; Shi, G.; Chen, H.; Fan, H.; Wang, J.; Jiang, J.; Wang, J.; et al. 2025 a . Seed1. 5-vl technical report. arXiv preprint arXiv:2505.07062

  14. [22]

    Guo, Y.; Ji, K.; Zhu, X.; Wang, J.; Wen, F.; Li, C.; Zhang, Z.; and Zhai, G. 2025 b . Human-Centric Evaluation for Foundation Models. arXiv:2506.01793

  15. [23]

    D.; Bouadjenek, M

    Huynh, N. D.; Bouadjenek, M. R.; Aryal, S.; Razzak, I.; and Hacid, H. 2025. Visual Question Answering: From Early Developments to Recent Advances -- A Survey. arXiv preprint arXiv:2501.03939. Cs.CV, cs.MM

  16. [24]

    Jiao, Z.; and Li, X. 2025. An End-to-End Deep Graph Clustering via Online Mutual Learning. IEEE Transactions on Neural Networks and Learning Systems, 36(2): 3847--3854

  17. [25]

    Jiao, Z.; Zhang, H.; and Li, X. 2025 a . Cnn2gnn: How to bridge cnn with gnn. IEEE Transactions on Pattern Analysis and Machine Intelligence

  18. [26]

    Jiao, Z.; Zhang, H.; and Li, X. 2025 b . Deep Graph Multi-View Representation Learning With Self-Augmented View Fusion. IEEE Transactions on Neural Networks and Learning Systems, 1--12

  19. [27]

    Li, B.; Ge, Y.; Ge, Y.; Wang, G.; Wang, R.; Zhang, R.; and Shan, Y. 2024 a . SEED-Bench: Benchmarking Multimodal Large Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13299--13308

  20. [28]

    Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; et al. 2024 b . Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326

  21. [29]

    Lin, J.; Yin, H.; Ping, W.; Lu, Y.; Molchanov, P.; Tao, A.; Mao, H.; Kautz, J.; Shoeybi, M.; and Han, S. 2023. VILA: On Pre-training for Visual Language Models. arXiv:2312.07533

  22. [30]

    Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2024 a . Improved Baselines with Visual Instruction Tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 26296--26306. IEEE. CVPR 2024 Highlight

  23. [31]

    Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual Instruction Tuning. In Advances in Neural Information Processing Systems, volume 36, 26924--26958. Curran Associates, Inc. NeurIPS 2023 Oral Presentation

  24. [32]

    Liu, X.; Zhu, Y.; Gu, J.; Lan, Y.; Yang, C.; and Qiao, Y. 2024 b . Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In European Conference on Computer Vision, 386--403. Springer

  25. [33]

    Liu, Y.; Duan, H.; Zhang, Y.; Li, B.; Zhang, S.; Zhao, W.; Yuan, Y.; Wang, J.; He, C.; Liu, Z.; et al. 2024 c . Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, 216--233. Springer

  26. [34]

    Liu, Y.; Zhao, Z.; Zhuang, Z.; Tian, L.; Zhou, X.; and Zhou, J. 2024 d . Points: Improving your vision-language model with affordable strategies. arXiv preprint arXiv:2409.04828

  27. [35]

    Lu, S.; Li, Y.; Chen, Q.-G.; Xu, Z.; Luo, W.; Zhang, K.; and Ye, H.-J. 2024. Ovis: Structural embedding alignment for multimodal large language model. arXiv preprint arXiv:2405.20797

  28. [36]

    MUG-U-7B. 2025. MUG-U. Https://github.com/Shopee-MUG/MUG-U/

  29. [37]

    OpenAI . 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/

  30. [38]

    OpenAI . 2025 a . Introducing OpenAI o3 and o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/

  31. [39]

    OpenAI . 2025 b . OpenAI Models Documentation. https://platform.openai.com/docs/models

  32. [40]

    Phan, L.; Gatti, A.; Han, Z.; Li, N.; Hu, J.; Zhang, H.; et al. 2025. Humanity's Last Exam. arXiv:2501.14249

  33. [41]

    H.; Wu, J.; Washington, C.; Sadler, B

    Song, C. H.; Wu, J.; Washington, C.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2023. LLM-Planner : Few-Shot Grounded Planning for Embodied Agents with Large Language Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

  34. [42]

    step-1o. 2024. Http://www.eecs.harvard.edu/mdw/ proj/codeblue/

  35. [43]

    Precision

    Streiner, D. L.; and Norman, G. R. 2006. “Precision” and “Accuracy”: Two Terms That Are Neither. Journal of Clinical Epidemiology, 59(4): 327--330

  36. [44]

    Team, K.; Du, A.; Yin, B.; Xing, B.; Qu, B.; Wang, B.; Chen, C.; Zhang, C.; Du, C.; Wei, C.; et al. 2025. Kimi-vl technical report. arXiv preprint arXiv:2504.07491

  37. [45]

    Wang, J.; Zhang, H.; Wang, H.; and Yuan, Y. 2025 a . Graph Convolutional Network With Self-Augmented Weights for Semi-Supervised Multi-View Learning. IEEE Transactions on Neural Networks and Learning Systems, 36(7): 12257--12270

  38. [46]

    Wang, J.; Zhang, H.; and Yuan, Y. 2025. Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks. In Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 21001--21010

  39. [47]

    Wang, J.; Zhang, Z.; Guo, Y.; Wen, F.; Shen, Y.; Liang, Y.; Wu, Y.; Li, W.; Li, C.; Chen, Z.; Jia, Q.; and Zhai, G. 2025 b . The Ever-Evolving Science Exam. arXiv:2507.16514

  40. [48]

    Wen, F.; Guo, Y.; Wang, J.; Xiao, J.; Zhou, Y.; Li, C.; Zhang, Z.; and Zhai, G. 2025. Improve MLLM Benchmark Efficiency through Interview. arXiv:2506.00883

  41. [49]

    Wu, C.; Chen, X.; Wu, Z.; Ma, Y.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; Ruan, C.; et al. 2024 a . Janus: Decoupling visual encoding for unified multimodal understanding and generation. arXiv preprint arXiv:2410.13848

  42. [50]

    Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Li, C.; Sun, W.; Yan, Q.; Zhai, G.; and Lin, W. 2024 b . Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision. arXiv:2309.14181

  43. [51]

    Wu, H.; Zhang, Z.; Zhang, W.; Chen, C.; Liao, L.; Li, C.; Gao, Y.; Wang, A.; Zhang, E.; Sun, W.; Yan, Q.; Min, X.; Zhai, G.; and Lin, W. 2023. Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels. arXiv:2312.17090

  44. [52]

    Wu, Z.; Chen, X.; Pan, Z.; Liu, X.; Liu, W.; Dai, D.; Gao, H.; Ma, Y.; et al. 2024 c . DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding. arXiv:2412.10302

  45. [53]

    Yao, Y.; Yu, T.; Zhang, A.; Wang, C.; Cui, J.; Zhu, H.; Cai, T.; Li, H.; Zhao, W.; He, Z.; et al. 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone. arXiv preprint arXiv:2408.01800

  46. [54]

    Ye, J.; Xu, H.; Liu, H.; Hu, A.; Yan, M.; Qian, Q.; Zhang, J.; Huang, F.; and Zhou, J. 2024. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models. arXiv:2408.04840

  47. [55]

    Ying, K.; Meng, F.; Wang, J.; Li, Z.; Lin, H.; Yang, Y.; Zhang, H.; et al. 2024. MMT-bench: a comprehensive multimodal benchmark for evaluating large vision-language models towards multitask AGI. In Proceedings of the 41st International Conference on Machine Learning, ICML'24....

  48. [56]

    Zhang, B.; Li, S.; Tian, R.; Yang, Y.; Tang, J.; Zhou, J.; and Ma, L. 2025 a . Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput. arXiv preprint arXiv:2505.09498

  49. [57]

    Zhang, Z.; Su, S.; Zhu, Y.; Yan, Q.; Sun, J.; and Zhang, Y. 2024. A-Bench: Are LMMs Masters at Evaluating AI-Generated Images? arXiv preprint arXiv:2406.03070

  50. [58]

    Zhang, Z.; Sun, W.; Min, X.; Wang, T.; Lu, W.; and Zhai, G. 2022. No-Reference Quality Assessment for 3D Colored Point Cloud and Mesh Models. IEEE Transactions on Circuits and Systems for Video Technology, 32(11): 7618--7631

  51. [59]

    Zhang, Z.; Wang, J.; Guo, Y.; Wen, F.; Chen, Z.; Wang, H.; Li, W.; Sun, L.; Zhou, Y.; Zhang, J.; Yan, B.; Jia, Z.; Xiao, J.; Tian, Y.; Zhu, X.; Zhang, K.; Li, C.; Liu, X.; Min, X.; Jia, Q.; and Zhai, G. 2025 b . AIBench: Towards Trustworthy Evaluation Under The 45° Law. https:...

  52. [60]

    Zhang, Z.; Zhao, X.; Fang, X.; Li, C.; Liu, X.; Min, X.; Duan, H.; Chen, K.; and Zhai, G. 2025 c . Redundancy Principles for MLLMs Benchmarks. arXiv preprint arXiv:2501.13953

  53. [61]

    Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023. MiniGPT-4 : Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv preprint arXiv:2304.10592. Originally announced April 2023

  54. [62]

    Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; et al. 2025. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. arXiv:2504.10479

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.