REVIEW 3 major objections 4 minor 62 references
Affordance Benchmark for MLLMs
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that even the best multimodal large language models perceive object affordances far below human level, with the top model scoring 18.05% exact match against 81.25–85.34% for humans on its 2,000-question A4Bench.
desk verdict A worthwhile affordance benchmark, but the headline 18-vs-85 gap is partly an artifact of adversarial filtering; still deserves review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is A4Bench, a benchmark constructed from 2,000 image-question-answer pairs organized along two dimensions: Constitutive Affordance (1,282 pairs across daily life, natural and rural, urban and industrial, and medical and scientific environments) and Transformative Affordance (718 pairs covering misleading, time-dependent, cross-cultural, and individual-variant affordances). The benchmark uses a multi-select question format where the number of correct answers is hidden, explicit object names are replaced with 'this object' in some items, and answers are randomized and evaluated by exact match. The key mechanism is the pairing of diverse, context-rich images with questions that require reasoning about action possibilities, combined with a human-MLLM iterative revision process intended to ensure that the questions are genuinely difficult for models.
What would settle it
Have a new, independent panel of human annotators re-label a random sample of the transformative-affordance questions, especially the cross-cultural and individual-variant subsets, without seeing the original answers. If inter-annotator agreement on those subsets is low or near chance, then the fixed ground-truth labels are not stable across people, and the reported human-model performance gap would not be a well-defined measurement.
Extended reading notes
Core claim
The central discovery is that current multimodal large language models, despite strong performance on many vision-language tasks, perceive affordance at a level dramatically below human performance. On A4Bench, the top-performing model (Gemini-2.0-Pro) achieves only 18.05% exact match accuracy, while human best and worst scores are 85.34% and 81.25%. Proprietary models generally outperform open-source models, yet all models fall far short of humans, particularly on transformative affordance categories such as misleading affordance, time-dependent affordance, cross-cultural affordance, and individual-variant affordance. The paper argues that this gap reveals a fundamental deficiency in models' understanding of object-human-environment interactions, and it positions A4Bench as a diagnostic tool to measure and guide progress toward more context-aware, safe multimodal AI.
Load-bearing premise
The benchmark's labeled correct answers are treated as objective and unambiguous even for affordances that the paper itself says vary by culture or by individual, and the paper does not explain how annotator disagreements were resolved.
Editorial extensions
If this is right
- If the benchmark is accurate, current MLLMs cannot be relied upon for tasks that require understanding what objects afford, such as robotic manipulation or navigation, without substantial further development.
- The proprietary-versus-open-source performance gap suggests that larger or more diverse training data helps affordance perception but is not sufficient to approach human-level ability.
- The very low performance on transformative affordance implies that models lack robust reasoning about context, culture, time, and individual differences, not just object recognition.
- A4Bench can serve as a standardized diagnostic for measuring progress in affordance perception across future multimodal models.
- The finding that all models outperform random guessing but remain far below humans indicates that models have partial, but incomplete, affordance knowledge.
Reading between the lines
- My inference: the exact-match scoring on multi-select questions with a hidden number of correct answers may understate partial knowledge; the paper reports moderate precision but low recall, so models may identify some correct affordances while missing others.
- My inference: the fixed 'correct answer' labels for cross-cultural and individual-variant affordance are inherently contestable, since the paper itself defines these affordances as varying by culture or person; a different panel of annotators might produce different labels and a different human-model gap.
- My inference: a testable extension would be to ask models to provide explanations or confidence scores alongside their answers, which could separate genuine affordance understanding from pattern matching and could inform training strategies.
- My inference: because transformative affordance failures are most severe, embodied AI systems should be evaluated on interactive affordance tasks (e.g., actually picking up or using objects) rather than only on image-based question answering.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces A4Bench, a benchmark of 2,000 multimodal question-answer pairs for evaluating affordance perception in multimodal large language models (MLLMs). The benchmark is split into Constitutive Affordance (1,282 pairs across four disciplines and nine sub-disciplines) and Transformative Affordance (718 pairs covering misleading, time-dependent, cross-cultural, and individual-variant affordances). The authors evaluate 17 MLLMs (nine proprietary and eight open-source) plus human participants, reporting exact-match accuracy and several auxiliary metrics. Their headline result is a large human-model gap: the best model, Gemini-2.0-Pro, reaches 18.05% overall exact match accuracy while human best and worst are 85.34% and 81.25%, respectively. The paper concludes that MLLMs are still poor at affordance perception, especially in transformative affordance.
Significance. If the reported results are taken at face value, A4Bench would be a useful diagnostic resource for a relatively under-explored capability of multimodal models, and the paper's breadth of evaluation—nine disciplines, four transformative dimensions, 17 models, multiple metrics, and a human study—is a strength. The benchmark construction is also thoughtful in several respects: it uses multiple choice formats with concealed answer counts, removes explicit object names to force visual grounding, and applies expert review to questions. However, the central claim that MLLMs are generally poor at affordance perception is undermined by the model-in-the-loop adversarial filtering used during benchmark construction, which selects items specifically until current models fail. Because the reported gap is therefore partly by construction, the paper's main conclusion overgeneralizes unless the authors provide unfiltered or control evidence. The absence of variance measures and significance testing further weakens the quantitative comparisons. With additional analysis and reframing, the contribution could be valuable for the MLLM benchmarking community.
major comments (3)
- [Constructing the A4Bench, Key Principles] The quality control mechanism described in the Key Principles section is an adversarial, model-in-the-loop selection procedure: items are revised by alternating model and expert input 'until the model fails to respond correctly.' This means the final 2,000 items are chosen specifically to make current MLLMs fail, so the reported 18.05% versus ~85% human gap is not an unbiased estimate of general affordance perception ability. The paper should report results on the pre-filtering or unfiltered question set, name the MLLMs used in the loop, and provide a random-sample control subset, or alternatively reframe the central claim to say that MLLMs perform poorly on an adversarially filtered benchmark rather than that they are poor at affordance perception generally.
- [Table 1 and Findings of Transformative Affordance] The metric aggregation is internally inconsistent. Table 1 computes the Transformative Overall as a weighted average of Time, Culture, and Individual dimensions, deliberately excluding Misleading, while the final Overall includes Misleading. The text in Findings of Transformative Affordance explains the exclusion by the high random-guess rate for Misleading, but the headline 18.05% overall and the per-dimension comparisons are affected by this choice. The paper should report results both with and without the Misleading dimension consistently, or use a baseline-subtracted score for Misleading, so that readers can compare subdimensions and the overall score on the same basis.
- [Human Performance and Human Expert Annotation] No variance or significance information is reported for either human or model scores, and the human study lacks key details: the number of participants, whether the same participants answered all questions, and how disagreement among the five expert annotators was resolved. This matters most for Cross-Cultural and Individual-Variant Affordance, where the correct answer is defined as observer- or culture-dependent; if the ground-truth labels are a single fixed answer, the paper needs to show inter-annotator agreement and explain whether the label represents majority vote, expert consensus, or something else. Without this, the claimed human-model gap cannot be distinguished from label instability.
minor comments (4)
- [Introduction] There are several typographical errors, including 'cross-cluture,' 'discplines,' 'Indenpendent,' 'quzlity,' 'choosen,' 'percception,' and duplicated words such as 'category category'; a careful proofreading pass is needed.
- [Figure 4] Figure 4(b) is described as excluding the Misleading dimension, which is consistent with the Transformative Overall definition in Table 1, but the text in Findings of Transformative Affordance and the Table 1 footnote should state this exclusion in the same place so readers are not confused about why Misleading appears in the final Overall but not in the Transformative Overall.
- [References] Several references are incomplete or inconsistently formatted, such as 'step-1o' with a non-descriptive URL and 'BailingMM-Pro-0120' as an entry without full author or venue information; these should be standardized before publication.
- [Experiment Results, Human Performance] The sentence 'This rigorous methodology ensures reliable and results' appears to be missing a word, and the paragraph would benefit from stating explicitly whether the human participants received the same prompts as the models and whether they were allowed to ask clarifying questions.
Circularity Check
The headline claim 'MLLMs are still poor at affordance perception' is partly by construction: A4Bench's quality-control loop keeps questions only after 'the model fails to respond correctly,' so low MLLM scores on the final set are an inclusion-criterion output rather than an independent estimate.
-
fitted input called prediction
[Constructing the A4Bench, Key Principles, 'Guaranteeing Benchmarking Difficulty']
"Quality control mechanism employs a human-MLLM mixed-obfuscation adaptation process. This iterative method starts with human-generated problems, followed by alternating revisions between models and experts until the model fails to respond correctly, ensuring a challenging benchmark."
The rule for keeping a revised question is that an unnamed MLLM fails it. The paper then presents exact-match accuracy on this same filtered set (Gemini-2.0-Pro 18.05% overall) as evidence that 'MLLMs are still poor at affordance perception.' Because the benchmark's difficulty was tuned to produce model failure, low scores are partly a restatement of the selection rule rather than an unbiased measurement of general affordance ability. The paper reports no pre-revision control subset, no random-sample control, and does not identify which MLLMs were in the revision loop, so the headline human-model gap cannot be separated from the filter. Human expert labels are independent, so the result is partially but not fully circular.
full rationale
The benchmark's ground truth is anchored by expert human annotation, so the human-model comparison is not definitionally forced: humans reach 85.34% best accuracy on the same items, showing the tasks are answerable. However, the central conclusion generalizes from a set whose construction criterion is 'the model fails to respond correctly' after alternating model-expert revisions. That makes the low MLLM scores partly an artifact of model-in-the-loop selection rather than a free empirical discovery. No self-citation chain or imported uniqueness theorem is load-bearing; the paper's many author-overlapping related-work citations (e.g., AIBench, The Ever-Evolving Science Exam) are contextual, not the basis of the derivation. The remaining concerns—unresolved label-subjectivity for Cross-Cultural and Individual-Variant affordance, and the exclusion of Misleading from the Transformative Overall score—are validity and metric-consistency issues, not circularity. Overall score 6 reflects one central reduction: the benchmark's difficulty is fitted to failure and the 'poor at affordance perception' claim is reported from that fitted set, giving partial circularity rather than full equivalence.
Assumptions & free parameters
assumptions (4)
- domain assumption Affordance perception can be validly measured by multi-select QA pairs where object names are replaced with 'this object'.
- domain assumption Expert annotator consensus provides objective ground truth, including for cross-cultural and individual-variant affordance.
- domain assumption The quality-control process, which revises questions until MLLMs fail, does not bias the benchmark against the evaluated models.
- domain assumption Zero-shot API evaluations with tailored prompts are comparable across models.
Cite this review
Pith. "Pith review of Affordance Benchmark for MLLMs." pith.science (2026). https://pith.science/paper/DOQPUIIX
@misc{pith2026250600893,
author = {Pith},
title = {Pith review of: Affordance Benchmark for MLLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/DOQPUIIX}},
note = {Machine review of arXiv:2506.00893}
}
read the original abstract
Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong performance in vision-language tasks, their ability to perceive affordance, which is crucial for intuitive and safe interactions, remains underexplored. To address this, we introduce **A4Bench**, a novel benchmark designed to evaluate the affordance perception abilities of MLLMs across two dimensions: 1) Constitutive Affordance, assessing understanding of inherent object properties through 1,282 questionanswer pairs spanning nine sub-disciplines, and 2) Transformative Affordance, probing dynamic and contextual nuances (e.g., misleading, time-dependent, cultural, or individual-specific affordance) with 718 challenging question-answer pairs. We evaluate 17 MLLMs (nine proprietary and eight open-source) and compare them to human performance. Results show that proprietary models generally outperform open-source ones, yet all models perform far below humans, especially in transformative affordance. Furthermore, even top-performing models, such as Gemini-2.0-Pro (18.05% overall exact match accuracy), significantly lag behind human performance (best: 85.34%, worst: 81.25%). These findings highlight critical gaps in environmental understanding of MLLMs and provide a foundation for advancing AI systems toward more robust, context-aware interactions.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[4]
B.; Bout, B.; Chaplot, D.; Chudnovsky, J.; Costa, D.; De Monicault, B.; Garg, S.; Gervet, T.; et al
Agrawal, P.; Antoniak, S.; Hanna, E. B.; Bout, B.; Chaplot, D.; Chudnovsky, J.; Costa, D.; De Monicault, B.; Garg, S.; Gervet, T.; et al. 2024. Pixtral 12B. arXiv preprint arXiv:2410.07073
arXiv 2024
-
[5]
Anthropic . 2024. The Claude 3 Model Family: Opus, Sonnet, Haiku. https://www.anthropic.com/news/claude-3-family
work page 2024
-
[6]
Anthropic . 2025. Claude 3.7 Sonnet. https://www.anthropic.com/news/claude-3-7-sonnet
work page 2025
-
[7]
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; et al. 2025. Qwen2.5-VL Technical Report. arXiv:2502.13923
arXiv 2025
-
[8]
BailingMM-Pro-0120. 2025. Https://github.com/wwbin2017/bailing/
work page 2025
Show all 62 references
-
[9]
Blake, R.; and Shiffrar, M. 2007. Perception of human motion. Annu. Rev. Psychol., 58(1): 47--73
2007
-
[10]
Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; and Zhao, F. 2024 a . Are We on the Right Way for Evaluating Large Vision-Language Models? In Globerson, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J.; and Zhang,...
2024
-
[11]
Chen, L.; Zhang, Y.; Ren, S.; Zhao, H.; Cai, Z.; Wang, Y.; Wang, P.; Meng, X.; Liu, T.; and Chang, B. 2024 b . PCA -Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Assoc...
2024
-
[12]
Chen, X.; Wu, Z.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; and Ruan, C. 2025. Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling. arXiv preprint arXiv:2501.17811
2025 arXiv
-
[13]
Dai, W.; Lee, N.; Wang, B.; Yang, Z.; Liu, Z.; Barker, J.; Rintamaki, T.; Shoeybi, M.; Catanzaro, B.; and Ping, W. 2024. Nvlm: Open frontier-class multimodal llms. arXiv preprint arXiv:2409.11402
2024 arXiv
-
[14]
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023. InstructBLIP : Towards General-purpose Vision-Language Models with Instruction Tuning. arXiv preprint arXiv:2305.06500. Originally announced May 2023
2023 arXiv
-
[15]
Davis, J.; and Goadrich, M. 2006. The relationship between Precision-Recall and ROC curves. ICML '06. New York, NY, USA: Association for Computing Machinery. ISBN 1595933832
2006
-
[16]
Dong, H.; Kang, Z.; Yin, W.; Liang, X.; Feng, C.; and Ran, J. 2025. Scalable vision language model training via high quality data curation. arXiv preprint arXiv:2501.05952
2025 arXiv
-
[17]
I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al
Georgiev, P.; Lei, V. I.; Burnell, R.; Bai, L.; Gulati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; et al. 2024. Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context. arXiv preprint arXiv:2403.05530
2024 arXiv
-
[18]
Gibson, J. J. 2014. The ecological approach to visual perception: classic edition. Psychology press
2014
-
[19]
Google . 2024 a . Gemini . https://gemini.google.com. Large language model
2024
-
[20]
Google . 2024 b . Introducing Gemini 2.0: Our New AI Model for the Agentic Era. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/
2024
-
[21]
Guo, D.; Wu, F.; Zhu, F.; Leng, F.; Shi, G.; Chen, H.; Fan, H.; Wang, J.; Jiang, J.; Wang, J.; et al. 2025 a . Seed1. 5-vl technical report. arXiv preprint arXiv:2505.07062
2025 arXiv
-
[22]
Guo, Y.; Ji, K.; Zhu, X.; Wang, J.; Wen, F.; Li, C.; Zhang, Z.; and Zhai, G. 2025 b . Human-Centric Evaluation for Foundation Models. arXiv:2506.01793
2025 arXiv
-
[23]
D.; Bouadjenek, M
Huynh, N. D.; Bouadjenek, M. R.; Aryal, S.; Razzak, I.; and Hacid, H. 2025. Visual Question Answering: From Early Developments to Recent Advances -- A Survey. arXiv preprint arXiv:2501.03939. Cs.CV, cs.MM
2025 arXiv
-
[24]
Jiao, Z.; and Li, X. 2025. An End-to-End Deep Graph Clustering via Online Mutual Learning. IEEE Transactions on Neural Networks and Learning Systems, 36(2): 3847--3854
2025
-
[25]
Jiao, Z.; Zhang, H.; and Li, X. 2025 a . Cnn2gnn: How to bridge cnn with gnn. IEEE Transactions on Pattern Analysis and Machine Intelligence
2025
-
[26]
Jiao, Z.; Zhang, H.; and Li, X. 2025 b . Deep Graph Multi-View Representation Learning With Self-Augmented View Fusion. IEEE Transactions on Neural Networks and Learning Systems, 1--12
2025
-
[27]
Li, B.; Ge, Y.; Ge, Y.; Wang, G.; Wang, R.; Zhang, R.; and Shan, Y. 2024 a . SEED-Bench: Benchmarking Multimodal Large Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13299--13308
2024
-
[28]
Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; et al. 2024 b . Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326
2024 arXiv
-
[29]
Lin, J.; Yin, H.; Ping, W.; Lu, Y.; Molchanov, P.; Tao, A.; Mao, H.; Kautz, J.; Shoeybi, M.; and Han, S. 2023. VILA: On Pre-training for Visual Language Models. arXiv:2312.07533
2023 arXiv
-
[30]
Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2024 a . Improved Baselines with Visual Instruction Tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 26296--26306. IEEE. CVPR 2024 Highlight
2024
-
[31]
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual Instruction Tuning. In Advances in Neural Information Processing Systems, volume 36, 26924--26958. Curran Associates, Inc. NeurIPS 2023 Oral Presentation
2023
-
[32]
Liu, X.; Zhu, Y.; Gu, J.; Lan, Y.; Yang, C.; and Qiao, Y. 2024 b . Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In European Conference on Computer Vision, 386--403. Springer
2024
-
[33]
Liu, Y.; Duan, H.; Zhang, Y.; Li, B.; Zhang, S.; Zhao, W.; Yuan, Y.; Wang, J.; He, C.; Liu, Z.; et al. 2024 c . Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, 216--233. Springer
2024
-
[34]
Liu, Y.; Zhao, Z.; Zhuang, Z.; Tian, L.; Zhou, X.; and Zhou, J. 2024 d . Points: Improving your vision-language model with affordable strategies. arXiv preprint arXiv:2409.04828
2024 arXiv
-
[35]
Lu, S.; Li, Y.; Chen, Q.-G.; Xu, Z.; Luo, W.; Zhang, K.; and Ye, H.-J. 2024. Ovis: Structural embedding alignment for multimodal large language model. arXiv preprint arXiv:2405.20797
2024 arXiv
-
[36]
MUG-U-7B. 2025. MUG-U. Https://github.com/Shopee-MUG/MUG-U/
2025
-
[37]
OpenAI . 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/
2024
-
[38]
OpenAI . 2025 a . Introducing OpenAI o3 and o4-mini. https://openai.com/index/introducing-o3-and-o4-mini/
2025
-
[39]
OpenAI . 2025 b . OpenAI Models Documentation. https://platform.openai.com/docs/models
2025
-
[40]
Phan, L.; Gatti, A.; Han, Z.; Li, N.; Hu, J.; Zhang, H.; et al. 2025. Humanity's Last Exam. arXiv:2501.14249
2025 arXiv
-
[41]
H.; Wu, J.; Washington, C.; Sadler, B
Song, C. H.; Wu, J.; Washington, C.; Sadler, B. M.; Chao, W.-L.; and Su, Y. 2023. LLM-Planner : Few-Shot Grounded Planning for Embodied Agents with Large Language Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2023
-
[42]
step-1o. 2024. Http://www.eecs.harvard.edu/mdw/ proj/codeblue/
2024
-
[43]
Precision
Streiner, D. L.; and Norman, G. R. 2006. “Precision” and “Accuracy”: Two Terms That Are Neither. Journal of Clinical Epidemiology, 59(4): 327--330
2006
-
[44]
Team, K.; Du, A.; Yin, B.; Xing, B.; Qu, B.; Wang, B.; Chen, C.; Zhang, C.; Du, C.; Wei, C.; et al. 2025. Kimi-vl technical report. arXiv preprint arXiv:2504.07491
2025 arXiv
-
[45]
Wang, J.; Zhang, H.; Wang, H.; and Yuan, Y. 2025 a . Graph Convolutional Network With Self-Augmented Weights for Semi-Supervised Multi-View Learning. IEEE Transactions on Neural Networks and Learning Systems, 36(7): 12257--12270
2025
-
[46]
Wang, J.; Zhang, H.; and Yuan, Y. 2025. Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks. In Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 21001--21010
2025
-
[47]
Wang, J.; Zhang, Z.; Guo, Y.; Wen, F.; Shen, Y.; Liang, Y.; Wu, Y.; Li, W.; Li, C.; Chen, Z.; Jia, Q.; and Zhai, G. 2025 b . The Ever-Evolving Science Exam. arXiv:2507.16514
2025
-
[48]
Wen, F.; Guo, Y.; Wang, J.; Xiao, J.; Zhou, Y.; Li, C.; Zhang, Z.; and Zhai, G. 2025. Improve MLLM Benchmark Efficiency through Interview. arXiv:2506.00883
2025
-
[49]
Wu, C.; Chen, X.; Wu, Z.; Ma, Y.; Liu, X.; Pan, Z.; Liu, W.; Xie, Z.; Yu, X.; Ruan, C.; et al. 2024 a . Janus: Decoupling visual encoding for unified multimodal understanding and generation. arXiv preprint arXiv:2410.13848
2024 arXiv
-
[50]
Wu, H.; Zhang, Z.; Zhang, E.; Chen, C.; Liao, L.; Wang, A.; Li, C.; Sun, W.; Yan, Q.; Zhai, G.; and Lin, W. 2024 b . Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision. arXiv:2309.14181
2024 arXiv
-
[51]
Wu, H.; Zhang, Z.; Zhang, W.; Chen, C.; Liao, L.; Li, C.; Gao, Y.; Wang, A.; Zhang, E.; Sun, W.; Yan, Q.; Min, X.; Zhai, G.; and Lin, W. 2023. Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels. arXiv:2312.17090
2023 arXiv
-
[52]
Wu, Z.; Chen, X.; Pan, Z.; Liu, X.; Liu, W.; Dai, D.; Gao, H.; Ma, Y.; et al. 2024 c . DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding. arXiv:2412.10302
2024 arXiv
-
[53]
Yao, Y.; Yu, T.; Zhang, A.; Wang, C.; Cui, J.; Zhu, H.; Cai, T.; Li, H.; Zhao, W.; He, Z.; et al. 2024. MiniCPM-V: A GPT-4V Level MLLM on Your Phone. arXiv preprint arXiv:2408.01800
2024 arXiv
-
[54]
Ye, J.; Xu, H.; Liu, H.; Hu, A.; Yan, M.; Qian, Q.; Zhang, J.; Huang, F.; and Zhou, J. 2024. mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models. arXiv:2408.04840
2024 arXiv
-
[55]
Ying, K.; Meng, F.; Wang, J.; Li, Z.; Lin, H.; Yang, Y.; Zhang, H.; et al. 2024. MMT-bench: a comprehensive multimodal benchmark for evaluating large vision-language models towards multitask AGI. In Proceedings of the 41st International Conference on Machine Learning, ICML'24....
2024
-
[56]
Zhang, B.; Li, S.; Tian, R.; Yang, Y.; Tang, J.; Zhou, J.; and Ma, L. 2025 a . Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput. arXiv preprint arXiv:2505.09498
2025 arXiv
-
[57]
Zhang, Z.; Su, S.; Zhu, Y.; Yan, Q.; Sun, J.; and Zhang, Y. 2024. A-Bench: Are LMMs Masters at Evaluating AI-Generated Images? arXiv preprint arXiv:2406.03070
2024 arXiv
-
[58]
Zhang, Z.; Sun, W.; Min, X.; Wang, T.; Lu, W.; and Zhai, G. 2022. No-Reference Quality Assessment for 3D Colored Point Cloud and Mesh Models. IEEE Transactions on Circuits and Systems for Video Technology, 32(11): 7618--7631
2022
-
[59]
Zhang, Z.; Wang, J.; Guo, Y.; Wen, F.; Chen, Z.; Wang, H.; Li, W.; Sun, L.; Zhou, Y.; Zhang, J.; Yan, B.; Jia, Z.; Xiao, J.; Tian, Y.; Zhu, X.; Zhang, K.; Li, C.; Liu, X.; Min, X.; Jia, Q.; and Zhai, G. 2025 b . AIBench: Towards Trustworthy Evaluation Under The 45° Law. https:...
2025
-
[60]
Zhang, Z.; Zhao, X.; Fang, X.; Li, C.; Liu, X.; Min, X.; Duan, H.; Chen, K.; and Zhai, G. 2025 c . Redundancy Principles for MLLMs Benchmarks. arXiv preprint arXiv:2501.13953
2025 arXiv
-
[61]
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023. MiniGPT-4 : Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv preprint arXiv:2304.10592. Originally announced April 2023
2023 arXiv
-
[62]
Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; et al. 2025. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models. arXiv:2504.10479
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.