Pith. sign in

REVIEW 4 major objections 6 minor 57 references

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A benchmark built from research-paper figures shows that even the strongest proprietary multimodal model reaches about half of human expert accuracy on architecture and civil engineering reasoning.

desk verdict A construction pipeline worth studying, but the empirical table has a data-integrity red flag that blocks the headline claim until raw scores are released. read the letter →

arxiv 2608.09281 v1 pith:P3LVEWFP submitted 2026-08-10 cs.AI

classification cs.AI
keywords multimodallargelanguagemodelsbenchmarkconstructionarchitectureandcivilengineeringvisualreasoningprinciple-groundedperceive-know-judgeshort-answerevaluationerroranalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MMArch is a benchmark claiming that professional architectural reasoning cannot be reduced to reading a single figure or recalling text, and that current multimodal models are far from it. It consists of 1,212 short-answer items built entirely from figures in peer-reviewed papers across ten architecture and civil-engineering subdomains; every item is designed so that answering requires jointly perceiving visual evidence, identifying a governing engineering principle, and applying it. On these items the strongest open-source model scores about 30 percent and the best proprietary system about 52 percent, while a panel of human architects and engineers reaches 94.6 percent, a gap of more than forty points. The authors argue this gap is genuine reasoning failure rather than shortcut exploitation, since construction-time screening removes text-only, caption-only, and easy full-input solutions. The paper also reports that a third of model failures are composition errors, in which models extract all pieces correctly but fail to combine evidence with the governing principle.

What carries the argument

The load-bearing mechanism is the Perceive-Know-Judge decomposition combined with the decoupled planner-writer pipeline. Perceive extracts and compares visual evidence across figure panels; Know identifies the engineering principle, convention, or constraint that governs the case; Judge applies that principle to the observations to reach a short answer. Because the benchmark's construction freezes the answer before the question is written, prevents any candidate solvable from text or captions alone, and requires three architects and engineers to approve an item unanimously, the authors argue that achieving a score requires all three stages, so the Perceive-Know-Judge chain functions as a diagnostic lens for error analysis rather than a set of task labels.

What would settle it

Take a random sample of MMArch items and have independent professional engineers, who did not participate in construction, re-derive the answers from the original source papers while blind to the reference answers; if a substantial fraction of reference answers are disputed or corrected, the reported human-expert 94.6 percent can no longer be read as evidence that the items are well-posed.

Watch

Extended reading notes

Core claim

The central discovery the paper argues for is that principle-grounded visual reasoning in architecture and civil engineering can be measured separately from perception and memorization, and that on such a measure all evaluated multimodal models fall far short of human professionals. Each MMArch item is formalized as a triple of figure set, question, and answer, with source-paper passages kept secret from evaluated models; the authors define the required capability as the Perceive-Know-Judge chain in which every valid item requires all three stages. Using a decoupled planner-writer construction, automated leakage screening, a blind adversarial audit, and unanimous expert review, they report 1,212 validated items, and evaluation of 18 open-weight and proprietary models with deterministic rule-based scoring. The best proprietary model reaches 51.73 percent, the best open-source model 29.94 percent, and the human panel 94.57 percent; the paper interprets this gap as isolating joint evidence-and-principle reasoning rather than measuring perception alone.

Load-bearing premise

The benchmark's reference answers are defined by what the source papers' text and figures say, not by independent engineering verification, so a wrong, ambiguous, or curator-misread source would make the corresponding correct answer wrong no matter how well-posed the question looks.

Editorial extensions

If this is right

  • If MMArch measures what it claims, then current multimodal models cannot yet perform professional-grade principle-grounded visual reasoning in architecture and civil engineering.
  • Scaling model size alone brings only marginal gains, so further progress would require changes in architecture or training rather than simply more parameters.
  • Composition errors dominate the failure mix, meaning the primary bottleneck is combining evidence and principle, not finding the evidence in the first place.
  • Chain-of-thought and panel-identifier prompting produce small and inconsistent gains, indicating that the deficit is domain knowledge and compositional reasoning rather than lack of explicit reasoning traces.
  • The high human scores across subdomains support the paper's claim that the items are well-posed rather than noisy, making the model-human gap a usable diagnostic signal.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the source of difficulty shifts with model class, MMArch could serve not only as a leaderboard but as a diagnostic for targeted training: a model weak on spatial items likely lacks general-purpose visual grounding rather than domain text.
  • The decoupled planner-writer pipeline with answer freezing could be adapted to other expert domains with dense figures and assembled evidence, potentially reducing hallucinated ground truth in benchmark construction.
  • If composition errors remain dominant as models improve, a testable prediction is that architecture changes enforcing cross-figure binding will move scores more than additional data or larger parameter counts.
  • A natural extension is to measure whether giving models access to a corpus of domain principles at inference time, or fine-tuning on principle-application traces, closes more of the gap than simple prompting does.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces MMArch, a 1,212-item multimodal short-answer benchmark for architecture and civil engineering, built entirely from figures in peer-reviewed papers. Items are generated by a decoupled planner–writer pipeline with answer freezing, then filtered by automated text/caption leakage checks, a blind adversarial audit, two-path verification, and unanimous expert review. The authors evaluate 18 open-weight and proprietary MLLMs against a panel of human experts and report a large gap: the strongest open-source model reaches about 30%, the best proprietary model 51.73%, and human experts 94.57%. An error analysis attributes most failures to composition errors, with principle and perception errors next. The central claim is that MMArch isolates principle-grounded multimodal reasoning rather than single-figure perception or textual recall.

Significance. If the reported measurements are reliable, MMArch fills a genuine gap in AEC and scientific-figure benchmarking: it requires joint perception, domain knowledge, and judgment, and it incorporates an unusually thorough construction-time quality stack (answer freezing, evidence leads, blind adversarial audit, dual-path verification, deterministic rule-based scoring, and unanimous expert review). These methodological choices are a real strength and are described in commendable detail. However, the empirical contribution rests entirely on the accuracy numbers in Table 2, and those numbers currently display a regularity that is not consistent with the stated counting procedure. The headline 30%/52%/95% gap, the human baseline, and the error-analysis distribution are therefore not yet auditable. The benchmark construction protocol could still be valuable even if the table requires correction, but the central empirical claims cannot be accepted without raw data and exact counts.

major comments (4)
  1. [Section 4.1, Table 2] The repeated decimal endings in Table 2 are not compatible with the stated scoring rule. Under Section 4.1, accuracy is computed as 100*k_i/N_i for integer counts k_i and fixed per-subdomain totals N_i. Yet the Qwen3.6-27B row has all ten subdomain values and the average ending in .33 (17.33, 25.33, ..., 20.33), and the Qwen3.5-27B row has all values ending in .05 (21.05, 14.05, ..., 19.05). For independently measured subdomains with different denominators, such uniform endings would require a very specific modular coincidence across all ten N_i and k_i, and repeating across multiple rows makes the coincidence vanishingly unlikely. The paper releases neither per-item predictions nor raw counts, reports no confidence intervals, and gives no per-subdomain item counts. Unless the integer counts are shown to reproduce the printed percentages, the headline 30%/52%/95% gap and the error-analysis distribution are unsupported. Please release the raw predictions and exact numerator/denominator values, or correct the table and the evaluation description.
  2. [Section 4.1, Table 2] The human baseline is essential to the paper's central claim of a >40-point gap and to the inference that items are well-posed, but the description is under-specified. The paper gives no panel size, no per-subdomain denominator for the Human row, no inter-annotator agreement, and no statement about whether the human experts answered under the same input conditions as the models (e.g., without captions or source text). The Urban subdomain is reported at 100.00%, which with a finite item count implies a perfect score; no confidence interval is provided. Please report the number of experts, per-expert accuracy, the exact answer-matching protocol used for humans, and whether the human panel had access to source-paper text.
  3. [Section 3.2, Section 4.2] The ground truth is defined relative to source-paper evidence leads, and the 94.6% human score is interpreted as confirming item well-posedness. However, because the source passages are curator-only metadata and are not independently verified, a source-paper error or a curator misreading would make the reference answer wrong regardless of expert unanimity. Human experts who solve an item by applying the same paper-derived reasoning could reproduce the same wrong answer, so the high human score cannot by itself distinguish 'well-posed item' from 'experts reproduced the paper-derived answer'. The authors should add an independent provenance audit on a sample of items (for example, checking answers against codes, standards, or independent calculations) and report how many items had to be revised or discarded.
  4. [Section 5] The error analysis is described as a manual inspection of 'every model failure', but the paper reports no inter-annotator reliability, no per-model error counts, and no explicit definitions of the boundaries between the five categories. For example, the distinction between Principle Error and Composition Error is likely to be difficult to apply consistently, and Figure 6 does not show error-category counts, only unspecified proportions. Please provide a codebook, reliability statistics, and raw per-item error annotations, or at least a per-model table of error counts.
minor comments (6)
  1. [Section 3.2] Please state explicitly what the writer receives when composing the question: does it see only the figure and the frozen answer, or also the evidence lead and source passage? This is needed to make the decoupling argument fully checkable.
  2. [Section 4.1] The scoring normalization is described as accounting for 'units, synonyms, numeric tolerance, and accepted answer sets', but the actual normalization rules and tolerances are not specified. Without these, the reported accuracies cannot be reproduced independently.
  3. [Section 3.1] In the sentence 'This yields 1,212 validated short-answer items, each anchored to a specific engineering principle and its supporting visual evidence', the preceding text says the pipeline 'authors and validates each item'; this should be 'author and validate' or similar.
  4. [Figure 2, Table 2] Figure 2 is said to report per-subdomain item counts and proportions, but the figure in the manuscript does not show numeric counts or proportions. Table 2 also lacks per-subdomain item counts, which would help contextualize the percentages.
  5. [Figure 6] Figure 6 has no axis labels or legend, making it difficult to interpret the per-model error mix. Please add labels and a key.
  6. [Section 3.2] The paper describes the automated screening thresholds (8 full-input trials, 4 question-only, 2 caption-only; Hard/Challenge/0-8 boundaries) but gives no funnel counts of how many candidates were discarded at each stage. Reporting a funnel would strengthen confidence in the yield of 1,212 items.

Circularity Check

1 steps flagged · score 1.0 of 10

MMArch's headline model–human gap is an independent empirical measurement; only minor circularity is the expert-validation loop used as evidence of item well-posedness.

  1. other [Section 4.2 (Main Results); compare Section 3.2 (Quality Control)]
    "The consistently high human scores, between 89% and 100% across subdomains, further confirm that the items are well-posed rather than noisy."

    Section 3.2 retains an item only after 'three professional architects and engineers review all remaining and deferred items and retain an item only by unanimous agreement,' and each question must be 'answerable in principle by a domain expert.' The later claim that high human scores confirm item well-posedness therefore re-asserts the inclusion criterion rather than providing independent confirmation. The human panel is separate from construction, and the central model-versus-human gap is measured on the same items for both groups, so this is a mild validation loop rather than a fitted prediction.

full rationale

The derivation chain in MMArch is largely self-contained. Ground truth comes from curator-selected evidence leads in source papers, the answer is frozen before question writing, models are evaluated with deterministic decoding, and the human panel is disjoint from construction. No fitted parameter is renamed as a prediction, and no load-bearing result rests on a self-citation chain: all references are to external prior work. The only mild circularity is interpretational: items were retained only through unanimous expert agreement, so using the later high human scores as evidence of well-posedness partly restates the construction criterion. This does not affect the empirical model–human gap, which remains an independent measurement. The unusual decimal endings in Table 2 are a data-integrity and correctness concern, not a circularity, under the rule that circularity must be exhibited as a reduction to the paper's own inputs.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on construction thresholds rather than fitted scientific constants. The main axioms are the validity of paper-derived ground truth, expert review as a correctness guarantee, and the effectiveness of the decoupled pipeline. The benchmark introduces no physical or mathematical entities.

free parameters (4)
  • Answer span limit = 10 tokens
    Hand-set cap on reference answers; affects item brevity and scoring granularity (Section 3.2).
  • Full-input easy threshold = >=5 of 8 trials solved
    Items solved in at least five of eight full-input trials are discarded as too easy; hand-set difficulty cutoff (Section 3.2).
  • Leakage trial counts = 8 full, 4 question-only, 2 caption-only
    Hand-set number of solver trials used to detect text/caption leakage (Section 3.2).
  • Two-path verification tolerance = unspecified fixed unit, rounding, tolerance
    Dual-path answer agreement criterion; tolerance is stated but not quantified (Section 3.2).
assumptions (4)
  • domain assumption Peer-reviewed AEC papers provide correct ground-truth answers for engineering questions derived from their figures.
    The benchmark is built entirely from figures in peer-reviewed papers and uses evidence leads from the source text (Section 3.2).
  • domain assumption Unanimous expert review guarantees item correctness and unambiguity.
    Three professional architects and engineers retain an item only by unanimous agreement (Section 3.2).
  • ad hoc to paper The planner-writer decoupling prevents question-answer contamination even when both roles are filled by LLMs.
    The paper asserts answer freezing prevents the two roles from tailoring to each other (Section 3.2); this is a design assumption rather than a proven property.
  • domain assumption Manual failure categorization is reliable without measuring annotator agreement.
    Error analysis manually inspects and categorizes every model failure but reports no inter-annotator agreement statistics (Section 5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence." pith.science (2026). https://pith.science/paper/P3LVEWFP

@misc{pith2026260809281,
  author       = {Pith},
  title        = {Pith review of: MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P3LVEWFP}},
  note         = {Machine review of arXiv:2608.09281}
}
abstract

Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion. We introduce MMArch, a benchmark for architecture and civil engineering spanning ten subdomains and built entirely from figures in peer-reviewed papers. Its $1{,}212$ short-answer items are produced by a decoupled planner--writer pipeline and validated through automated screening, a blind adversarial audit, and expert review, so that answering requires perceiving the relevant evidence, identifying the governing principle, and applying it, not exploiting textual or single-figure shortcuts. Evaluating $18$ open-weight and proprietary MLLMs against a domain-expert panel, we find a wide gap: the strongest open-source model attains about $30\%$ and the best proprietary system $52\%$, while human experts reach $95\%$, more than forty points ahead. Our error analysis shows that failures concentrate in applying principles and combining evidence across figures rather than in locating it, pointing to substantial headroom for future research. Code and data are available at https://dcx-swjtu.github.io/MMArch/.

Figures

Figures reproduced from arXiv: 2608.09281 by the authors.

Figure 1
Figure 1. Overview of MMArch. Left: representative items that require reading evidence jointly across figure panels and applying [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The ten subdomains covered by MMArch, with [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Illustration of the MMArch construction pipeline: images are collected from peer-reviewed academic papers, relevant [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 6
Figure 6. Figure 6: Per-model distribution of correct answers and error [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 5
Figure 5. Figure 5: Illustration of four error types identified in MLLM [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 53 canonical work pages

  1. [1]

    Chen, L.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Wang, J.; Qiao, Y.; Lin, D.; and Zhao, F. 2024. Are We on the Right Way for Evaluating Large Vision-Language Models? In Advances in Neural Information Processing Systems (NeurIPS), volume 37

  2. [2]

    C.; Grandi, D.; Tomich, R.; Alam, M

    Doris, A. C.; Grandi, D.; Tomich, R.; Alam, M. F.; Ataei, M.; Cheong, H.; and Ahmed, F. 2025. DesignQA : A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation. Journal of Computing and Information Science in Engineering, 25(2): 021009

  3. [3]

    Fan, Z.; Zhu, L.; Li, H.; Chen, X.; Zhu, S.; and Tan, P. 2021. FloorPlanCAD : A Large-Scale CAD Drawing Dataset for Panoptic Symbol Spotting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10128--10137

  4. [4]

    Fu, C.; Chen, P.; Shen, Y.; Qin, Y.; Zhang, M.; Lin, X.; Yang, J.; Zheng, X.; Li, K.; Sun, X.; Wu, Y.; Ji, R.; Shan, C.; and He, R. 2025. MME : A Comprehensive Evaluation Benchmark for Multimodal Large Language Models. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, volume 38

  5. [5]

    E.; Michalski, V.; Atkinson, A.; K \'a d \'a r, \'A .; Trischler, A.; and Bengio, Y

    Kahou, S. E.; Michalski, V.; Atkinson, A.; K \'a d \'a r, \'A .; Trischler, A.; and Bengio, Y. 2017. FigureQA : An Annotated Figure Dataset for Visual Reasoning. arXiv:1710.07300

  6. [6]

    Kalervo, A.; Ylioinas, J.; H \"a iki \"o , M.; Karhu, A.; and Kannala, J. 2019. CubiCasa5K : A Dataset and an Improved Multi-task Model for Floorplan Image Analysis. In Image Analysis: 21st Scandinavian Conference, SCIA 2019, volume 11482 of Lecture Notes in Computer Science, 28--40. Springer

  7. [7]

    Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; and Farhadi, A. 2016. A Diagram is Worth a Dozen Images. In Computer Vision -- ECCV 2016, volume 9908 of Lecture Notes in Computer Science, 235--251. Springer

  8. [8]

    Kiela, D.; Bartolo, M.; Nie, Y.; Kaushik, D.; Geiger, A.; Wu, Z.; Vidgen, B.; Prasad, G.; Singh, A.; Ringshia, P.; Ma, Z.; Thrush, T.; Riedel, S.; Waseem, Z.; Stenetorp, P.; Jia, R.; Bansal, M.; Potts, C.; and Williams, A. 2021. Dynabench: Rethinking Benchmarking in NLP . In Proceedings of the 2021 Conference of the North American Chapter of the Associati...

Show all 57 references
  1. [10]

    Kunz, L.; Klostermeier, M.; Thanabalan, K.; Legler, T.; and Ruskowski, M. 2025. TechMB : Exploring the Potential of Vision Language Models for Interpreting Technical Drawings. In DS 140: Proceedings of the 36th Symposium Design for X (DFX2025), 179--188. The Design Society

  2. [11]

    Li, B.; Ge, Y.; Ge, Y.; Wang, G.; Wang, R.; Zhang, R.; and Shan, Y. 2024 a . SEED-Bench : Benchmarking Multimodal Large Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13299--13308

  3. [13]

    Liang, C.; Huang, Z.; Wang, H.; Chai, F.; Yu, C.; Wei, H.; Liu, Z.; Li, Y.; Wang, H.; Luo, R.; and Zhao, X. 2026. AECBench : A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field. Advanced Engineering Informatics, 71: 104314

  4. [14]

    Liu, Y.; Duan, H.; Zhang, Y.; Li, B.; Zhang, S.; Zhao, W.; Yuan, Y.; Wang, J.; He, C.; Liu, Z.; Chen, K.; and Lin, D. 2024. MMBench : Is Your Multi-modal Model an All-Around Player? In European Conference on Computer Vision (ECCV), 216--233. Springer

  5. [15]

    Lu, P.; Bansal, H.; Xia, T.; Liu, J.; Li, C.; Hajishirzi, H.; Cheng, H.; Chang, K.-W.; Galley, M.; and Gao, J. 2024. MathVista : Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. In International Conference on Learning Representations (ICLR)

  6. [16]

    Lu, P.; Mishra, S.; Xia, T.; Qiu, L.; Chang, K.-W.; Zhu, S.-C.; Tafjord, O.; Clark, P.; and Kalyan, A. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In Advances in Neural Information Processing Systems (NeurIPS), volume 35

  7. [17]

    Luo, R.; Liu, Z.; Cheng, T.; Wang, J.; Wang, T.; Wei, X.; Wang, H.; Li, Y.; Chai, F.; Cheng, F.; Ye, S.; Wang, W.; Zhang, Y.; Qiao, Y.; Zhang, H.; and Zhao, X. 2025. ArchCAD-400K : A Large-Scale CAD Drawings Dataset and New Baseline for Panoptic Symbol Spotting. In Advances in...

  8. [19]

    X.; Tan, J

    Masry, A.; Long, D. X.; Tan, J. Q.; Joty, S.; and Hoque, E. 2022. ChartQA : A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, 2263--2279. Dublin, Ireland: Association for C...

  9. [20]

    Mathew, M.; Bagal, V.; Tito, R.; Karatzas, D.; Valveny, E.; and Jawahar, C. V. 2022. InfographicVQA . In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 1697--1706

  10. [21]

    Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021. DocVQA : A Dataset for VQA on Document Images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2200--2209

  11. [22]

    M.; and Kumar, P

    Methani, N.; Ganguly, P.; Khapra, M. M.; and Kumar, P. 2020. PlotQA : Reasoning over Scientific Plots. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), 1516--1525

  12. [23]

    Nie, Y.; Williams, A.; Dinan, E.; Bansal, M.; Weston, J.; and Kiela, D. 2020. Adversarial NLI : A New Benchmark for Natural Language Understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 4885--4901. Online: Association for C...

  13. [24]

    Pramanick, S.; Chellappa, R.; and Venugopalan, S. 2024. SPIQA : A Dataset for Multimodal Question Answering on Scientific Papers. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, volume 37

  14. [25]

    Roberts, J.; Han, K.; Houlsby, N.; and Albanie, S. 2024. SciFIBench : Benchmarking Large Multimodal Models for Scientific Figure Interpretation. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, volume 37

  15. [26]

    Wang, Z.; Xia, M.; He, L.; Chen, H.; Liu, Y.; Zhu, R.; Liang, K.; Wu, X.; Liu, H.; Malladi, S.; Chevalier, A.; Arora, S.; and Chen, D. 2024. CharXiv : Charting Gaps in Realistic Chart Understanding in Multimodal LLMs . In Advances in Neural Information Processing Systems (Neur...

  16. [27]

    Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; Wei, C.; Yu, B.; Yuan, R.; Sun, R.; Yin, M.; Zheng, B.; Yang, Z.; Liu, Y.; Huang, W.; Sun, H.; Su, Y.; and Chen, W. 2024. MMMU : A Massive Multi-discipline Multimodal Understan...

  17. [28]

    Yue, X.; Zheng, T.; Ni, Y.; Wang, Y.; Zhang, K.; Tong, S.; Sun, Y.; Yu, B.; Zhang, G.; Sun, H.; Su, Y.; Chen, W.; and Neubig, G. 2025. MMMU-Pro : A More Robust Multi-discipline Multimodal Understanding Benchmark. In Proceedings of the 63rd Annual Meeting of the Association for...

  18. [29]

    Zellers, R.; Bisk, Y.; Schwartz, R.; and Choi, Y. 2018. SWAG : A Large-Scale Adversarial Dataset for Grounded Commonsense Inference. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), 93--104. Brussels, Belgium: Association for C...

  19. [30]

    Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. HellaSwag : Can a Machine Really Finish Your Sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 4791--4800. Florence, Italy: Association for Computational Li...

  20. [31]

    Image Analysis: 21st Scandinavian Conference, SCIA 2019 , series =

    Ahti Kalervo and Juha Ylioinas and Markus H. Image Analysis: 21st Scandinavian Conference, SCIA 2019 , series =

  21. [32]

    Zhiwen Fan and Lingjie Zhu and Honghua Li and Xiaohao Chen and Siyu Zhu and Ping Tan , booktitle =

  22. [33]

    Ruifeng Luo and Zhengjie Liu and Tianxiao Cheng and Jie Wang and Tongjie Wang and Xingguang Wei and Haomin Wang and YanPeng Li and Fu Chai and Fei Cheng and Shenglong Ye and Wenhai Wang and Yanting Zhang and Yu Qiao and Hongjie Zhang and Xianzhong Zhao , booktitle =

  23. [34]

    Hsain and Guido Maciocci , year =

    Aleksei Kondratenko and Mussie Birhane and Houssame E. Hsain and Guido Maciocci , year =. 2601.04819 , archivePrefix =

  24. [35]

    Chen Liang and Zhaoqi Huang and Haofen Wang and Fu Chai and Chunying Yu and Huanhuan Wei and Zhengjie Liu and Yanpeng Li and Hongjun Wang and Ruifeng Luo and Xianzhong Zhao , journal =

  25. [36]

    Doris and Daniele Grandi and Ryan Tomich and Md Ferdous Alam and Mohammadmehdi Ataei and Hyunmin Cheong and Faez Ahmed , journal =

    Anna C. Doris and Daniele Grandi and Ryan Tomich and Md Ferdous Alam and Mohammadmehdi Ataei and Hyunmin Cheong and Faez Ahmed , journal =

  26. [37]

    Leonhard Kunz and Mario Klostermeier and Kokulan Thanabalan and Tatjana Legler and Martin Ruskowski , booktitle =

  27. [38]

    2603.29199 , archivePrefix =

    Harsh Mankodiya and Chase Gallik and Theodoros Galanos and Andriy Mulyar , year =. 2603.29199 , archivePrefix =

  28. [39]

    2017 , eprint =

    Samira Ebrahimi Kahou and Vincent Michalski and Adam Atkinson and. 2017 , eprint =

  29. [40]

    Khapra and Pratyush Kumar , booktitle =

    Nitesh Methani and Pritha Ganguly and Mitesh M. Khapra and Pratyush Kumar , booktitle =

  30. [41]

    Ahmed Masry and Do Xuan Long and Jia Qing Tan and Shafiq Joty and Enamul Hoque , booktitle =

  31. [42]

    Minesh Mathew and Dimosthenis Karatzas and C. V. Jawahar , booktitle =

  32. [43]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages =

    Minesh Mathew and Viraj Bagal and Rub. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pages =

  33. [44]

    Computer Vision -- ECCV 2016 , series =

    A Diagram is Worth a Dozen Images , author =. Computer Vision -- ECCV 2016 , series =

  34. [45]

    Zirui Wang and Mengzhou Xia and Luxi He and Howard Chen and Yitao Liu and Richard Zhu and Kaiqu Liang and Xindi Wu and Haotian Liu and Sadhika Malladi and Alexis Chevalier and Sanjeev Arora and Danqi Chen , booktitle =

  35. [46]

    Shraman Pramanick and Rama Chellappa and Subhashini Venugopalan , booktitle =

  36. [47]

    Jonathan Roberts and Kai Han and Neil Houlsby and Samuel Albanie , booktitle =

  37. [48]

    Wilson and Woosang Lim and William Yang Wang , year =

    Zekun Li and Xianjun Yang and Kyuri Choi and Wanrong Zhu and Ryan Hsieh and HyeonJung Kim and Jin Hyuk Lim and Sungyoung Ji and Byungju Lee and Xifeng Yan and Linda Ruth Petzold and Stephen D. Wilson and Woosang Lim and William Yang Wang , year =. 2407.04903 , archivePrefix =

  38. [49]

    Advances in Neural Information Processing Systems (NeurIPS) , volume =

    Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering , author =. Advances in Neural Information Processing Systems (NeurIPS) , volume =

  39. [50]

    Pan Lu and Hritik Bansal and Tony Xia and Jiacheng Liu and Chunyuan Li and Hannaneh Hajishirzi and Hao Cheng and Kai-Wei Chang and Michel Galley and Jianfeng Gao , booktitle =

  40. [51]

    Xiang Yue and Yuansheng Ni and Kai Zhang and Tianyu Zheng and Ruoqi Liu and Ge Zhang and Samuel Stevens and Dongfu Jiang and Weiming Ren and Yuxuan Sun and Cong Wei and Botao Yu and Ruibin Yuan and Renliang Sun and Ming Yin and Boyuan Zheng and Zhenzhu Yang and Yibo Liu and We...

  41. [52]

    Chaoyou Fu and Peixian Chen and Yunhang Shen and Yulei Qin and Mengdan Zhang and Xu Lin and Jinrui Yang and Xiawu Zheng and Ke Li and Xing Sun and Yunsheng Wu and Rongrong Ji and Caifeng Shan and Ran He , booktitle =

  42. [53]

    Yuan Liu and Haodong Duan and Yuanhan Zhang and Bo Li and Songyang Zhang and Wangbo Zhao and Yike Yuan and Jiaqi Wang and Conghui He and Ziwei Liu and Kai Chen and Dahua Lin , booktitle =

  43. [54]

    Bohao Li and Yuying Ge and Yixiao Ge and Guangzhi Wang and Rui Wang and Ruimao Zhang and Ying Shan , booktitle =

  44. [55]

    Advances in Neural Information Processing Systems (NeurIPS) , volume =

    Are We on the Right Way for Evaluating Large Vision-Language Models? , author =. Advances in Neural Information Processing Systems (NeurIPS) , volume =

  45. [56]

    Xiang Yue and Tianyu Zheng and Yuansheng Ni and Yubo Wang and Kai Zhang and Shengbang Tong and Yuxuan Sun and Botao Yu and Ge Zhang and Huan Sun and Yu Su and Wenhu Chen and Graham Neubig , booktitle =

  46. [57]

    Rowan Zellers and Yonatan Bisk and Roy Schwartz and Yejin Choi , booktitle =

  47. [58]

    Rowan Zellers and Ari Holtzman and Yonatan Bisk and Ali Farhadi and Yejin Choi , booktitle =

  48. [59]

    Adversarial

    Yixin Nie and Adina Williams and Emily Dinan and Mohit Bansal and Jason Weston and Douwe Kiela , booktitle =. Adversarial

  49. [60]

    Dynabench: Rethinking Benchmarking in

    Douwe Kiela and Max Bartolo and Yixin Nie and Divyansh Kaushik and Atticus Geiger and Zhengxuan Wu and Bertie Vidgen and Grusha Prasad and Amanpreet Singh and Pratik Ringshia and Zhiyi Ma and Tristan Thrush and Sebastian Riedel and Zeerak Waseem and Pontus Stenetorp and Robin ...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.