Pith. sign in

REVIEW 4 major objections 5 minor 74 references

Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper claims that a multimodal large language model, prompted with named perceptual principles, can rank images by visual complexity without any human labels for training, and that the resulting comparisons expose a systematic bias…

desk verdict Useful benchmark result with an overreaching cognitive claim; the principle-judgment bridge needs validation before the bias interpretation holds. read the letter →

arxiv 2504.12511 v1 pith:774N4IHZ submitted 2025-04-16 cs.HC cs.AIcs.CVcs.LG

classification cs.HCcs.AIcs.CVcs.LG
keywords multimodallargelanguagemodelsvisualcomplexityGestaltprinciplesclutterlawofsimplicitypairwisecomparisonhumanannotationbiasinterpretability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes that a multimodal large language model, prompted with psychology's Gestalt principles, can serve as an annotation-free cognitive assistant for visual complexity analysis. The authors ask the model to compare pairs of images along eight interpretable dimensions—six Gestalt laws plus visual clutter and visual symmetry—and convert those pairwise judgments into per-image scores. Across the SAVOIAS and IC9600 datasets, scores for visual clutter and the law of simplicity correlate most strongly with human complexity annotations. The authors read this as evidence that human annotators in those datasets weight clutter and simplicity heavily and neglect other principles from the perception literature, and that a prompt-constrained MLLM can expose such biases without needing a training set.

What carries the argument

The load-bearing object is the pairwise comparison matrix $S=\{s_{i,j}\}$, where $s_{i,j}$ is the MLLM's binary judgment of which image in a pair better exemplifies an explainable parameter, aggregated into $\hat{s}_i = \frac{1}{n}\sum_{j=1}^{n} s_{i,j}$. The parameters are six Gestalt principles (similarity, proximity, simplicity, closure, continuity, figure/ground) plus visual clutter and visual symmetry, defined in the prompt as they would be given to a human annotator. The matrix converts free-form MLLM reasoning into ranked scores per principle, which are then correlated with human complexity labels via Pearson and Spearman coefficients. Pairwise comparison is chosen over absolute ratings to avoid annotator scale bias and context limits, and the model is run at temperature 0.01 with requested justifications to stabilize outputs.

What would settle it

Take a random sample of images from SAVOIAS and IC9600, collect human pairwise judgments on each of the eight principles, and compute an agreement matrix among principles and with human complexity ratings. If humans' own clutter ratings track their overall complexity ratings as strongly as the model's do, or if the model's eight scores collapse into a single general-complexity factor, then the claimed demonstration of annotator bias toward clutter and simplicity is unsupported.

Watch

Extended reading notes

Core claim

The paper's central claim is that an MLLM can reason about visual complexity through named perceptual principles, and that doing so reveals a human bias. Questioned pairwise with definitions of eight principles, Claude Sonnet 3.0 produces rankings whose clutter and simplicity scores correlate with human complexity ratings at roughly 0.62–0.82 (PLCC/SROCC) across categories in SAVOIAS and IC9600, consistently higher than similarity, proximity, closure, continuity, figure/ground, or symmetry. The authors conclude that human annotators behind these datasets are biased toward visual clutter and visual simplicity, neglecting other reasonings proposed in psychology and cognitive science. They also report category-specific effects, such as law-of-closure correlations that are weak for advertisements but stronger for suprematism and paintings, and stable advertisement-category results across both datasets. The framework is presented as a scalable, annotation-free alternative to deep-learning complexity predictors, aimed at HCI tasks rather than at forecasting complexity scores.

Load-bearing premise

The load-bearing premise is that the MLLM's 'clutter' and 'simplicity' pairwise judgments measure those named principles the way a human would and are not just the model's overall complexity impression; if that bridge fails, the correlation cannot be read as evidence about human annotator bias.

Editorial extensions

If this is right

  • If the bias claim holds, human-annotated visual complexity datasets should not be treated as measuring general complexity; they largely measure clutter and simplicity.
  • The same annotation-free protocol can be reused on new image categories without retraining, because the MLLM is constrained by prompt definitions rather than by dataset labels.
  • Practical HCI applications follow directly: designers and content creators can receive explainable per-principle feedback on visual balance, clutter, and symmetry, and platforms can adjust search-result presentation toward principles that reduce cognitive load.
  • Category-level differences imply that an interface tuned for one domain (e.g., advertisements) may not transfer to another (e.g., paintings), because the relevant perceptual principle changes.
  • The finding suggests that datasets like SAVOIAS and IC9600 carry systematic annotator bias that downstream models trained on them will inherit.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper never validates the model's eight principle judgments against human judgments of those same eight principles; a direct experiment collecting per-principle human pairwise labels would separate the claim about human annotators from the claim about the model's own perceptual axis.
  • If future work finds that the model's clutter and simplicity scores are nearly collinear with its general-complexity score, the 'bias' result reduces to the weaker statement that one complexity factor drives both; that would be a confound the current correlation table cannot rule out.
  • Because prompt sensitivity is explicitly left unexamined, the quantitative rankings plausibly depend on the exact wording and model version; re-running with prompts that instruct the model to weigh symmetry or closure deliberately would test whether the clutter/simplicity dominance is stable.
  • A natural extension is to apply the same pairwise protocol to other subjective annotations—aesthetic appeal, trust, cognitive load—to ask whether clutter and simplicity dominate human judgment there as well, or whether the bias is specific to complexity.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript proposes an annotation-free framework for assessing whether multimodal LLMs can reason about visual complexity using eight explainable principles drawn from Gestalt psychology and visual perception: similarity, proximity, simplicity, closure, continuity, figure/ground, visual clutter, and visual symmetry. Claude Sonnet 3.0 performs pairwise comparisons of images under each principle; the binary results are aggregated into per-image scores (Section 4.2) and correlated with human complexity annotations on the SAVOIAS and IC9600 datasets across several image categories (Section 4.3, Tables 1-2). The paper reports that visual clutter and law of simplicity consistently yield the highest PLCC/SROCC values and interprets this pattern as evidence that human annotators of those datasets are biased toward clutter and simplicity while neglecting other psychological principles (Section 5). It also outlines HCI applications and future work (Section 6).

Significance. The framework is a useful low-cost benchmarking approach: it involves no fitted parameters, uses externally supplied human complexity labels, and makes the principle-level comparisons transparent. The consistent top ranking of clutter and simplicity across two datasets and most categories is an interesting empirical pattern that corroborates prior work on visual complexity and is worth reporting. The main load-bearing weakness is conceptual rather than computational: the inference from MLLM score-complexity correlations to human annotator bias requires validating that the MLLM's 'clutter' and 'simplicity' judgments track those specific constructs as humans would judge them, and that they are distinct from a general complexity impression. That validation is absent. The paper also ships no uncertainty quantification, which weakens the comparative claims.

major comments (4)
  1. [§5, with §4.2-4.3] The central claim that human annotators are biased toward visual clutter and simplicity is an interpretive bridge that is not tested. The scores s_i in Section 4.2 are derived from binary comparisons made by Claude Sonnet 3.0 'based on' each principle, but the only external validation in Section 4.3 is against human complexity labels. The paper neither collects human judgments of the same eight principles nor reports an agreement or discriminability matrix among the principle scores. Since visual clutter and simplicity are semantically close to complexity itself, and the paper cites [54] in connecting complexity to clutter, the high correlations in Tables 1-2 could be produced by a model that is essentially ranking images by overall complexity under different prompt phrasings; in that case the conclusion about human cognition would not follow. The paper's own Section 6 concession that prompt sensitivity is unexplored points to the same gap. A concrete remedy is to validate the principle judgments against human ratings of those principles and to include prompt-ablation controls demonstrating that the eight axes do not collapse into one.
  2. [Tables 1 and 2, §4.3/§5] The reported PLCC and SROCC values are point estimates without sample sizes, confidence intervals, or significance tests. The statement in Section 5 that clutter and simplicity have 'highest correlation' is a ranking of noisy estimates; for example, in Table 2 the advertisement row separates Visual Clutter (0.76/0.76) from Law of Simplicity (0.70/0.76) by amounts that may be within sampling variability. The related claim that these measures are 'consistent across different categories' needs an explicit test or at least per-category uncertainty to be supported. Without this, the empirical foundation for the paper's strongest comparative conclusion is incomplete.
  3. [§4.2 and §4.4] The aggregate scores rest on a single MLLM at temperature 0.01, with no repeated sampling, no reported agreement between runs, and no description of how ties or invalid comparison outputs are handled. Because the analysis compares small differences in correlations, the reliability of the pairwise-comparison step should be demonstrated. Additionally, conclusions about human annotators are drawn from one model; Section 4.4 explains why other tested MLLMs were not usable, but the human-bias claim should be explicitly framed as a property of this model unless corroborated by additional models or by direct human-principle ratings.
  4. [§6 (Applications & Future Work)] The manuscript explicitly states that 'There is room for improvement in understanding the sensitivity of MLLM methods to the quality and design of the prompts.' This limitation is load-bearing for the Section 5 conclusion: if the ranking of principles changes under different prompt wording, the observed correlation pattern may reflect prompt design rather than a stable property of human annotations. The paper should either provide a prompt-sensitivity analysis or restrict the conclusion to the specific prompt protocol used.
minor comments (5)
  1. [§4] The list of Gestalt principles contains a duplicate: 'laws of similarity, proximity, simplicity, simplicity, closure, continuity and figure vs. ground.' The second 'simplicity' should be removed; the set should be six principles plus clutter and symmetry.
  2. [§2 and §5] The contribution states the framework is 'not susceptible to annotator biases,' while Section 5 argues that the human-annotated datasets are biased. This is not an outright contradiction if read as 'does not inherit training-set annotation bias,' but the wording should be clarified to avoid the apparent tension.
  3. [Throughout] There are repeated typos and inconsistent naming ('SAVIOAS' vs 'SAVOIAS', 'Explainabile AI', 'populalrly', 'reat time', 'interprete', 'monotocity'). A careful proofread is needed.
  4. [Figure 5 and Tables 1-2] The sample sizes for each category are not reported, and the figures and captions do not indicate uncertainty; the reader cannot judge how many images underlie each coefficient or whether the visual ordering of bars is statistically meaningful.
  5. [Title and framing] The title uses 'Interpretable' to describe the framework, but the paper measures correlations between principle-based scores and human labels; this is better described as principle-aligned scoring rather than interpretability of the model's internal representations. Clarifying this terminology would make the contribution easier to position.

Circularity Check

1 steps flagged · score 3.0 of 10

Main correlations rest on external human labels and are not circular, but the human-bias conclusion is partially self-definitional because the MLLM's clutter axis is prompted as a parameter of complexity evaluation and never validated against human clutter judgments.

  1. self definitional [Section 4.1 (Prompt Design) and Section 5 (Results & Discussion)]
    "We first develop a text prompt to describe the role of the MLLM agent as an HCI researcher whose goal is to evaluate the complexity of the input images. We then provide and define the explainable parameters of evaluating complexity in visual perception ... These principles include the 6 gestalt principles along with visual clutter and visual symmetry. ... We find that scores for visual clutter and law of simplicity based comparisons have highest correlation with human annotations. This demonstrates that most of the human annotators ..."

    The prompt frames visual clutter and simplicity as 'explainable parameters of evaluating complexity' and instructs the model to make its comparisons in that framing. Section 5 then treats the resulting MLLM clutter/simplicity scores' correlation with human complexity labels as evidence that human annotators are biased toward those principles. The paper never validates the MLLM's principle-specific judgments against human judgments of the same principles, and it reports no agreement matrix among the eight axes. Consequently, the high correlation can be explained by the construction of the prompt itself: the model's 'clutter' axis may simply be its general complexity estimate under another name.

full rationale

The numerical core of the paper is not circular: no parameters are fitted, the MLLM is not trained on SAVOIAS or IC9600, and the reported PLCC/SROCC values are correlations against external human complexity labels. There is also no load-bearing self-citation chain; the cited Gestalt and clutter literature is independent of the authors. The circularity concern is limited to the interpretive step in Section 5. Because the prompt explicitly defines visual clutter and simplicity as parameters for evaluating complexity, a high correlation between those parameter scores and human complexity labels is partly built into the measurement setup unless the MLLM's axes are shown to be distinct from an overall complexity impression. The paper does not supply that validation. The finding is therefore an interesting empirical correlation, but the 'demonstrates human annotator bias' claim is partially self-definitional and should be downgraded from a demonstrated result to a hypothesis requiring human principle-level judgments. A score of 3 reflects this partial, interpretive circularity rather than circularity in the main correlation results.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters are fitted: the pipeline is prompt-driven comparison, row-mean aggregation, and correlation, with no learned weights. The cost moves to domain assumptions: the eight principles are assumed complete and distinct, and the single model's judgments are assumed to be valid human-like measurements of each principle. No new entities are postulated. The paper's own Section 6 concedes prompt sensitivity is unexamined, which points at the same assumptions.

assumptions (4)
  • domain assumption The six Gestalt principles plus visual clutter and visual symmetry form a sufficient decomposition of visual complexity perception.
    Section 4 selects these eight parameters from the psychology literature and measures only them. The completeness and mutual independence of the list is assumed, and the headline result depends on clutter and simplicity being well-defined, non-overlapping operations for the model.
  • domain assumption Claude Sonnet 3.0's pairwise judgments validly measure each named principle as a human annotator would perceive it.
    Section 4.2 turns binary model comparisons into scores, and Section 5 interprets the correlation of those scores with human complexity labels as evidence about human annotator cognition. No human ground truth for the individual principles is used or reported.
  • domain assumption Row means of pairwise comparison outcomes are a valid global score; single-sample pairwise judgments are reliable enough to be treated as measurements.
    Section 4.2 defines s_i as the mean of binary outcomes and cites [3,19,34] for the pairwise-to-global approach. The paper reports no transitivity audit, no repeat-sampling variance, and no consistency check at temperature 0.01.
  • standard math Pearson and Spearman correlations of the aggregated scores against human labels correctly capture the association.
    Section 4.3 defines the metrics. The computation is standard, but no significance testing, confidence intervals, or error propagation are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis." pith.science (2026). https://pith.science/paper/774N4IHZ

@misc{pith2026250412511,
  author       = {Pith},
  title        = {Pith review of: Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/774N4IHZ}},
  note         = {Machine review of arXiv:2504.12511}
}
read the original abstract

In this paper, we advance the study of AI-augmented reasoning in the context of Human-Computer Interaction (HCI), psychology and cognitive science, focusing on the critical task of visual perception. Specifically, we investigate the applicability of Multimodal Large Language Models (MLLMs) in this domain. To this end, we leverage established principles and explanations from psychology and cognitive science related to complexity in human visual perception. We use them as guiding principles for the MLLMs to compare and interprete visual content. Our study aims to benchmark MLLMs across various explainability principles relevant to visual perception. Unlike recent approaches that primarily employ advanced deep learning models to predict complexity metrics from visual content, our work does not seek to develop a mere new predictive model. Instead, we propose a novel annotation-free analytical framework to assess utility of MLLMs as cognitive assistants for HCI tasks, using visual perception as a case study. The primary goal is to pave the way for principled study in quantifying and evaluating the interpretability of MLLMs for applications in improving human reasoning capability and uncovering biases in existing perception datasets annotated by humans.

Figures

Figures reproduced from arXiv: 2504.12511 by the authors.

Figure 1
Figure 1. Images ordered based on visual complexity across different categories. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Visual perception analysis framework using multimodal LLMs [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Various important components of the prompt to compare advertisement images based on visual clutter. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Sample MLLM (Claude Sonnet 3.0) output comparing two interior design images. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Results summary for SAVIOAS and IC9600 dataset. Detailed results can be found in tables 1 and 2 in the appendix. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 43 canonical work pages

  1. [54]

    Ruth Rosenholtz, Yuanzhen Li, and Lisa Nakano. 2007. Measuring visual clutter. Journal of Vision 7, 2 (2007), 17–17

  2. [1]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...

  3. [2]

    Lawrence Zitnick, and Devi Parikh

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual Question Answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)

  4. [3]

    K. J. Arrow. 1950. A difficulty in the concept of social welfare. Journal of Political Economy 58, 4 (1950), 328–346

  5. [4]

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv:2308.12966 [cs.CV] https://arxiv.org/abs/2308.12966

  6. [5]

    Amanda Baughan, Tal August, Naomi Yamashita, and Katharina Reinecke. 2020. Keep it Simple: How Visual Complexity and Preferences Impact Search Efficiency on Websites. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–10. https://doi.org/1...

  7. [6]

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)

  8. [7]

    Steven Bradley. 2019. The Role of Gestalt Principles in Data Visualization. Topcoder Blog (2019). https://www.topcoder.com/blog/gestalt-principles- for-data-visualization/

Show all 74 references
  1. [8]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  2. [9]

    Luigi Celona, Gianluigi Ciocca, and Raimondo Schettini. 2024. On the Use of Visual Transformer for Image Complexity Assessment. In Proceedings of the 19th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications - Volume 3: VISAP...

  3. [10]

    Soravit Changpinyo, Piyush Kumar Sharma, Nan Ding, and Radu Soricut. 2021. Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021), 3557–3567. https://ap...

  4. [11]

    Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li, Zechun Liu, Pengchuan Zhang, Raghuraman Krishnamoorthi, Vikas Chandra, Yunyang Xiong, and Mohamed Elhoseiny. 2023. MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning. ArXiv abs/2310.0947...

  5. [12]

    Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023. ShareGPT4V: Improving Large Multi-Modal Models with Better Captions. arXiv:2311.12793 [cs.CV] https://arxiv.org/abs/2311.12793

  6. [13]

    Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shakeri, Mostafa dehghani, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang, and Radu Soricut

  7. [14]

    Susan F Chipman. 1977. Complexity and structure in visual patterns. Journal of Experimental Psychology: General 106, 3 (1977), 269

  8. [15]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  9. [16]

    Mia Cinelli. 2020. Gestalt Principles and Their Application in UX Design. Maze Blog (2020). https://maze.co/blog/gestalt-principles/

  10. [17]

    Claude.ai. [n. d.]. Claude.ai article. ([n. d.]). https://support.anthropic.com/en/articles/9002500-what-kinds-of-images-can-i-upload-to-claude-ai Manuscript submitted to ACM Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis 9

  11. [18]

    Silvia Elena Corchs, Gianluigi Ciocca, Emanuela Bricolo, and Francesca Gasparini. 2016. Predicting complexity perception of real world images. PloS One 11, 6 (2016), e0157986

  12. [19]

    H. A. David. 1963. The Method of Paired Comparisons . Vol. 12. London

  13. [20]

    Mohsen Fayyaz et al. 2024. Evaluating Human Alignment and Model Faithfulness of LLM Rationale. arXiv preprint arXiv:2407.00219 (2024). https://arxiv.org/abs/2407.00219

  14. [21]

    Wang et al. 2024. Towards Reasoning in Large Language Models: A Survey.Journal of Machine Learning Research(2024). https://www.topbots.com/llm- reasoning-research-papers/

  15. [22]

    Zhen Bao Fan, Yi-Na Li, Jinhui Yu, and Kang Zhang. 2017. Visual complexity of Chinese ink paintings. In Proceedings of the ACM Symposium on Applied Perception. 1–8

  16. [23]

    Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. 2023. EVA: Exploring the Limits of Masked Visual Representation Learning at Scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...

  17. [24]

    Jacob Feldman. 2016. The simplicity principle in perception and cognition. WIREs Cognitive Science 7, 5 (2016), 330–340. https://doi.org/10.1002/wcs. 1406 arXiv:https://wires.onlinelibrary.wiley.com/doi/pdf/10.1002/wcs.1406

  18. [25]

    Tinglei Feng, Yingjie Zhai, Jufeng Yang, Jie Liang, Deng-Ping Fan, Jing Zhang, Ling Shao, and Dacheng Tao. 2022. IC9600: A benchmark dataset for automatic image complexity assessment. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)

  19. [26]

    T. Feng, Y. Zhai, J. Yang, J. Liang, D. P. Fan, J. Zhang, L. Shao, and D. Tao. 2023. IC9600: A Benchmark Dataset for Automatic Image Complexity Assessment. IEEE Trans Pattern Anal Mach Intell 45, 7 (Jul 2023), 8577–8593

  20. [27]

    Xingyu Fu, Sheng Zhang, Gukyeong Kwon, Pramuditha Perera, Henghui Zhu, Yuhao Zhang, Alexander Hanbo Li, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, Dan Roth, and Bing Xiang. 2023. Generate then Select: Open-ended Visual Question Answering Guided by World Kno...

  21. [28]

    Xingyu Fu, Ben Zhou, Ishaan Chandratreya, Carl Vondrick, and Dan Roth. 2022. There’s a Time and Place for Reasoning Beyond the Image. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav ...

  22. [29]

    Gauthier

    Thomas D. Gauthier. 2001. Detecting Trends Using Spearman’s Rank Correlation Coefficient. Environmental Forensics 2, 4 (2001), 359–362. https://doi.org/10.1006/enfo.2001.0061

  23. [30]

    Eline Van Geert and Johan Wagemans. 2020. Order, complexity, and aesthetic appreciation. Psychology of Aesthetics, Creativity, and the Arts 14, 2 (2020), 135

  24. [31]

    Yash Goyal, Tejas Khot, Aishwarya Agrawal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2019. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. Int. J. Comput. Vision 127, 4 (April 2019), 398–414. https://doi.org/10.1007...

  25. [32]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  26. [33]

    Heaps and S

    C. Heaps and S. Handel. 1999. Similarity and Features of Natural Textures. Journal of Experimental Psychology: Human Perception and Performance 25, 2 (1999), 299–320

  27. [34]

    M. G. Kendall and B. B. Smith. 1940. On the Method of Paired Comparisons. Biometrika 31, 3/4 (1940), 324–345

  28. [35]

    Andy J King, Allison J Lazard, and Shawna R White. 2020. The influence of visual complexity on initial user impressions: Testing the persuasive model of web design. Behaviour & Information Technology 39, 5 (2020), 497–510

  29. [36]

    Kurt Koffka. 1935. The Principles of Gestalt Psychology . Harcourt Brace

  30. [37]

    Shamma, Michael S

    Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotatio...

  31. [38]

    Cameron Kyle-Davidson and Karla Evans. 2023. Complexity & Memorability have a Nonlinear Relationship when Remembering Scenes. Journal of Vision 23 (08 2023), 5251. https://doi.org/10.1167/jov.23.9.5251

  32. [39]

    Wolfgang Köhler. 1947. Gestalt Psychology: An Introduction to New Concepts in Modern Psychology . Harcourt Brace

  33. [40]

    Bo Li, Hao Zhang, Kaichen Zhang, Dong Guo, Yuanhan Zhang, Renrui Zhang, Feng Li, Ziwei Liu, and Chunyuan Li. 2024. LLaVA-NeXT: What Else Influences Visual Instruction Tuning Beyond Data? https://llava-vl.github.io/blog/2024-05-25-llava-next-ablations/

  34. [41]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. arXiv:2301.12597 [cs.CV] https://arxiv.org/abs/2301.12597

  35. [42]

    T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Proceedings of European Conference on Computer Vision (ECCV) . 740–755. Manuscript submitted to ACM Multimodal LLM Augmented Reas...

  36. [43]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023. Improved Baselines with Visual Instruction Tuning. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023), 26286–26296. https://api.semanticscholar.org/CorpusID:263672058

  37. [44]

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024. LLaVA-NeXT: Improved reasoning, OCR, and world knowledge. https://llava-vl.github.io/blog/2024-01-30-llava-next/

  38. [45]

    Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024. DeepSeek-VL: Towards Real-World Vision-Language Understanding. arXiv:2403.05525 [cs.AI] htt...

  39. [46]

    G. A. Miller. 1956. The Magical Number Seven, Plus or Minus Two: Some Limits on Our Capacity for Processing Information. Psychological Review 63, 2 (1956), 81

  40. [47]

    Aliaksei Miniukovich and Maurizio Marchese. 2020. Relationship Between Visual Complexity and Aesthetics of Webpages. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY...

  41. [48]

    Surabhi S Nath, Kevin Shen, Aenne Annelie Brielmann, and Peter Dayan. 2024. Simplicity in Complexity. In ICLR 2024 Workshop on Representational Alignment. https://openreview.net/forum?id=DHvVdakpqO

  42. [49]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  43. [50]

    Letizia Palumbo, Ruth Ogden, Alexis DJ Makin, and Marco Bertamini. 2014. Examining visual complexity and its influence on perceived duration. Journal of Vision 14, 14 (2014), 3–3

  44. [51]

    Rik Pieters, Michel Wedel, and Rajeev Batra. 2010. The stopping power of advertising: Measures and effects of visual complexity. Journal of Marketing 74, 5 (2010), 48–60

  45. [52]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings...

  46. [53]

    Katharina Reinecke, Tom Yeh, Luke Miratrix, Rahmatri Mardiko, Yuechen Zhao, Jenny Liu, and Krzysztof Z Gajos. 2013. Predicting users’ first impressions of website aesthetics with a quantification of perceived visual complexity and colorfulness. In Proceedings of the SIGCHI Con...

  47. [55]

    Mathias Sable-Meyer, Kevin Ellis, Josh Tenenbaum, and Stanislas Dehaene. 2022. A language of thought for the mental representation of geometric shapes. Cognitive Psychology 139 (2022), 101527

  48. [56]

    Elham Saraee, Mona Jalal, and Margrit Betke. 2018. SAVOIAS: A Diverse, Multi-Category Visual Complexity Dataset. arXiv:1810.01771 [cs.CV] https://arxiv.org/abs/1810.01771

  49. [57]

    Elham Saraee, Mona Jalal, and Margrit Betke. 2020. Visual complexity analysis using deep intermediate-layer features. Computer Vision and Image Understanding 195 (2020), 102949. https://doi.org/10.1016/j.cviu.2020.102949

  50. [58]

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021. LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs. CoRR abs/2111.02114 (2021). arXiv:2111.02114 h...

  51. [59]

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge. In Computer Vision – ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23–27, 2022, Pr...

  52. [60]

    J. G. Snodgrass and M. Vanderwart. 1980. A Standardized Set of 260 Pictures: Norms for Name Agreement, Image Agreement, Familiarity, and Visual Complexity. Journal of Experimental Psychology: Human Learning and Memory 6, 2 (1980), 174–215

  53. [61]

    Peter X.-K. Song. 2007. Correlated data analysis : modeling, analytics, and applications. https://api.semanticscholar.org/CorpusID:117757997

  54. [62]

    Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. 2023. EVA-CLIP: Improved Training Techniques for CLIP at Scale. arXiv:2303.15389 [cs.CV] https://arxiv.org/abs/2303.15389

  55. [63]

    Zekun Sun and Chaz Firestone. 2021. Curious Objects: How Visual Complexity Guides Attention and Engagement. Cognitive science 45 4 (2021), e12933. https://api.semanticscholar.org/CorpusID:233309494

  56. [64]

    Yi Tang, Chia-Ming Chang, and Xi Yang. 2024. PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs. In Proceedings of the 29th International Conference on Intelligent User Interfaces (Greenville, SC, USA) (IUI ’24). Association fo...

  57. [65]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lil...

  58. [66]

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, Soroosh Mariooryad, Yifan Ding, Xinyang Geng, Fred Alcober, Roy Frostig, Mark Omernick, Lexi Walker, Cosmin Paduraru, Christina Sorokin, A...

  59. [67]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  60. [68]

    Max Wertheimer. 1923. Gestalt Principles in Visual Perception. Psychological Research 4 (1923), 301–350

  61. [69]

    Kewen Wu, Julita Vassileva, Yuxiang Zhao, Zeinab Noorian, Wesley Waldner, and Ifeoma Adaji. 2016. Complexity or simplicity? Designing product pictures for advertising in online marketplaces. Journal of Retailing and Consumer Services 28 (2016), 17–27

  62. [70]

    Zelinsky

    Chen-Ping Yu, Dimitris Samaras, and Gregory J. Zelinsky. 2014. Modeling visual clutter perception using proto-object segmentation.Journal of Vision 14, 7 (06 2014), 4–4. https://doi.org/10.1167/14.7.4 arXiv:https://arvojournals.org/arvo/content_public/journal/jov/933548/i1534-...

  63. [71]

    Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2018. From Recognition to Cognition: Visual Commonsense Reasoning. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2018), 6713–6724. https://api.semanticscholar.org/CorpusID:53734356

  64. [72]

    Ze Yu Zhang. 2024. Understanding the Relationship between Prompts and Response Uncertainty in Large Language Models. arXiv preprint arXiv:2407.14845 (2024). https://arxiv.org/abs/2407.14845

  65. [73]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Thirty-seventh Conference on Neural...

  66. [2023]

    https://doi.org/10.48550/arXiv.2305.18565

    PaLI-X: On Scaling up a Multilingual Vision and Language Model. https://doi.org/10.48550/arXiv.2305.18565

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.