{"id":"e23fa7c1-98fc-4ed8-bdc9-350c52f4a75a","arxiv_id":"2504.12511","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A multimodal LLM's pairwise judgments on visual clutter and Gestalt simplicity correlate most strongly with human visual complexity ratings across the SAVOIAS and IC9600 datasets.","lead":"The authors asked a commercial multimodal AI model to compare images on eight perceptual principles, from visual clutter and symmetry to the Gestalt laws, and matched the model's judgments against human complexity ratings from two public datasets. Clutter and simplicity consistently agreed with human ratings most strongly, which the authors interpret as a bias in how human annotators judged complexity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-bias conclusion rests on an unvalidated bridge: MLLM principle scores are never compared with human judgments of those same principles, so 'clutter' may just be the model's general complexity estimate.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing point: the paper validates MLLM outputs only against human complexity labels and never against human judgments of the eight principles, so the step from 'model clutter scores correlate with human complexity' to 'humans are biased toward clutter' is unsupported. I agree with the CONDITIONAL verdict: the descriptive result (one proprietary MLLM, on two datasets, produces clutter/simplicity scores that track human complexity labels) is internally consistent and is a reasonable benchmark observation, but it does not demonstrate human annotator bias. The concern is not a dispute with prior psychology literature; it is a missing control. Credit is due for the pairwise-comparison design, the use of public datasets, and the detailed per-category tables, but none of those supplies the missing human-principle validation. The one concrete check that would settle the matter is a human annotation study of the same principles. Because the reader already conditions acceptance on this kind of validation, I recommend no change to the verdict. I would not reject the paper: the framework and benchmark are useful, and the overclaim is localized to the interpretation in Section 5.","tokens_in":28936,"tokens_out":4679,"duration_ms":48929,"concrete_test":"Collect human pairwise judgments on all eight principles plus overall complexity for a stratified sample of image pairs from SAVOIAS and IC9600; compute agreement between human and Claude principle scores, and which human principles best predict human complexity labels. If Claude's clutter/simplicity judgments do not match human clutter/simplicity judgments, or if those human principles do not predict human complexity better than other principles, the bias claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 5) is that because Claude Sonnet 3.0's pairwise 'visual clutter' and 'law of simplicity' scores correlate most strongly with human complexity labels on SAVOIAS and IC9600, human annotators of those datasets are biased toward clutter and simplicity. For that inference to work, the MLLM's principle-specific judgments must (a) track the named psychological constructs the way human judgments do, and (b) be distinct from an overall complexity impression. Neither condition is tested. Section 4.2 converts binary comparisons into scores s_i, but the only external validation anywhere in the paper is against human complexity labels (Section 4.3, Tables 1 and 2). No human judgments of the eight principles are collected, and no agreement matrix among principles is reported. This matters because 'visual clutter' and 'simplicity' are semantically close to 'complexity' itself; the paper even cites [54] in defining complexity partly through clutter. If the model's 'clutter' axis is essentially its general complexity estimate, the high correlation is near-tautological and says nothing about human annotators' cognitive biases. The paper's own Section 6 concedes that prompt sensitivity is unexplored, which is the same weak point: different wording may collapse the eight axes into one. The result pattern is interesting, but the cognitive interpretation in Section 5 overreaches until the principle-judgment bridge is validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes an annotation-free framework for assessing whether multimodal LLMs can reason about visual complexity using eight explainable principles drawn from Gestalt psychology and visual perception: similarity, proximity, simplicity, closure, continuity, figure/ground, visual clutter, and visual symmetry. Claude Sonnet 3.0 performs pairwise comparisons of images under each principle; the binary results are aggregated into per-image scores (Section 4.2) and correlated with human complexity annotations on the SAVOIAS and IC9600 datasets across several image categories (Section 4.3, Tables 1-2). The paper reports that visual clutter and law of simplicity consistently yield the highest PLCC/SROCC values and interprets this pattern as evidence that human annotators of those datasets are biased toward clutter and simplicity while neglecting other psychological principles (Section 5). It also outlines HCI applications and future work (Section 6).","tokens_in":29026,"tokens_out":7283,"duration_ms":73202,"significance":"The framework is a useful low-cost benchmarking approach: it involves no fitted parameters, uses externally supplied human complexity labels, and makes the principle-level comparisons transparent. The consistent top ranking of clutter and simplicity across two datasets and most categories is an interesting empirical pattern that corroborates prior work on visual complexity and is worth reporting. The main load-bearing weakness is conceptual rather than computational: the inference from MLLM score-complexity correlations to human annotator bias requires validating that the MLLM's 'clutter' and 'simplicity' judgments track those specific constructs as humans would judge them, and that they are distinct from a general complexity impression. That validation is absent. The paper also ships no uncertainty quantification, which weakens the comparative claims.","major_comments":[{"comment":"The central claim that human annotators are biased toward visual clutter and simplicity is an interpretive bridge that is not tested. The scores s_i in Section 4.2 are derived from binary comparisons made by Claude Sonnet 3.0 'based on' each principle, but the only external validation in Section 4.3 is against human complexity labels. The paper neither collects human judgments of the same eight principles nor reports an agreement or discriminability matrix among the principle scores. Since visual clutter and simplicity are semantically close to complexity itself, and the paper cites [54] in connecting complexity to clutter, the high correlations in Tables 1-2 could be produced by a model that is essentially ranking images by overall complexity under different prompt phrasings; in that case the conclusion about human cognition would not follow. The paper's own Section 6 concession that prompt sensitivity is unexplored points to the same gap. A concrete remedy is to validate the principle judgments against human ratings of those principles and to include prompt-ablation controls demonstrating that the eight axes do not collapse into one.","section":"§5, with §4.2-4.3"},{"comment":"The reported PLCC and SROCC values are point estimates without sample sizes, confidence intervals, or significance tests. The statement in Section 5 that clutter and simplicity have 'highest correlation' is a ranking of noisy estimates; for example, in Table 2 the advertisement row separates Visual Clutter (0.76/0.76) from Law of Simplicity (0.70/0.76) by amounts that may be within sampling variability. The related claim that these measures are 'consistent across different categories' needs an explicit test or at least per-category uncertainty to be supported. Without this, the empirical foundation for the paper's strongest comparative conclusion is incomplete.","section":"Tables 1 and 2, §4.3/§5"},{"comment":"The aggregate scores rest on a single MLLM at temperature 0.01, with no repeated sampling, no reported agreement between runs, and no description of how ties or invalid comparison outputs are handled. Because the analysis compares small differences in correlations, the reliability of the pairwise-comparison step should be demonstrated. Additionally, conclusions about human annotators are drawn from one model; Section 4.4 explains why other tested MLLMs were not usable, but the human-bias claim should be explicitly framed as a property of this model unless corroborated by additional models or by direct human-principle ratings.","section":"§4.2 and §4.4"},{"comment":"The manuscript explicitly states that 'There is room for improvement in understanding the sensitivity of MLLM methods to the quality and design of the prompts.' This limitation is load-bearing for the Section 5 conclusion: if the ranking of principles changes under different prompt wording, the observed correlation pattern may reflect prompt design rather than a stable property of human annotations. The paper should either provide a prompt-sensitivity analysis or restrict the conclusion to the specific prompt protocol used.","section":"§6 (Applications & Future Work)"}],"minor_comments":[{"comment":"The list of Gestalt principles contains a duplicate: 'laws of similarity, proximity, simplicity, simplicity, closure, continuity and figure vs. ground.' The second 'simplicity' should be removed; the set should be six principles plus clutter and symmetry.","section":"§4"},{"comment":"The contribution states the framework is 'not susceptible to annotator biases,' while Section 5 argues that the human-annotated datasets are biased. This is not an outright contradiction if read as 'does not inherit training-set annotation bias,' but the wording should be clarified to avoid the apparent tension.","section":"§2 and §5"},{"comment":"There are repeated typos and inconsistent naming ('SAVIOAS' vs 'SAVOIAS', 'Explainabile AI', 'populalrly', 'reat time', 'interprete', 'monotocity'). A careful proofread is needed.","section":"Throughout"},{"comment":"The sample sizes for each category are not reported, and the figures and captions do not indicate uncertainty; the reader cannot judge how many images underlie each coefficient or whether the visual ordering of bars is statistically meaningful.","section":"Figure 5 and Tables 1-2"},{"comment":"The title uses 'Interpretable' to describe the framework, but the paper measures correlations between principle-based scores and human labels; this is better described as principle-aligned scoring rather than interpretability of the model's internal representations. Clarifying this terminology would make the contribution easier to position.","section":"Title and framing"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is workshop-paper depth, and its empirical pattern is worth publishing if reframed. The main risk is over-interpretation of MLLM correlations as human cognitive bias. I would ask the authors either to add direct human-principle validation and significance testing, or to soften the Section 5 claim to a statement about MLLM ranking behavior. The framework itself is not circular in the parameters sense, since no parameters are fitted and the labels are external, but the interpretive bridge from model scores to human cognition is untested."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core is a benchmark result: on SAVOIAS and IC9600, pairwise comparisons from one MLLM, guided by eight perceptual principles, give a consistent ranking of principle-to-complexity correlations, with visual clutter and law of simplicity on top across most categories. That is new and practically informative for anyone thinking about MLLMs as perception annotators.\n\nThe design is mostly clean. Pairwise comparisons are converted to scores, evaluated against external human complexity labels, no parameters are fitted, so the main correlations are not circular. The cross-category consistency strengthens the observation. The paper also openly reports that other MLLMs failed and that prompt sensitivity is unexplored.\n\nThe soft spot is the interpretation. Section 5 reads the correlation pattern as evidence that human annotators of SAVOIAS and IC9600 were biased toward clutter and simplicity. That step requires the MLLM's principle-specific judgments to track the named psychological constructs the way human judgments do. The paper never validates that. It compares the model's outputs only against human complexity labels, not against human judgments of the same eight principles. 'Visual clutter' and 'simplicity' are semantically close to 'complexity' itself, so the high correlation could be the model's general complexity impression under a principle-colored prompt. Section 6's admission that prompt sensitivity is unexplored points at the same gap. Missing sample sizes, confidence intervals, significance tests, and any baseline comparison to existing feature-based predictors further weaken the empirical claims. These gaps do not ruin the benchmark observation, but they do undercut the bias conclusion.\n\nWho should read it: people building interpretable complexity proxies for HCI or e-commerce will find the framework useful; people citing the human-bias claim should wait.\n\nMy recommendation: send to peer review. A serious referee should ask for human validation of the principle judgments, stated N and confidence intervals, and at least one feature-based baseline. With those, this becomes a solid benchmark paper.","headline":"Useful benchmark result with an overreaching cognitive claim; the principle-judgment bridge needs validation before the bias interpretation holds.","tokens_in":29711,"tokens_out":2556,"would_cite":false,"duration_ms":24805,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a multimodal large language model, prompted with named perceptual principles, can rank images by visual complexity without any human labels for training, and that the resulting comparisons expose a systematic bias…","keywords":["multimodal large language models","visual complexity","Gestalt principles","visual clutter","law of simplicity","pairwise comparison","human annotation bias","interpretability"],"falsifier":"Take a random sample of images from SAVOIAS and IC9600, collect human pairwise judgments on each of the eight principles, and compute an agreement matrix among principles and with human complexity ratings. If humans' own clutter ratings track their overall complexity ratings as strongly as the model's do, or if the model's eight scores collapse into a single general-complexity factor, then the claimed demonstration of annotator bias toward clutter and simplicity is unsupported.","tokens_in":28581,"feed_emoji":"👁","tokens_out":6710,"duration_ms":67380,"temperature":0.7,"pith_summary":"This paper proposes that a multimodal large language model, prompted with psychology's Gestalt principles, can serve as an annotation-free cognitive assistant for visual complexity analysis. The authors ask the model to compare pairs of images along eight interpretable dimensions—six Gestalt laws plus visual clutter and visual symmetry—and convert those pairwise judgments into per-image scores. Across the SAVOIAS and IC9600 datasets, scores for visual clutter and the law of simplicity correlate most strongly with human complexity annotations. The authors read this as evidence that human annotators in those datasets weight clutter and simplicity heavily and neglect other principles from the perception literature, and that a prompt-constrained MLLM can expose such biases without needing a training set.","feed_headline":"MLLM comparisons tie human complexity ratings to clutter and simplicity","feed_subtitle":"A multimodal LLM prompted with Gestalt principles matches human complexity labels best on clutter and simplicity.","key_machinery":"The load-bearing object is the pairwise comparison matrix $S=\\{s_{i,j}\\}$, where $s_{i,j}$ is the MLLM's binary judgment of which image in a pair better exemplifies an explainable parameter, aggregated into $\\hat{s}_i = \\frac{1}{n}\\sum_{j=1}^{n} s_{i,j}$. The parameters are six Gestalt principles (similarity, proximity, simplicity, closure, continuity, figure/ground) plus visual clutter and visual symmetry, defined in the prompt as they would be given to a human annotator. The matrix converts free-form MLLM reasoning into ranked scores per principle, which are then correlated with human complexity labels via Pearson and Spearman coefficients. Pairwise comparison is chosen over absolute ratings to avoid annotator scale bias and context limits, and the model is run at temperature 0.01 with requested justifications to stabilize outputs.","core_discovery":"The paper's central claim is that an MLLM can reason about visual complexity through named perceptual principles, and that doing so reveals a human bias. Questioned pairwise with definitions of eight principles, Claude Sonnet 3.0 produces rankings whose clutter and simplicity scores correlate with human complexity ratings at roughly 0.62–0.82 (PLCC/SROCC) across categories in SAVOIAS and IC9600, consistently higher than similarity, proximity, closure, continuity, figure/ground, or symmetry. The authors conclude that human annotators behind these datasets are biased toward visual clutter and visual simplicity, neglecting other reasonings proposed in psychology and cognitive science. They also report category-specific effects, such as law-of-closure correlations that are weak for advertisements but stronger for suprematism and paintings, and stable advertisement-category results across both datasets. The framework is presented as a scalable, annotation-free alternative to deep-learning complexity predictors, aimed at HCI tasks rather than at forecasting complexity scores.","pith_inferences":["The paper never validates the model's eight principle judgments against human judgments of those same eight principles; a direct experiment collecting per-principle human pairwise labels would separate the claim about human annotators from the claim about the model's own perceptual axis.","If future work finds that the model's clutter and simplicity scores are nearly collinear with its general-complexity score, the 'bias' result reduces to the weaker statement that one complexity factor drives both; that would be a confound the current correlation table cannot rule out.","Because prompt sensitivity is explicitly left unexamined, the quantitative rankings plausibly depend on the exact wording and model version; re-running with prompts that instruct the model to weigh symmetry or closure deliberately would test whether the clutter/simplicity dominance is stable.","A natural extension is to apply the same pairwise protocol to other subjective annotations—aesthetic appeal, trust, cognitive load—to ask whether clutter and simplicity dominate human judgment there as well, or whether the bias is specific to complexity."],"forward_implications":["If the bias claim holds, human-annotated visual complexity datasets should not be treated as measuring general complexity; they largely measure clutter and simplicity.","The same annotation-free protocol can be reused on new image categories without retraining, because the MLLM is constrained by prompt definitions rather than by dataset labels.","Practical HCI applications follow directly: designers and content creators can receive explainable per-principle feedback on visual balance, clutter, and symmetry, and platforms can adjust search-result presentation toward principles that reduce cognitive load.","Category-level differences imply that an interface tuned for one domain (e.g., advertisements) may not transfer to another (e.g., paintings), because the relevant perceptual principle changes.","The finding suggests that datasets like SAVOIAS and IC9600 carry systematic annotator bias that downstream models trained on them will inherit."],"supporting_citations":[{"why":"SAVOIAS dataset whose human complexity labels ground the correlation analysis.","marker":"[56]"},{"why":"IC9600 dataset supplying the second benchmark of human complexity labels.","marker":"[25]"},{"why":"Claude Sonnet 3.0, the multimodal model that performs the pairwise perceptual comparisons.","marker":"[17]"},{"why":"Source of the Gestalt principles used as the named perceptual laws in the prompt.","marker":"[68]"},{"why":"Provides the notion of visual clutter used to define one of the eight comparison parameters.","marker":"[54]"},{"why":"Prior study on visual clutter perception cited as corroborating evidence that clutter shapes human judgment.","marker":"[70]"},{"why":"Prior statement of the simplicity principle in perception, cited as supporting the second strongest correlate.","marker":"[24]"}],"fun_headline_variants":["MLLM reasoning uncovers human bias toward clutter and simplicity","LLM with Gestalt principles exposes clutter bias in complexity ratings","Clutter and simplicity drive human complexity judgments, LLM reveals","Multimodal LLM spots bias: Visual complexity hinges on clutter, simplicity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the MLLM's 'clutter' and 'simplicity' pairwise judgments measure those named principles the way a human would and are not just the model's overall complexity impression; if that bridge fails, the correlation cannot be read as evidence about human annotator bias.","fun_headline_variants_meta":{"raw":{"variants":["MLLM reasoning uncovers human bias toward clutter and simplicity","LLM with Gestalt principles exposes clutter bias in complexity ratings","Clutter and simplicity drive human complexity judgments, LLM reveals","Multimodal LLM spots bias: Visual complexity hinges on clutter, simplicity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000527,"raw_usage":{"total_tokens":2539,"prompt_tokens":937,"completion_tokens":1602,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":1528}},"tokens_in":553,"tokens_out":1602,"duration_ms":13096,"temperature":1.0,"reasoning_tokens":1528,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:30:29.307110+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of images from SAVOIAS and IC9600, collect human pairwise judgments on each of the eight principles, and compute an agreement matrix among principles and with human complexity ratings. If humans' own clutter ratings track their overall complexity ratings as strongly as the model's do, or if the model's eight scores collapse into a single general-complexity factor, then the claimed demonstration of annotator bias toward clutter and simplicity is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"IC9600 dataset supplying the second benchmark of human complexity labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Claude Sonnet 3.0, the multimodal model that performs the pairwise perceptual comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the Gestalt principles used as the named perceptual laws in the prompt."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the notion of visual clutter used to define one of the eight comparison parameters."},{"cited_title":"Zelinsky","cited_arxiv_id":null,"evidence_quote":"Prior study on visual clutter perception cited as corroborating evidence that clutter shapes human judgment."}],"review_version":1}