Pith. sign in

REVIEW 1 major objections 3 minor 41 references

More Accurate, Less Human: Gestalt Grouping in Vision Models

T0 review · 1 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Accuracy does not buy human-like grouping in 45 vision models

desk verdict A genuinely reusable Gestalt battery with a real silhouette dissociation, held back by an overclaimed series-counting result and no released artifacts. read the letter →

arxiv 2608.10195 v1 pith:2FG7XWCN submitted 2026-08-10 cs.CV cs.LG

classification cs.CVcs.LG
keywords GraphicalperceptionGestaltgroupingvisionsciencepsychophysicsbehavioralevaluationerrorconsistencymodelshuman–modelalignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a vision model's benchmark accuracy says little about whether it organizes visual input the way humans do, and that agreement with published human perception data is a measurable, reusable yardstick for that difference. Across 45 models on four Gestalt grouping tasks, accuracy and human-likeness dissociate: 34 of 45 models recognize silhouettes more accurately than human observers, yet only four match human–human error consistency, and the two measures rank models discordantly in 37% of pairs (ρ=0.38). The authors build a behavioral battery from existing perception studies—silhouette recognition, mark-color odd-one-out, color-series counting, and object odd-one-out—so no new user study is needed. The upshot for visualization research is concrete: models entering chart-reading pipelines cannot be assumed to see marks and objects the way their human audience does, and which model looks human-like depends on which grouping task is measured.

What carries the argument

The carrying object is the behavioral battery itself: four forced-choice tasks reusing stimuli and human data from published perception studies, with two natural-image control tasks (silhouette closure, object odd-one-out) and two chart-track tasks (mark-color odd-one-out, color-series counting). Three metrics do the scoring: Behavioral Agreement B (how often the model's response equals the human reference), trial-level Behavioral Error Consistency κ (a chance-corrected covariance between model correctness and per-item human accuracy, scored against the human–human ceiling), and Behavioral Effect Replication a(k) (whether accuracy as a function of series count bends where human ensemble capacity does). The battery makes encoders and foundation models directly comparable by reducing every output to one scored response, and it anchors each score to source-study floors and ceilings rather than to thresholds chosen by the authors.

What would settle it

Run a fresh psychophysics experiment on the identical color-series counting stimuli, collecting per-trial human responses; if human accuracy by series count does not show a capacity-limited decline or plateau across the 6–12 range, then the claimed failure of Behavioral Effect Replication in every model family is not measurable against the human reference the paper assumes.

Watch

Extended reading notes

Core claim

The central claim is that human-likeness in perceptual organization is a distinct, task-specific property from task performance, and that it can be measured behaviorally by scoring models against published human responses. On their battery, human agreement B and error consistency κ against a human–human ceiling reveal dissociations invisible to accuracy: the best color-similarity encoder is mid-pack on semantic odd-one-out, the best semantic encoder is unremarkable on color, and open-weight VLMs express little of the grouping their encoder backbones carry. At the frontier, however, a few closed foundation models pair above-human accuracy with error consistency at or beyond the human–human ceiling, showing that human-like grouping is achievable, but it is earned by training recipe rather than by scale or accuracy.

Load-bearing premise

The load-bearing premise is that the published human data remain valid for the exact stimuli shown to the models, most fragilely the 6–12 color-series capacity band, which rests on qualitative findings with no per-trial human responses; if that band is not the right signature of human ensemble segmentation, the Behavioral Effect Replication result loses its anchor.

Editorial extensions

If this is right

  • Benchmark accuracy alone cannot certify a model as a human proxy for design evaluation; a task-indexed behavioral audit is needed.
  • A model that groups chart marks human-likely on one task can diverge sharply on another, so human-perception guidelines do not automatically transfer to machine readers.
  • Open-weight VLMs can lose the perceptual organization of their encoder towers, so the instruction-tuning stage itself may be reshaping grouping behavior.
  • The best closed foundation models achieve human-level error consistency on silhouettes, showing the dissociation is not an inherent limit of current architectures.
  • Published psychophysics data can be reused as a no-new-user-study yardstick, and the same construction extends to other Gestalt laws and chart-context stimuli.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether the language-alignment stage actively destroys perceptual organization; a testable follow-up would probe intermediate encoder layers of the open VLMs to see where grouping information is lost.
  • If the dissociation generalizes to complete charts with axes and legends, visualization guidelines validated on human viewers may need model-specific versions, not just a single human-likeness score.
  • The rank discordance between accuracy and κ suggests human-aligned grouping could be optimized as its own training target, with the prediction that such models would score higher on this battery without necessarily changing benchmark accuracy.
  • The color-series capacity band is the weakest anchor; a fresh per-trial human study on the exact stimuli would tighten or refute the Behavioral Effect Replication result.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 3 minor

Summary. The paper introduces a behavioral battery that scores vision models against published human psychophysics on four Gestalt tasks: silhouette recognition (closure), mark-color odd-one-out, color-series counting, and object odd-one-out (similarity). It evaluates 15 encoders and 30 foundation models, computing Behavioral Agreement (B), trial-level Behavioral Error Consistency (kappa), and Behavioral Effect Replication (a(k)). The central finding is that silhouette accuracy and human error consistency dissociate: many models exceed human accuracy while failing to share human error patterns, and only a few closed frontier models approach or exceed the human-human consistency ceiling. The authors conclude that human-likeness is task-specific, not a global trait, and propose the battery as a reusable, user-study-free yardstick for auditing models in visualization pipelines. The manuscript includes reproducibility details, per-trial records, bootstrap CIs, and a validation of the kappa ceiling, making the core dissociation claim transparent and testable.

Significance. If the core dissociation result holds, this is a valuable contribution to model evaluation in visualization and beyond. The paper introduces a general methodology for reusing published perception data as evaluation targets, which is cheap and principled. The silhouette dissociation is supported by per-trial human data, a chance-corrected kappa with a validated ceiling, and bootstrap confidence intervals. The manuscript also ships reproducible stimulus generation, a schema-checked record format, and per-trial records for every model, which is exemplary practice. The main weakness is the series-counting Behavioral Effect Replication claim, which rests on a qualitative human capacity band rather than a per-trial human curve; this specific claim is unfalsifiable as stated. The remaining evidence for task-specific human-likeness is credible, and the methodological contribution is significant.

major comments (1)
  1. [Sec. 4; Table 1; Appendix D.1] The claim in Sec. 4 that "Behavioral Effect Replication fails in every family" is not supported by the evidence presented. The human anchor for color-series counting is a "derived 6–12 capacity band" (Table 1, Sec. 3.1), which is a qualitative range from other studies, not a per-count human accuracy curve. Sec. 3.3 states that replication means "the profile bends where human capacity does," but the paper never defines a test statistic, a null distribution, or a criterion for deciding that a model's a(k) is "unconstrained by the band." Moreover, Appendix D.1 explicitly limits series counting to "capacity-band membership" and reports a bootstrap CI half-width of 0.155 for n=40, which is inconsistent with the strong, cross-family failure asserted in Sec. 4. Without a defined human a(k) profile, the claim is unfalsifiable. This is load-bearing because it is used as evidence for the task-specificity of human-likeness, so it should be removed or downgraded to a descriptive observation until a per-trial human reference or a clearly specified criterion is available.
minor comments (3)
  1. [Abstract; Sec. 4] The phrase "benchmark accuracy" is ambiguous. The only accuracy measure used in the accuracy-vs-human-likeness dissociation is within-battery silhouette class accuracy, not a general-purpose benchmark such as ImageNet or chart-QA accuracy. Please replace "benchmark accuracy" with "silhouette accuracy" or "within-battery accuracy" to avoid overstatement.
  2. [Fig. 1 legend; Sec. 1; Abstract] The contrastive vision-language family is denoted with a bullet "•" in the abstract and text, but with a filled circle "●" in the Fig. 1 legend. Please use one symbol consistently across the paper.
  3. [Sec. 3.3] The metric a(k) is described verbally but never defined with an equation, and the paper does not show a plot or table of accuracy versus series count k. Adding an explicit definition (even a simple one) and a supplementary figure would allow readers to assess the claimed threefold spread across encoder objectives.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: all human targets are external published data and all probes are fit on disjoint splits.

full rationale

The battery's targets are human data from prior perception studies (Geirhos silhouettes, Demiralp perceptual kernels, THINGS odd-one-out, and the published 6-12 capacity band), none produced by the authors or fitted to the model outputs. B, kappa, and a(k) are computed against fixed external references; no equation in the paper defines a predicted quantity in terms of those references. The only probe parameters (class prototypes, ridge regression readout) are fit on disjoint probe-train splits and applied frozen to test splits, so they are not fitted inputs renamed as predictions. The paper explicitly notes that for odd-one-out tasks B collapses to accuracy because the ground truth is the human choice; that is an acknowledged definitional identity, not a hidden circularity, and the dissociation claims rest on the silhouette task where ground truth and human target are independent. The series-counting 'a(k) unconstrained by the 6-12 band' claim is weakly specified—Appendix D.1 limits it to capacity-band membership and no test statistic is given—but the band is an external qualitative human reference, so this is a correctness/evidentiary weakness, not a circular reduction. No load-bearing self-citation occurs; cited prior work is independent and mostly external to the authors.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical or mathematical entities. Its free parameters are standard readout probes fit on disjoint training splits; the core assumptions are about the validity of reusing published human data as ground truth and about the faithfulness of readouts and text parsing.

free parameters (2)
  • Silhouette class prototypes
    Fitted on probe-train split of Geirhos silhouettes to read encoder features; applied frozen to the test split. Standard linear readout, not a parameter of the theoretical claim.
  • Series-counting ridge regression weights
    Fitted on disjoint probe-train split to map encoder features to rendered counts; used only for the series counting task.
assumptions (5)
  • domain assumption Published human datasets (Geirhos silhouettes, Demiralp kernels, THINGS triplets) are valid, representative measures of human Gestalt grouping and can serve as ground truth for model comparison.
    Entire battery is built on reusing these datasets; if the human data are not valid for the rendered stimuli, all B/kappa/a(k) comparisons lose their reference.
  • domain assumption The re-rendered chart stimuli match the source studies' stimuli closely enough that the published human responses remain applicable.
    Sec. 3.1 'Construction'; exact stimulus matching is what licenses error-level scoring.
  • domain assumption Linear probes on frozen encoder features estimate what is linearly decodable from the representation and cannot introduce organization that is absent.
    Sec. 3.2 'Readouts'; this assumption underpins using probes for silhouettes and series counting.
  • domain assumption Rule-based parsing of free-text answers and greedy/zero-temperature decoding yield a deterministic, meaningful single response per foundation-model stimulus.
    Sec. 3.2 'Protocol' and Appendix D.4; unparseable answers are excluded, which assumes missing responses are ignorable.
  • standard math Cohen's kappa and the displayed covariance identity give a valid chance-corrected measure of trial-level error consistency.
    Appendix D.2 derives kappa = 2Cov/(1-c_exp).

how reviews work

0 comments
Cite this review

Pith. "Pith review of More Accurate, Less Human: Gestalt Grouping in Vision Models." pith.science (2026). https://pith.science/paper/2FG7XWCN

@misc{pith2026260810195,
  author       = {Pith},
  title        = {Pith review of: More Accurate, Less Human: Gestalt Grouping in Vision Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2FG7XWCN}},
  note         = {Machine review of arXiv:2608.10195}
}
read the original abstract

Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.

Figures

Figures reproduced from arXiv: 2608.10195 by the authors.

Figure 1
Figure 1. Human-likeness B on the four battery tasks for all 45 models (tasks in rows, models in columns), grouped by family and ordered within each family by model size: parameter count for encoders and open VLMs, vendor and tier for ⋆ closed models, whose sizes are undisclosed. mance [9]. Several of these studies published raw per-trial data; our battery is built directly on those releases. Model–human comparison. Behaviora… view at source ↗
Figure 2
Figure 2. The behavioral battery. Each stimulus (left) is answered by every model and by human observers (center). The paired responses are then scored by three battery metrics (right) [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Accuracy and human-likeness are different properties. One [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

41 extracted references · 36 canonical work pages

  1. [1]

    Bendeck and J

    A. Bendeck and J. Stasko. An empirical evaluation of the GPT-4 mul- timodal language model on visualization literacy tasks.IEEE Trans- actions on Visualization and Computer Graphics, 31(1):1105–1115,

  2. [2]

    Binz and E

    M. Binz and E. Schulz. Using cognitive psychology to under- stand GPT-3.Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023. 2

  3. [3]

    W. S. Cleveland and R. McGill. Graphical perception: Theory, ex- perimentation, and application to the development of graphical meth- ods.Journal of the American Statistical Association, 79(387):531– 554, 1984. 1

  4. [4]

    J. Cohen. A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960. 1, 7

  5. [5]

    Demiralp, M

    C ¸ . Demiralp, M. S. Bernstein, and J. Heer. Learning perceptual ker- nels for visualization design.IEEE Transactions on Visualization and Computer Graphics, 20(12):1933–1942, 2014. 1, 2, 6

  6. [6]

    Geirhos, K

    R. Geirhos, K. Meding, and F. A. Wichmann. Beyond accuracy: quan- tifying trial-by-trial behaviour of CNNs and humans by measuring er- ror consistency. InProc. NeurIPS, 2020. 1, 2, 3, 7

  7. [7]

    Geirhos, K

    R. Geirhos, K. Narayanappa, B. Mitzkus, T. Thieringer, M. Bethge, F. A. Wichmann, and W. Brendel. Partial success in closing the gap between human and machine vision. InProc. NeurIPS, 2021. 2

  8. [8]

    Geirhos, P

    R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel. ImageNet-trained CNNs are biased towards texture; in- creasing shape bias improves accuracy and robustness. InProc. ICLR,

Show all 41 references
  1. [9]

    C. C. Gramazio, K. B. Schloss, and D. H. Laidlaw. The relation between visualization size, grouping, and user performance.IEEE Transactions on Visualization and Computer Graphics, 20(12):1953– 1962, 2014. 2

  2. [10]

    Haehn, J

    D. Haehn, J. Tompkin, and H. Pfister. Evaluating ‘graphical percep- tion’ with CNNs.IEEE Transactions on Visualization and Computer Graphics, 25(1):641–650, 2019. 2

  3. [11]

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick. Masked autoencoders are scalable vision learners. InProc. CVPR, 2022. 8

  4. [12]

    C. G. Healey. Choosing effective colours for data visualization.Proc. IEEE Visualization, pp. 263–270, 1996. 1, 2, 6

  5. [13]

    M. N. Hebart, O. Contier, L. Teichmann, A. H. Rockter, C. Y . Zheng, A. Kidder, A. Corriveau, M. Vaziri-Pashkam, and C. I. Baker. THINGS-data: A multimodal collection of large-scale datasets for in- vestigating object representations in human brain and behavior.eLife, 12:e8258...

  6. [14]

    M. N. Hebart, C. Y . Zheng, F. Pereira, and C. I. Baker. Revealing the multidimensional mental representations of natural objects underly- ing human similarity judgements.Nature Human Behaviour, 4:1173– 1185, 2020. 2, 3, 6

  7. [15]

    Heer and M

    J. Heer and M. Bostock. Crowdsourcing graphical perception: Using mechanical turk to assess visualization design. InProc. ACM CHI, pp. 203–212, 2010. 1

  8. [16]

    T. Li, Z. Wen, L. Song, J. Liu, Z. Jing, and T. S. Lee. From local cues to global percepts: Emergent gestalt organization in self-supervised vision models.arXiv preprint arXiv:2506.00718, 2025. 2

  9. [17]

    H. Liu, C. Li, Y . Li, and Y . J. Lee. Improved baselines with visual instruction tuning. InProc. CVPR, 2024. 8

  10. [18]

    Marafioti, O

    A. Marafioti, O. Zohar, M. Farr ´e, et al. SmolVLM: Redefining small and efficient multimodal models.arXiv:2504.05299, 2025. 8

  11. [19]

    Masry, D

    A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguis- tics: ACL 2022, pp. 2263–2279, 2022. 1

  12. [20]

    C. M. McColeman, F. Yang, T. F. Brady, and S. Franconeri. Rethink- ing the ranks of visual channels.IEEE Transactions on Visualization and Computer Graphics, 28(1):707–717, 2022. 1

  13. [21]

    Muttenthaler, J

    L. Muttenthaler, J. Dippel, L. Linhardt, R. A. Vandermeulen, and S. Kornblith. Human alignment of neural network representations. InProc. ICLR, 2023. 2, 3

  14. [22]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, et al. DINOv2: Learning robust visual features without supervision.Transactions on Machine Learn- ing Research, 2024. 8

  15. [23]

    Pandey and A

    S. Pandey and A. Ottley. Benchmarking visual language models on standardized visualization literacy tests.Computer Graphics Forum, 44(3):e70137, 2025. 1, 2

  16. [24]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervi- sion. InProc. ICML, 2021. 8

  17. [25]

    Schrimpf, J

    M. Schrimpf, J. Kubilius, M. J. Lee, N. A. R. Murty, R. Ajemian, and J. J. DiCarlo. Integrative benchmarking to advance neurally mecha- nistic models of human intelligence.Neuron, 108(3):413–423, 2020. 2

  18. [26]

    J. Sha, H. Shindo, K. Kersting, and D. S. Dhami. Gestalt vision: A dataset for evaluating gestalt principles in visual perception. In Proceedings of the 19th International Conference on Neurosymbolic Learning and Reasoning, vol. 284 ofPMLR, pp. 873–890, 2025. 2

  19. [27]

    Skau and R

    D. Skau and R. Kosara. Arcs, angles, or areas: Individual data encod- ings in pie and donut charts.Computer Graphics Forum, 35(3):121– 130, 2016. 1

  20. [28]

    Sucholutsky, L

    I. Sucholutsky, L. Muttenthaler, A. Weller, A. Peng, A. Bobu, B. Kim, B. C. Love, E. Grant, I. Groen, J. Achterberg, J. B. Tenenbaum, et al. Getting aligned on representational alignment.arXiv preprint arXiv:2310.13018, 2023. 2

  21. [29]

    D. A. Szafir. Modeling color difference for visualization design.IEEE Transactions on Visualization and Computer Graphics, 24(1):392– 401, 2018. 1, 2, 6

  22. [30]

    Talbot, V

    J. Talbot, V . Setlur, and A. Anand. Four experiments on the percep- tion of bar charts.IEEE Transactions on Visualization and Computer Graphics, 20(12):2152–2160, 2014. 1

  23. [31]

    Torfs, K

    K. Torfs, K. Vancleef, C. Lafosse, J. Wagemans, and L. de-Wit. The Leuven Perceptual Organization Screening Test (L-POST), an online test to assess mid-level visual perception.Behavior Research Meth- ods, 46(2):472–487, 2014. 2

  24. [32]

    Tseng, A

    C. Tseng, A. Z. Wang, G. J. Quadri, and D. A. Szafir. Revisiting categorical color perception in scatterplots: Sequential, diverging, and categorical palettes. InEuroVis Short Papers, 2024. 1

  25. [33]

    Verma, K

    A. Verma, K. Mukherjee, C. Potts, E. Kreiss, and J. E. Fan. CHART- 6: Human-centered evaluation of data visualization understanding in vision-language models.arXiv preprint arXiv:2505.17202, 2025. 2

  26. [34]

    P. Wang, S. Bai, S. Tan, et al. Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution.arXiv:2409.12191,

  27. [35]

    Ware.Information Visualization: Perception for Design

    C. Ware.Information Visualization: Perception for Design. Morgan Kaufmann, 3rd ed., 2013. 1, 2, 6

  28. [36]

    Wertheimer

    M. Wertheimer. Untersuchungen zur Lehre von der Gestalt. II.Psy- chologische Forschung, 4:301–350, 1923. 1

  29. [37]

    Whitney and A

    D. Whitney and A. Yamanashi Leib. Ensemble perception.Annual Review of Psychology, 69:105–129, 2018. 2

  30. [38]

    X. Yue, Y . Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, et al. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. InProc. IEEE/CVF CVPR,

  31. [39]

    X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre-training. InProc. ICCV, 2023. 8

  32. [40]

    Zhang, J

    K. Zhang, J. Yang, J. P. Inala, C. Singh, J. Gao, Y . Su, and C. Wang. Towards understanding graphical perception in large multimodal mod- els.arXiv preprint arXiv:2503.10857, 2025. 2

  33. [41]

    trial_id

    Y . Zhang, J. Pan, Y . Zhou, R. Pan, and J. Chai. Grounding visual illusions in language: Do vision-language models perceive illusions like humans? InProc. EMNLP, pp. 5718–5728, 2023. 1 APPENDIX A EXAMPLE STIMULI Every stimulus below is the exact image presented to the foundat...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.