REVIEW 1 major objections 3 minor 41 references
More Accurate, Less Human: Gestalt Grouping in Vision Models
T0 review · 1 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Accuracy does not buy human-like grouping in 45 vision models
desk verdict A genuinely reusable Gestalt battery with a real silhouette dissociation, held back by an overclaimed series-counting result and no released artifacts. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the behavioral battery itself: four forced-choice tasks reusing stimuli and human data from published perception studies, with two natural-image control tasks (silhouette closure, object odd-one-out) and two chart-track tasks (mark-color odd-one-out, color-series counting). Three metrics do the scoring: Behavioral Agreement B (how often the model's response equals the human reference), trial-level Behavioral Error Consistency κ (a chance-corrected covariance between model correctness and per-item human accuracy, scored against the human–human ceiling), and Behavioral Effect Replication a(k) (whether accuracy as a function of series count bends where human ensemble capacity does). The battery makes encoders and foundation models directly comparable by reducing every output to one scored response, and it anchors each score to source-study floors and ceilings rather than to thresholds chosen by the authors.
What would settle it
Run a fresh psychophysics experiment on the identical color-series counting stimuli, collecting per-trial human responses; if human accuracy by series count does not show a capacity-limited decline or plateau across the 6–12 range, then the claimed failure of Behavioral Effect Replication in every model family is not measurable against the human reference the paper assumes.
Extended reading notes
Core claim
The central claim is that human-likeness in perceptual organization is a distinct, task-specific property from task performance, and that it can be measured behaviorally by scoring models against published human responses. On their battery, human agreement B and error consistency κ against a human–human ceiling reveal dissociations invisible to accuracy: the best color-similarity encoder is mid-pack on semantic odd-one-out, the best semantic encoder is unremarkable on color, and open-weight VLMs express little of the grouping their encoder backbones carry. At the frontier, however, a few closed foundation models pair above-human accuracy with error consistency at or beyond the human–human ceiling, showing that human-like grouping is achievable, but it is earned by training recipe rather than by scale or accuracy.
Load-bearing premise
The load-bearing premise is that the published human data remain valid for the exact stimuli shown to the models, most fragilely the 6–12 color-series capacity band, which rests on qualitative findings with no per-trial human responses; if that band is not the right signature of human ensemble segmentation, the Behavioral Effect Replication result loses its anchor.
Editorial extensions
If this is right
- Benchmark accuracy alone cannot certify a model as a human proxy for design evaluation; a task-indexed behavioral audit is needed.
- A model that groups chart marks human-likely on one task can diverge sharply on another, so human-perception guidelines do not automatically transfer to machine readers.
- Open-weight VLMs can lose the perceptual organization of their encoder towers, so the instruction-tuning stage itself may be reshaping grouping behavior.
- The best closed foundation models achieve human-level error consistency on silhouettes, showing the dissociation is not an inherent limit of current architectures.
- Published psychophysics data can be reused as a no-new-user-study yardstick, and the same construction extends to other Gestalt laws and chart-context stimuli.
Reading between the lines
- The paper leaves open whether the language-alignment stage actively destroys perceptual organization; a testable follow-up would probe intermediate encoder layers of the open VLMs to see where grouping information is lost.
- If the dissociation generalizes to complete charts with axes and legends, visualization guidelines validated on human viewers may need model-specific versions, not just a single human-likeness score.
- The rank discordance between accuracy and κ suggests human-aligned grouping could be optimized as its own training target, with the prediction that such models would score higher on this battery without necessarily changing benchmark accuracy.
- The color-series capacity band is the weakest anchor; a fresh per-trial human study on the exact stimuli would tighten or refute the Behavioral Effect Replication result.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a behavioral battery that scores vision models against published human psychophysics on four Gestalt tasks: silhouette recognition (closure), mark-color odd-one-out, color-series counting, and object odd-one-out (similarity). It evaluates 15 encoders and 30 foundation models, computing Behavioral Agreement (B), trial-level Behavioral Error Consistency (kappa), and Behavioral Effect Replication (a(k)). The central finding is that silhouette accuracy and human error consistency dissociate: many models exceed human accuracy while failing to share human error patterns, and only a few closed frontier models approach or exceed the human-human consistency ceiling. The authors conclude that human-likeness is task-specific, not a global trait, and propose the battery as a reusable, user-study-free yardstick for auditing models in visualization pipelines. The manuscript includes reproducibility details, per-trial records, bootstrap CIs, and a validation of the kappa ceiling, making the core dissociation claim transparent and testable.
Significance. If the core dissociation result holds, this is a valuable contribution to model evaluation in visualization and beyond. The paper introduces a general methodology for reusing published perception data as evaluation targets, which is cheap and principled. The silhouette dissociation is supported by per-trial human data, a chance-corrected kappa with a validated ceiling, and bootstrap confidence intervals. The manuscript also ships reproducible stimulus generation, a schema-checked record format, and per-trial records for every model, which is exemplary practice. The main weakness is the series-counting Behavioral Effect Replication claim, which rests on a qualitative human capacity band rather than a per-trial human curve; this specific claim is unfalsifiable as stated. The remaining evidence for task-specific human-likeness is credible, and the methodological contribution is significant.
major comments (1)
- [Sec. 4; Table 1; Appendix D.1] The claim in Sec. 4 that "Behavioral Effect Replication fails in every family" is not supported by the evidence presented. The human anchor for color-series counting is a "derived 6–12 capacity band" (Table 1, Sec. 3.1), which is a qualitative range from other studies, not a per-count human accuracy curve. Sec. 3.3 states that replication means "the profile bends where human capacity does," but the paper never defines a test statistic, a null distribution, or a criterion for deciding that a model's a(k) is "unconstrained by the band." Moreover, Appendix D.1 explicitly limits series counting to "capacity-band membership" and reports a bootstrap CI half-width of 0.155 for n=40, which is inconsistent with the strong, cross-family failure asserted in Sec. 4. Without a defined human a(k) profile, the claim is unfalsifiable. This is load-bearing because it is used as evidence for the task-specificity of human-likeness, so it should be removed or downgraded to a descriptive observation until a per-trial human reference or a clearly specified criterion is available.
minor comments (3)
- [Abstract; Sec. 4] The phrase "benchmark accuracy" is ambiguous. The only accuracy measure used in the accuracy-vs-human-likeness dissociation is within-battery silhouette class accuracy, not a general-purpose benchmark such as ImageNet or chart-QA accuracy. Please replace "benchmark accuracy" with "silhouette accuracy" or "within-battery accuracy" to avoid overstatement.
- [Fig. 1 legend; Sec. 1; Abstract] The contrastive vision-language family is denoted with a bullet "•" in the abstract and text, but with a filled circle "●" in the Fig. 1 legend. Please use one symbol consistently across the paper.
- [Sec. 3.3] The metric a(k) is described verbally but never defined with an equation, and the paper does not show a plot or table of accuracy versus series count k. Adding an explicit definition (even a simple one) and a supplementary figure would allow readers to assess the claimed threefold spread across encoder objectives.
Circularity Check
No significant circularity: all human targets are external published data and all probes are fit on disjoint splits.
full rationale
The battery's targets are human data from prior perception studies (Geirhos silhouettes, Demiralp perceptual kernels, THINGS odd-one-out, and the published 6-12 capacity band), none produced by the authors or fitted to the model outputs. B, kappa, and a(k) are computed against fixed external references; no equation in the paper defines a predicted quantity in terms of those references. The only probe parameters (class prototypes, ridge regression readout) are fit on disjoint probe-train splits and applied frozen to test splits, so they are not fitted inputs renamed as predictions. The paper explicitly notes that for odd-one-out tasks B collapses to accuracy because the ground truth is the human choice; that is an acknowledged definitional identity, not a hidden circularity, and the dissociation claims rest on the silhouette task where ground truth and human target are independent. The series-counting 'a(k) unconstrained by the 6-12 band' claim is weakly specified—Appendix D.1 limits it to capacity-band membership and no test statistic is given—but the band is an external qualitative human reference, so this is a correctness/evidentiary weakness, not a circular reduction. No load-bearing self-citation occurs; cited prior work is independent and mostly external to the authors.
Assumptions & free parameters
free parameters (2)
- Silhouette class prototypes
- Series-counting ridge regression weights
assumptions (5)
- domain assumption Published human datasets (Geirhos silhouettes, Demiralp kernels, THINGS triplets) are valid, representative measures of human Gestalt grouping and can serve as ground truth for model comparison.
- domain assumption The re-rendered chart stimuli match the source studies' stimuli closely enough that the published human responses remain applicable.
- domain assumption Linear probes on frozen encoder features estimate what is linearly decodable from the representation and cannot introduce organization that is absent.
- domain assumption Rule-based parsing of free-text answers and greedy/zero-temperature decoding yield a deterministic, meaningful single response per foundation-model stimulus.
- standard math Cohen's kappa and the displayed covariance identity give a valid chance-corrected measure of trial-level error consistency.
Cite this review
Pith. "Pith review of More Accurate, Less Human: Gestalt Grouping in Vision Models." pith.science (2026). https://pith.science/paper/2FG7XWCN
@misc{pith2026260810195,
author = {Pith},
title = {Pith review of: More Accurate, Less Human: Gestalt Grouping in Vision Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/2FG7XWCN}},
note = {Machine review of arXiv:2608.10195}
}
read the original abstract
Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Bendeck and J. Stasko. An empirical evaluation of the GPT-4 mul- timodal language model on visualization literacy tasks.IEEE Trans- actions on Visualization and Computer Graphics, 31(1):1105–1115,
-
[2]
M. Binz and E. Schulz. Using cognitive psychology to under- stand GPT-3.Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023. 2
work page 2023
-
[3]
W. S. Cleveland and R. McGill. Graphical perception: Theory, ex- perimentation, and application to the development of graphical meth- ods.Journal of the American Statistical Association, 79(387):531– 554, 1984. 1
work page 1984
-
[4]
J. Cohen. A coefficient of agreement for nominal scales.Educational and Psychological Measurement, 20(1):37–46, 1960. 1, 7
work page 1960
-
[5]
C ¸ . Demiralp, M. S. Bernstein, and J. Heer. Learning perceptual ker- nels for visualization design.IEEE Transactions on Visualization and Computer Graphics, 20(12):1933–1942, 2014. 1, 2, 6
work page 1933
-
[6]
R. Geirhos, K. Meding, and F. A. Wichmann. Beyond accuracy: quan- tifying trial-by-trial behaviour of CNNs and humans by measuring er- ror consistency. InProc. NeurIPS, 2020. 1, 2, 3, 7
work page 2020
-
[7]
R. Geirhos, K. Narayanappa, B. Mitzkus, T. Thieringer, M. Bethge, F. A. Wichmann, and W. Brendel. Partial success in closing the gap between human and machine vision. InProc. NeurIPS, 2021. 2
work page 2021
-
[8]
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel. ImageNet-trained CNNs are biased towards texture; in- creasing shape bias improves accuracy and robustness. InProc. ICLR,
Show all 41 references
-
[9]
C. C. Gramazio, K. B. Schloss, and D. H. Laidlaw. The relation between visualization size, grouping, and user performance.IEEE Transactions on Visualization and Computer Graphics, 20(12):1953– 1962, 2014. 2
1953
-
[10]
Haehn, J
D. Haehn, J. Tompkin, and H. Pfister. Evaluating ‘graphical percep- tion’ with CNNs.IEEE Transactions on Visualization and Computer Graphics, 25(1):641–650, 2019. 2
2019
-
[11]
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick. Masked autoencoders are scalable vision learners. InProc. CVPR, 2022. 8
2022
-
[12]
C. G. Healey. Choosing effective colours for data visualization.Proc. IEEE Visualization, pp. 263–270, 1996. 1, 2, 6
1996
-
[13]
M. N. Hebart, O. Contier, L. Teichmann, A. H. Rockter, C. Y . Zheng, A. Kidder, A. Corriveau, M. Vaziri-Pashkam, and C. I. Baker. THINGS-data: A multimodal collection of large-scale datasets for in- vestigating object representations in human brain and behavior.eLife, 12:e8258...
2023
-
[14]
M. N. Hebart, C. Y . Zheng, F. Pereira, and C. I. Baker. Revealing the multidimensional mental representations of natural objects underly- ing human similarity judgements.Nature Human Behaviour, 4:1173– 1185, 2020. 2, 3, 6
2020
-
[15]
Heer and M
J. Heer and M. Bostock. Crowdsourcing graphical perception: Using mechanical turk to assess visualization design. InProc. ACM CHI, pp. 203–212, 2010. 1
2010
-
[16]
T. Li, Z. Wen, L. Song, J. Liu, Z. Jing, and T. S. Lee. From local cues to global percepts: Emergent gestalt organization in self-supervised vision models.arXiv preprint arXiv:2506.00718, 2025. 2
2025 arXiv
-
[17]
H. Liu, C. Li, Y . Li, and Y . J. Lee. Improved baselines with visual instruction tuning. InProc. CVPR, 2024. 8
2024
-
[18]
Marafioti, O
A. Marafioti, O. Zohar, M. Farr ´e, et al. SmolVLM: Redefining small and efficient multimodal models.arXiv:2504.05299, 2025. 8
2025 arXiv
-
[19]
Masry, D
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguis- tics: ACL 2022, pp. 2263–2279, 2022. 1
2022
-
[20]
C. M. McColeman, F. Yang, T. F. Brady, and S. Franconeri. Rethink- ing the ranks of visual channels.IEEE Transactions on Visualization and Computer Graphics, 28(1):707–717, 2022. 1
2022
-
[21]
Muttenthaler, J
L. Muttenthaler, J. Dippel, L. Linhardt, R. A. Vandermeulen, and S. Kornblith. Human alignment of neural network representations. InProc. ICLR, 2023. 2, 3
2023
-
[22]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, et al. DINOv2: Learning robust visual features without supervision.Transactions on Machine Learn- ing Research, 2024. 8
2024
-
[23]
Pandey and A
S. Pandey and A. Ottley. Benchmarking visual language models on standardized visualization literacy tests.Computer Graphics Forum, 44(3):e70137, 2025. 1, 2
2025
-
[24]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervi- sion. InProc. ICML, 2021. 8
2021
-
[25]
Schrimpf, J
M. Schrimpf, J. Kubilius, M. J. Lee, N. A. R. Murty, R. Ajemian, and J. J. DiCarlo. Integrative benchmarking to advance neurally mecha- nistic models of human intelligence.Neuron, 108(3):413–423, 2020. 2
2020
-
[26]
J. Sha, H. Shindo, K. Kersting, and D. S. Dhami. Gestalt vision: A dataset for evaluating gestalt principles in visual perception. In Proceedings of the 19th International Conference on Neurosymbolic Learning and Reasoning, vol. 284 ofPMLR, pp. 873–890, 2025. 2
2025
-
[27]
Skau and R
D. Skau and R. Kosara. Arcs, angles, or areas: Individual data encod- ings in pie and donut charts.Computer Graphics Forum, 35(3):121– 130, 2016. 1
2016
-
[28]
Sucholutsky, L
I. Sucholutsky, L. Muttenthaler, A. Weller, A. Peng, A. Bobu, B. Kim, B. C. Love, E. Grant, I. Groen, J. Achterberg, J. B. Tenenbaum, et al. Getting aligned on representational alignment.arXiv preprint arXiv:2310.13018, 2023. 2
2023 arXiv
-
[29]
D. A. Szafir. Modeling color difference for visualization design.IEEE Transactions on Visualization and Computer Graphics, 24(1):392– 401, 2018. 1, 2, 6
2018
-
[30]
Talbot, V
J. Talbot, V . Setlur, and A. Anand. Four experiments on the percep- tion of bar charts.IEEE Transactions on Visualization and Computer Graphics, 20(12):2152–2160, 2014. 1
2014
-
[31]
Torfs, K
K. Torfs, K. Vancleef, C. Lafosse, J. Wagemans, and L. de-Wit. The Leuven Perceptual Organization Screening Test (L-POST), an online test to assess mid-level visual perception.Behavior Research Meth- ods, 46(2):472–487, 2014. 2
2014
-
[32]
Tseng, A
C. Tseng, A. Z. Wang, G. J. Quadri, and D. A. Szafir. Revisiting categorical color perception in scatterplots: Sequential, diverging, and categorical palettes. InEuroVis Short Papers, 2024. 1
2024
-
[33]
Verma, K
A. Verma, K. Mukherjee, C. Potts, E. Kreiss, and J. E. Fan. CHART- 6: Human-centered evaluation of data visualization understanding in vision-language models.arXiv preprint arXiv:2505.17202, 2025. 2
2025 arXiv
-
[34]
P. Wang, S. Bai, S. Tan, et al. Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution.arXiv:2409.12191,
-
[35]
Ware.Information Visualization: Perception for Design
C. Ware.Information Visualization: Perception for Design. Morgan Kaufmann, 3rd ed., 2013. 1, 2, 6
2013
-
[36]
Wertheimer
M. Wertheimer. Untersuchungen zur Lehre von der Gestalt. II.Psy- chologische Forschung, 4:301–350, 1923. 1
1923
-
[37]
Whitney and A
D. Whitney and A. Yamanashi Leib. Ensemble perception.Annual Review of Psychology, 69:105–129, 2018. 2
2018
-
[38]
X. Yue, Y . Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, et al. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. InProc. IEEE/CVF CVPR,
-
[39]
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer. Sigmoid loss for language image pre-training. InProc. ICCV, 2023. 8
2023
-
[40]
Zhang, J
K. Zhang, J. Yang, J. P. Inala, C. Singh, J. Gao, Y . Su, and C. Wang. Towards understanding graphical perception in large multimodal mod- els.arXiv preprint arXiv:2503.10857, 2025. 2
2025 arXiv
-
[41]
trial_id
Y . Zhang, J. Pan, Y . Zhou, R. Pan, and J. Chai. Grounding visual illusions in language: Do vision-language models perceive illusions like humans? InProc. EMNLP, pp. 5718–5728, 2023. 1 APPENDIX A EXAMPLE STIMULI Every stimulus below is the exact image presented to the foundat...
2023
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.