REVIEW 3 major objections 6 minor 1 cited by
Representations in vision and language converge in a shared, multidimensional space of perceived similarities
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that human similarity judgments for natural scene images and for sentence captions of those images converge on a shared representational structure, captured by LLM embeddings and predictive of visual brain responses.
desk verdict A useful behavioral extension with a real selection-bias concern; the shared-space claim needs tempering before it's general. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is representational similarity analysis (RSA) built on multiple-arrangement (MA) similarity judgments, together with cross-validated non-negative least squares (NNLS) mapping from behavioural RDMs to brain and model RDMs. The MA task converts pairwise perceived similarity into a representational dissimilarity matrix (RDM) per modality; the fixed-effects Spearman correlation between the visual and linguistic RDMs is the headline evidence of convergence. NNLS then asks whether a weighted combination of participants' behavioural RDMs can predict held-out brain and network RDMs, and the comparison between LLM-trained and category-trained recurrent convolutional networks isolates the effect of the training objective on explaining human similarity.
What would settle it
Run the same multiple-arrangement task on a new set of images and captions selected at random, or matched for category balance, rather than by greedy maximisation of semantic spread in LLM embedding space. If the visual–linguistic RDM correlation falls far below 0.781, or caption-based RDMs no longer predict occipitotemporal searchlights in the Natural Scenes Dataset, then the shared-space result is an artefact of stimulus selection rather than a general property of scene representations.
Extended reading notes
Core claim
The authors' central claim is that the mind represents natural scenes and their verbal descriptions in a shared, multidimensional similarity space. At the behavioural level, the rank-transformed average dissimilarity matrix from image arrangements correlates at $\rho = 0.781$ with the corresponding matrix from caption arrangements. Both behavioural geometries predict visually evoked brain responses with a similar cortical profile, peaking in mid- and high-level visual areas, and the linguistic prediction survives when participants judged captions before ever seeing the images. Computational models trained to map images onto MPNet caption embeddings reproduce the human similarity geometry better than category-trained networks with identical architecture and training data, indicating that the shared space is embedding-like rather than category-like.
Load-bearing premise
The result depends on the 100 images and captions being representative: they were chosen by a greedy algorithm to spread maximally across the semantic space of LLM caption embeddings, and if this selection favours highly nameable, semantically distinct scenes, the cross-modal agreement and brain predictions may be inflated.
Editorial extensions
If this is right
- LLM embeddings of scene captions can serve as a practical proxy for how humans judge the similarity of natural scenes, useful for generating behavioural predictions without collecting new similarity ratings.
- Linguistic similarity structure alone can predict visual brain responses, suggesting that language-based behavioural measures can stand in for visual experience in some brain-encoding models.
- Models trained to map images onto LLM embeddings are better aligned with human similarity judgments than models trained on discrete category labels, pointing to continuous semantic structure rather than categorical structure as the organizing principle.
- The shared relational geometry found here could be a general principle of conceptual representation, potentially extending beyond vision and language to sound, action, and other modalities.
- The finding that visual judgments favour the image-computable LLM-trained model while linguistic judgments favour pure caption embeddings implies that the shared space retains complementary modality-specific information, motivating multimodal models of similarity.
Reading between the lines
- One step beyond the paper: the same multiple-arrangement design could be run with audio or video stimuli and their verbal descriptions to test whether the shared similarity space generalizes across non-visual sensory domains.
- If the shared space is truly modality-agnostic, then populations without visual experience, such as congenitally blind individuals, should still produce caption-arrangement geometries that predict sighted viewers' visual brain responses; this is a testable prediction the paper does not run.
- The modest random-effects correlation ($\rho = 0.161$) relative to the fixed-effects correlation ($\rho = 0.781$) suggests substantial individual variation in how people map language onto visual similarity; probing which participants show high cross-modal alignment could reveal cognitive or linguistic factors behind the shared space.
- A direct testable extension would be to compare similarity judgments of images to judgments of captions that systematically vary in how much perceptual versus conceptual detail they mention, which would separate the contribution of low-level visual features from high-level semantic content in driving the cross-modal agreement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents new behavioral similarity judgments for 100 natural scene images and their corresponding sentence captions, collected with a multiple-arrangements task. The authors report a high fixed-effects Spearman correlation (0.781) between the image and caption dissimilarity matrices, and show that both visual and linguistic judgments predict fMRI responses in occipitotemporal cortex using the Natural Scenes Dataset. They further compare recurrent convolutional neural networks trained on category labels versus networks trained on MPNet caption embeddings, finding that the LLM-trained networks better match the behavioral similarity structure. The central claim is that visual and linguistic similarity judgments converge in a shared, modality-agnostic representational space that is reflected in the brain.
Significance. If the central claim holds, the study provides important behavioral evidence for a shared representational format between vision and language, and it strengthens the case for using LLM embeddings as models of human semantic similarity. The study has notable strengths: it includes a hidden-image control condition, Session 1 analyses that rule out cross-modal familiarity, a no-NNLS control for the brain-prediction procedure, and model comparisons that hold architecture, training data, and random seeds constant. The behavioral dataset, once released, will be a valuable resource for the field. The weakness identified below concerns the stimulus selection, which directly affects the generality of the main claims.
major comments (3)
- [Methods, Stimuli] The 100 stimuli were selected by a greedy algorithm to maximally span a semantic space defined by MPNet caption embeddings. This same embedding space is used to define the linguistic RDM (via the captions) and to train the LLM-trained RCNN. Consequently, the headline fixed-effects correlation (0.781) and the superiority of the LLM-trained RCNN over category-trained models may be inflated by construction: the stimulus set was deliberately chosen to be far apart on the very dimensions that the captions and the LLM encode. The paper provides no control analysis that varies the stimulus selection, such as a random subset of the 515-image pool, a set selected to maximize low-level visual diversity, or a statistical control for caption-embedding distance. Because the central claim is that visual and linguistic similarity relations converge in a general, modality-agnostic space, the selection effect must be addressed; otherwise the claim is only established for a stimulus set chosen to maximize semantic spread in one specific embedding space.
- [Results, Uncovering a shared similarity space] The paper reports a fixed-effects Spearman correlation of 0.781 between the averaged visual and linguistic RDMs, and a random-effects correlation of 0.161 (p < 0.0001). The fixed-effects value is presented as showing 'strong representational overlap,' but the random-effects value indicates that the shared geometry accounts for a very small fraction of the variance in individual participants' RDMs. The authors do not report the distribution of individual participant correlations, nor do they discuss why the fixed-effects and random-effects values differ so markedly. This discrepancy is load-bearing for the strength of the behavioral claim: a consensus structure can be high even when individuals show little alignment, and the paper's language does not convey this nuance. The authors should report individual-level statistics and temper the interpretation of the 0.781 value accordingly.
- [Brain prediction analyses (Results and Supplementary Figures)] The same stimulus-selection concern applies to the fMRI analyses. The NSD brain RDMs are computed for the same 100 images selected to maximize caption-embedding spread, so mid- and high-level visual responses may be dominated by exactly the semantic distinctions that the caption RDMs also encode. The Session 1 and no-NNLS controls are valuable because they rule out some alternative explanations, but they do not address the stimulus-selection inflation. The paper should explicitly acknowledge this limitation and, if possible, provide a control analysis using a more diverse stimulus set or a partial-correlation approach that removes variance related to low-level image features. At minimum, the Discussion should contextualize the claim of a general shared space in light of the selection procedure.
minor comments (6)
- [Methods, Stimuli] The text refers to the 'Special100 database,' but this is a stimulus set rather than a database; consider using 'Special100 set' for clarity.
- [Introduction / Methods] The specific language model (MPNet) is not introduced until the Methods section; defining it in the Introduction would help readers interpret the 'LLM embeddings' used throughout.
- [Methods, Estimating the representational dissimilarity matrices] The expression '((100 items2 - 100)/2 = 4950 pairs)' is awkward; it should be written as '100 choose 2 = 4950 pairs'.
- [Supplementary Figure 8] The caption mentions 'cosine similarities' while the main text and other figures refer to Pearson correlations; please ensure consistent terminology.
- [Code and data availability] The code and data are stated to be available 'upon publication'; for a journal emphasizing reproducibility, providing a permanent repository link from the outset would be preferable.
- [General] The spelling 'Alexnet' appears in multiple places; the standard spelling is 'AlexNet'.
Circularity Check
Minor circularity from LLM-based stimulus selection and self-cited model; central behavioral and brain results remain independent.
-
other
[Methods, Stimuli]
"A greedy algorithm (Allen et al., 2022) selected the set of 100 images from a special set of 515 images such that they maximally span the high-level semantic space, which was defined by embeddings of the images’ sentence captions."
The same LLM caption-embedding space used to select the 100 images later defines the linguistic model (MPNet RDM) and serves as the training target for the LLM-trained RCNN that the paper reports best captures human similarity. Thus the strong visual-linguistic RDM correlation (fixed-effects Spearman = 0.781) and the favorable model comparisons are estimated on a stimulus set deliberately constructed to maximize spread along the very semantic dimensions being tested. This selection can inflate the apparent convergence between modalities and makes the result less a neutral discovery of a shared space than a confirmation of structure built into stimulus selection.
full rationale
The paper's central behavioral result is a direct fixed-effects Spearman correlation between averaged human visual and linguistic RDMs; no parameter is fitted to the target and the correlation is not guaranteed by construction. The brain-prediction analysis uses cross-validated NNLS and is reproduced without NNLS in Supplementary Figures 5 and 6, so the reweighting procedure is not circular. The LLM-trained RCNN was pre-trained on COCO in prior work (Doerig et al., 2024) and evaluated on new human behavioral data, so that self-citation is externally falsifiable and does not reduce to the present target. The only substantive circularity concern is the stimulus-selection step: the 100 images were chosen to maximize span in MPNet caption-embedding space, and that same space is later used as the linguistic model and as the RCNN training objective. This can inflate the headline visual-linguistic and model correlations, so the generality of the shared-space claim is weaker than presented. The paper also candidly notes a limitation: it does not test neural activity during sentence reading, but that is not a circular step. Overall, no derivation reduces to its own inputs; the self-citation and stimulus-selection issues are minor and partial, warranting a low score.
Assumptions & free parameters
free parameters (1)
- NNLS participant weights =
per participant, per fold
assumptions (4)
- domain assumption Multiple arrangements similarity judgments reflect stable perceptual and semantic similarity relations.
- domain assumption A linear (NNLS) mapping from behavioral to brain/RCNN RDMs is an appropriate model of representational alignment.
- domain assumption Searchlight spheres of radius 6 voxels capture local representational geometry.
- ad hoc to paper The greedy selection of the 100 stimuli based on caption embeddings yields a representative sample.
Cite this review
Pith. "Pith review of Representations in vision and language converge in a shared, multidimensional space of perceived similarities." pith.science (2026). https://pith.science/paper/GE64ZTCK
@misc{pith2026250721871,
author = {Pith},
title = {Pith review of: Representations in vision and language converge in a shared, multidimensional space of perceived similarities},
year = {2026},
howpublished = {\url{https://pith.science/paper/GE64ZTCK}},
note = {Machine review of arXiv:2507.21871}
}
read the original abstract
Humans can effortlessly describe what they see, yet establishing a shared representational format between vision and language remains a significant challenge. Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs). This raises the possibility that sensory systems converge in their inherent ability to transform their inputs onto shared, embedding-like representational space. However, it remains unclear how such a space manifests in human behaviour. To investigate this, sixty-three participants performed behavioural similarity judgements separately on 100 natural scene images and 100 corresponding sentence captions from the Natural Scenes Dataset. We found that visual and linguistic similarity judgements not only converge at the behavioural level but also predict a remarkably similar network of fMRI brain responses evoked by viewing the natural scene images. Furthermore, computational models trained to map images onto LLM-embeddings outperformed both category-trained and AlexNet controls in explaining the behavioural similarity structure. These findings demonstrate that human visual and linguistic similarity judgements are grounded in a shared, modality-agnostic representational structure that mirrors how the visual system encodes experience. The convergence between sensory and artificial systems suggests a common capacity of how conceptual representations are formed-not as arbitrary products of first order, modality-specific input, but as structured representations that reflect the stable, relational properties of the external world.
Figures
Forward citations
Cited by 1 Pith paper
-
Disentangling the Factors of Convergence between Brains and Computer Vision Models
By systematically varying model size, training amount, and image type in DINOv3 vision transformers, this paper shows that brain similarity increases with scale and human-centric data and emerges in a characteristic t...
Reference graph
Works this paper leans on
-
[1]
Representations in vision and language converge in a shared, multidimensional space of perceived similarities Katerina Marie Simkova (katerina.m.simkova@dartmouth.edu) Department of Psychological and Brain Sciences, Dartmouth College Hanover, NH, USA Adrien Doerig (adrien.doerig@fu-berlin.de) Department of Education and Psychology, Freie Universität Berli...
work page 2024
-
[2]
Representational alignment of behaviour-predicted and observed brain RDMs in natural scene viewing. (A) Cross-validated non-negative least squares regression was used to model the brain RDMs at every searchlight location using the behavioural RDMs derived from our MA tasks. (B) We averaged Pearson correlations across folds and determined significance acro...
work page 2024
-
[3]
Representational alignment between behaviour-predicted RCNN RDMs and observed RCNN RDMs. Bars indicate Pearson correlation between behaviour-predicted RCNN RDMs and observed RCNN RDMs for each layer and time step, averaged across 10 models trained with different random seeds. Black bars represent Pearson correlation between the behaviour-predicted MPNet R...
work page 1970
-
[4]
and used a computer mouse to arrange the images according to their similarity, positioning the images such that the distance between them reflects their similarity. In arrangement of sentence captions (henceforth “linguistic modality”), participants followed the same procedure, but each sentence initially appeared as an asterisk with the text revealed onl...
work page 2012
-
[5]
Experimental design of the MA tasks. (A) Participants completed the MA task either on 100 natural scene images (visual modality left) or 100 sentence captions 8 describing the images (linguistic modality right). To handle the large set of sentence captions, we adapted the MA method such that every caption was depicted as an asterisk and only the item curr...
work page 2022
-
[23]
A., Feder, A., Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, S
Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A., Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, S. C., Casto, C., Fanda, L., Doyle, W., Friedman, D., … Hasson, U. (2021). Thinking ahead: spontaneous prediction in context as a keystone of language in humans and machines. In bioRxiv...
arXiv 2021
-
[128]
VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment
Nikolaus, M., Mozafari, M., Asher, N., Reddy, L., & VanRullen, R. (2024). Modality-agnostic fMRI decoding of vision and language. In arXiv [cs.CV]. arXiv. https://scholar.google.com/citations?view_op=view_citation&hl=en&citation_for_view=1pwyaYgAAAAJ:4xDN1ZYqzskC Oosterhof, N. N., Wiggett, A. J., Diedrichsen, J., Tipper, S. P., & Downing, P. E. (06/2010)....
work page Pith review arXiv 2024
-
[134]
28 Charest, I., Kievit, R. A., Schmitz, T. W., Deca, D., & Kriegeskorte, N. (2014). Unique semantic space in the brain of each beholder predicts perceived similarity. Proceedings of the National Academy of Sciences of the United States of America, 111(40), 14565–14570. Cichy, R. M., Kriegeskorte, N., Jozwik, K. M., van den Bosch, J. J. F., & Charest, I. (...
work page 2014
Show all 19 references
-
[245]
A., Kiani, R., Bodurka, J., Esteky, H., Tanaka, K., & Bandettini, P
Kriegeskorte, N., Mur, M., Ruff, D. A., Kiani, R., Bodurka, J., Esteky, H., Tanaka, K., & Bandettini, P. A. (2008). Matching categorical object representations in inferior temporal cortex of man and monkey. Neuron, 60(6), 1126–1141. Lahner, B., Dwivedi, K., Iamshchinina, P., G...
2008
-
[2012]
visual modality
on stimuli obtained from the Special100 database of the Natural Scenes Dataset (NSD; Allen et al., 2022). These 100 natural scene images, each with a corresponding sentence caption describing the image, were selected to maximise semantic diversity (see Allen et al., 2022 for s...
2022
-
[2013]
We show that a similar relational structure emerges for both linguistic and visual inputs
- its role in other modalities has remained underexplored. We show that a similar relational structure emerges for both linguistic and visual inputs. This finding is consistent with that of Dima et al. (2024) who showed overlap between similarity judgements of actions in video...
2024
-
[2015]
The significance of correlations was tested using one-sided t-test across participants and corrected for multiple comparisons at FDR p < 0.05
and nilearn (https://github.com/nilearn/nilearn) libraries. The significance of correlations was tested using one-sided t-test across participants and corrected for multiple comparisons at FDR p < 0.05. For RCNN comparisons, Pearson correlations were averaged across 10 seeds f...
2022
-
[2019]
and represent this information in the visual cortex (Bedny et al., 2011). Altogether, this converging evidence suggests that what the visual system ultimately cares about is a stable, relational structure, and this promotes the alignment between language and vision. Further an...
2011
-
[2021]
It may be that the visual system translates sensory inputs into modality-agnostic representations that reflect stable, relational patterns observed in the real-world environment
or LLM embeddings (Doerig et al., 2024). It may be that the visual system translates sensory inputs into modality-agnostic representations that reflect stable, relational patterns observed in the real-world environment. Reactivation of this relational structure in language may...
2024
-
[2022]
The sentence captions were collected from five human annotators as part of the Microsoft Common Objects in Context database (Lin et al., 2014)
selected the set of 100 images from a special set of 515 images such that they maximally span the high-level semantic space, which was defined by embeddings of the images’ sentence captions. The sentence captions were collected from five human annotators as part of the Microso...
2012
-
[2024]
showed that encoding models trained on visual brain activity generalise to activity in natural language comprehension and vice versa. This converging evidence from language and vision motivates the possibility that various modalities all transform their low-level sensory input...
2024
-
[4081]
Caucheteux, C., Gramfort, A., & King, J.-R. (2022). Deep language algorithms predict semantic comprehension from brain activity. Scientific Reports, 12(1), 16327. Caucheteux, C., & King, J.-R. (2022). Brains and algorithms partially converge in natural language processing. Com...
2022
-
[6241]
T., Pan, B., Jin, S
Lahner, B., Dwivedi, K., Iamshchinina, P., Graumann, M., Lascelles, A., Roig, G., Gifford, A. T., Pan, B., Jin, S. Y., Apurva Ratan Murty, N., Kay, K., Oliva, A., & Cichy, R. (2023). BOLD Moments: modeling short visual events 2 through a video fMRI dataset and metadata. bioRxi...
2023 arXiv
-
[9383]
C., Janarthanan, S., Culham, J
29 Dima, D. C., Janarthanan, S., Culham, J. C., & Mohsenzadeh, Y. (2024). Shared representations of human actions across vision and language. Neuropsychologia, 202(108962), 108962. Doerig, A., Kietzmann, T. C., Allen, E., Wu, Y., Naselaris, T., Kay, K., & Charest, I. (2022). S...
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.