REVIEW 5 cited by
Visual representations in the human brain are aligned with large language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The human brain extracts complex information from visual inputs, including objects, their spatial and semantic interrelations, and their interactions with the environment. However, a quantitative approach for studying this information remains elusive. Here, we test whether the contextual information encoded in large language models (LLMs) is beneficial for modelling the complex visual information extracted by the brain from natural scenes. We show that LLM embeddings of scene captions successfully characterise brain activity evoked by viewing the natural scenes. This mapping captures selectivities of different brain areas, and is sufficiently robust that accurate scene captions can be reconstructed from brain activity. Using carefully controlled model comparisons, we then proceed to show that the accuracy with which LLM representations match brain representations derives from the ability of LLMs to integrate complex information contained in scene captions beyond that conveyed by individual words. Finally, we train deep neural network models to transform image inputs into LLM representations. Remarkably, these networks learn representations that are better aligned with brain representations than a large number of state-of-the-art alternative models, despite being trained on orders-of-magnitude less data. Overall, our results suggest that LLM embeddings of scene captions provide a representational format that accounts for complex information extracted by the brain from visual inputs.
Forward citations
Cited by 5 Pith papers
-
Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.
-
TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction
A transformer-based encoder that combines text, audio, and video embeddings predicts whole-brain fMRI responses to movies across subjects and won the Algonauts 2025 competition.
-
Representations in vision and language converge in a shared, multidimensional space of perceived similarities
Similarity judgments of natural scene images and their sentence captions are aligned with each other, with visual brain responses, and with LLM-trained visual models.
-
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
Instruction-tuned multimodal LLMs predict fMRI responses to natural images better than vision-only models and on par with CLIP, though most explained variance is shared across instructions.
-
Multi-modal brain encoding models for multi-modal stimuli
On movie-watching fMRI data, multi-modal vision-audio transformers predict brain activity better than unimodal video or speech models, with video dominating cross-modal alignment and video plus audio jointly contribut...
Discussion (0). Continue with ORCID to comment.