Pith. sign in

REVIEW 5 cited by

Visual representations in the human brain are aligned with large language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.11737 v2 pith:VUKQSLKU submitted 2022-09-23 cs.CV cs.LGq-bio.NC

classification cs.CVcs.LGq-bio.NC
keywords braininformationrepresentationscaptionscomplexmodelsscenevisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The human brain extracts complex information from visual inputs, including objects, their spatial and semantic interrelations, and their interactions with the environment. However, a quantitative approach for studying this information remains elusive. Here, we test whether the contextual information encoded in large language models (LLMs) is beneficial for modelling the complex visual information extracted by the brain from natural scenes. We show that LLM embeddings of scene captions successfully characterise brain activity evoked by viewing the natural scenes. This mapping captures selectivities of different brain areas, and is sufficiently robust that accurate scene captions can be reconstructed from brain activity. Using carefully controlled model comparisons, we then proceed to show that the accuracy with which LLM representations match brain representations derives from the ability of LLMs to integrate complex information contained in scene captions beyond that conveyed by individual words. Finally, we train deep neural network models to transform image inputs into LLM representations. Remarkably, these networks learn representations that are better aligned with brain representations than a large number of state-of-the-art alternative models, despite being trained on orders-of-magnitude less data. Overall, our results suggest that LLM embeddings of scene captions provide a representational format that accounts for complex information extracted by the brain from visual inputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning

    q-bio.NC 2025-08 unverdicted novelty 6.0 of 10

    Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.

  2. TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A transformer-based encoder that combines text, audio, and video embeddings predicts whole-brain fMRI responses to movies across subjects and won the Algonauts 2025 competition.

  3. Representations in vision and language converge in a shared, multidimensional space of perceived similarities

    q-bio.NC 2025-07 conditional novelty 6.0 of 10

    Similarity judgments of natural scene images and their sentence captions are aligned with each other, with visual brain responses, and with LLM-trained visual models.

  4. Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

    q-bio.NC 2025-05 conditional novelty 6.0 of 10

    Instruction-tuned multimodal LLMs predict fMRI responses to natural images better than vision-only models and on par with CLIP, though most explained variance is shared across instructions.

  5. Multi-modal brain encoding models for multi-modal stimuli

    q-bio.NC 2025-05 conditional novelty 5.0 of 10

    On movie-watching fMRI data, multi-modal vision-audio transformers predict brain activity better than unimodal video or speech models, with video dominating cross-modal alignment and video plus audio jointly contribut...

Pith tools