Pith. sign in

REVIEW 1 cited by

Brain encoding models based on multimodal transformers can transfer across language and vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.12248 v1 pith:O7KDBI22 submitted 2023-05-20 cs.CL cs.CV

classification cs.CLcs.CV
keywords encodingmodelsmultimodalbrainlanguagerepresentationstransformersvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Encoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained and tested on brain responses to each modality in isolation. Recent advances in multimodal pretraining have produced transformers that can extract aligned representations of concepts in language and vision. In this work, we used representations from multimodal transformers to train encoding models that can transfer across fMRI responses to stories and movies. We found that encoding models trained on brain responses to one modality can successfully predict brain responses to the other modality, particularly in cortical regions that represent conceptual meaning. Further analysis of these encoding models revealed shared semantic dimensions that underlie concept representations in language and vision. Comparing encoding models trained using representations from multimodal and unimodal transformers, we found that multimodal transformers learn more aligned representations of concepts in language and vision. Our results demonstrate how multimodal transformers can provide insights into the brain's capacity for multimodal processing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 16 citations worldwide. Full citation record

  1. Disentangling the Factors of Convergence between Brains and Computer Vision Models

    cs.AI 2025-08 unverdicted novelty 7.0 of 10

    By systematically varying model size, training amount, and image type in DINOv3 vision transformers, this paper shows that brain similarity increases with scale and human-centric data and emerges in a characteristic t...

Pith tools