Pith. sign in

Paper Citation Record · LEDGER

Representations in vision and language converge in a shared, multidimensional space of perceived similarities

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2507.21871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21871 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:21:22.544303Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:32:28.422930Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T16:32:28.671394Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9bdf7fb3-8a6f-4f18-8ba3-ce6f87e63edf · outbound

This paper cites Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs).

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.778065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.476705Z digest=sha256:d7e9d5db3d0b7e069e5fb608ee8602c3ab73f0a437e1843904d1df89aea22145

Observation 808fd061-f75c-4eae-8e7f-de22674eed3a · outbound

This paper cites (A) Cross-validated non-negative least squares regression was used to model the brain RDMs at every searchlight location using the behavioural RDMs derived from our MA tasks.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities (A) Cross-validated non-negative least squares regression was used to model the brain RDMs at every searchlight location using the behavioural RDMs derived from our MA tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.727176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.496414Z digest=sha256:72bb8030e3dc19faae047dc103399e4280d368dc483fc8e75955a6e7262695a2

Observation 5bbd874c-07bd-406d-97ab-218d9a5d3fd2 · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.716796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.500243Z digest=sha256:7e7df68343d637f7b81e6cc9b79d384b63e837a5868d51478379ca8e14624f8d

Observation 09891198-0646-410c-acc0-115c75458ddb · outbound

This paper cites linguistic modality.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities linguistic modality

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.748159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.488582Z digest=sha256:a9e46c2e8f9479a7028554c87aee7e8abd7bea87e428b7997af7acf6f97d2817

Observation 2a6fb438-672d-4b6a-b27c-01a0d165e43e · outbound

This paper cites (A) Participants completed the MA task either on 100 natural scene images (visual modality left) or 100 sentence captions 8 describing the images (linguistic modality right).

Representations in vision and language converge in a shared, multidimensional space of perceived similarities (A) Participants completed the MA task either on 100 natural scene images (visual modality left) or 100 sentence captions 8 describing the images (linguistic modality right)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.738082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.492359Z digest=sha256:4a90f6cf331b1ced834246475634cece4acd491d68bbd59b66159d75b21309b9

Observation 6cb16c37-3857-4953-adea-9511c086b9a4 · outbound

This paper cites Scaling Laws for Transfer.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Scaling Laws for Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.532478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.532478Z digest=sha256:d525c8c49d2e0dc1da4e18a1a53011193eede3ba5218b63d6f2978ffb680be9d

Observation 698df9c9-edc8-4cb8-b09b-0509435d986b · outbound

This paper cites VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:21:22.581138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.544303Z digest=sha256:52658976ce3fdb7c037c03a7aef46ebcd279e09e1a0199d42e60f700fab61d08

Observation 678bc90a-5ea0-4743-acde-4b632dd97752 · outbound

This paper cites A., Schmitz, T.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities A., Schmitz, T

Reference 134

Resolution
verified exact
doi, observed 2026-08-06T12:21:22.636069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.525364Z digest=sha256:994b768c8fb54c1af4281b1c3b5f03e827943ec0c3fc38e37297fa0a9e3b763d

Observation 24dcb9c7-e65c-45d3-be8e-6ace6e4c0c1b · outbound

This paper cites A., Kiani, R., Bodurka, J., Esteky, H., Tanaka, K., & Bandettini, P.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities A., Kiani, R., Bodurka, J., Esteky, H., Tanaka, K., & Bandettini, P

Reference 245

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.624190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.537199Z digest=sha256:c250fe45a964ae2b5d75e7c60a4ff6ab2eb4a2625041c15c1ab7878aae84e05f

Observation 0f02e926-5b21-49b1-aaf4-c65258533fd1 · outbound

This paper cites visual modality.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities visual modality

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.758598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.485065Z digest=sha256:370c1e678ff651b99409388a7b4d639617b148525bd99cec5c7ce903bf6e7873

Observation c8bfdb37-929c-4e63-94b7-f280d91f9d21 · outbound

This paper cites We show that a similar relational structure emerges for both linguistic and visual inputs.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities We show that a similar relational structure emerges for both linguistic and visual inputs

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.706659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.503950Z digest=sha256:ffde289769c25d13124227bb4a2d54a286cde8e3a265765aef01d43760573b42

Observation 5545cf45-c302-4c49-ac71-85dcceb97a61 · outbound

This paper cites The significance of correlations was tested using one-sided t-test across participants and corrected for multiple comparisons at FDR p < 0.05.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities The significance of correlations was tested using one-sided t-test across participants and corrected for multiple comparisons at FDR p < 0.05

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.659171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.518202Z digest=sha256:2b807571e066af6c0d20c6c1a5e0ade19180b492a0fdb68d21e722e5ec40d28a

Observation b50d41b6-b50e-4c2d-bff9-78862206375c · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.682747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.510706Z digest=sha256:eaec69df252dd93e545d0bb38c3c8baf62399927381b8f180478c8381705d176

Observation 044c86dc-3609-493a-a98d-061e1f87a0a4 · outbound

This paper cites It may be that the visual system translates sensory inputs into modality-agnostic representations that reflect stable, relational patterns observed in the real-world environment.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities It may be that the visual system translates sensory inputs into modality-agnostic representations that reflect stable, relational patterns observed in the real-world environment

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.695493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.507342Z digest=sha256:7e9678fae1ed651d71ecd4ccd28d68189bee10338addc867e770273e0ca84975

Observation c1b187da-df2f-4334-a4a4-b53ba3204fdc · outbound

This paper cites The sentence captions were collected from five human annotators as part of the Microsoft Common Objects in Context database (Lin et al., 2014).

Representations in vision and language converge in a shared, multidimensional space of perceived similarities The sentence captions were collected from five human annotators as part of the Microsoft Common Objects in Context database (Lin et al., 2014)

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:21:22.671282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.514391Z digest=sha256:19970b576b0b24c543fdedc3967a7c376fd91c3c0b4f3d13deab19f6243d8856

Observation 8746e166-da76-43e8-851d-94d81dc74e1e · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.767941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.481066Z digest=sha256:23318918c41832f0dc6beee4b34d41c297e6fd4c8400e31c8ab8805886d74b36

Observation be82ce1a-0acd-48f6-9705-f40062c5cb0f · outbound

This paper cites an unresolved cited work.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Unresolved cited work

Reference 4081

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:21:22.647477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T12:21:22.522047Z digest=sha256:df8867d9a00c2ab3a157c057094326872383b3be6b8b9f0a3b40e0969238702f

Observation 233de5ba-3820-4e13-b61c-6c4c11479083 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 6241

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.540405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.540405Z digest=sha256:62a9b7b4e5997e3173f2889a42eb16ea2b174cb2f699135c519c950b07ec41c1

Observation 2112a0ed-47c6-4cc6-acf6-2f903dea97d9 · outbound

This paper cites Visual representations in the human brain are aligned with large language models.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities Visual representations in the human brain are aligned with large language models

Reference 9383

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.528735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.528735Z digest=sha256:3bf4c06321be8b9631829bc40d34ffb7e7c92e86a346339d0de6ec697854ea87

Pith citing papers

Observation e30cad62-9bbc-4c50-a7ce-b698b07ceb1e · inbound

Disentangling the Factors of Convergence between Brains and Computer Vision Models cites this paper.

Disentangling the Factors of Convergence between Brains and Computer Vision Models Representations in vision and language converge in a shared, multidimensional space of perceived similarities

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T16:32:28.674050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T16:32:28.422930Z digest=sha256:4286d12af14dfa72972dc9dd689665a89b0de98dd48ecbcf91c2c186910b07ac