Pith. sign in

Paper Citation Record · LEDGER

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

As of 15 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.20029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20029 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:53.714175Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6008b149-f63a-466c-a1c2-7c0720922511 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.463444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.463444Z digest=sha256:ed40fe121d510d2de3bf8ff73a40c927df9949b82ca9ff343b9af5b3bbe51231

Observation 5a3b562f-ae32-4eff-851d-50c54f913806 · outbound

This paper cites InstructBLIP (Dai et al.,.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) InstructBLIP (Dai et al.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.237159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.323012Z digest=sha256:4a1c3e781ae535a0386e4d52e01c7834e25900ebec2386731b4f823144a4c7e6

Observation 4272d690-7d15-4aae-97b4-ef46650da518 · outbound

This paper cites Yes,” “No,.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Yes,” “No,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:53.950680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.714175Z digest=sha256:26e0dd482ba9d35a0c198c2bdc07cb1f74740a263a428e61e4070c3c45b4f262

Observation 1a9ad026-2a66-4868-8c9b-e735d6bf5ea1 · outbound

This paper cites The Llama 3 Herd of Models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.923241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.923241Z digest=sha256:586ebfd2267c58478c4f6be3bc4c05c923eca4504d94154d6b7ddddf7ec683ca

Observation bd294af3-9fcd-4021-b3eb-bc9e08df4189 · outbound

This paper cites Microsoft coco: Common objects in context.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Microsoft coco: Common objects in context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.296677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.296677Z digest=sha256:86789d19e1b30a4a8da1ab9c5a5aec09a7b73e06780f52d96fa6e1b898e7d4d3

Observation ef5420d5-bf60-4a57-aec7-a18bbfa206fc · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:56.492727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:52.486214Z digest=sha256:af582216213ae342e16084566cf83a826044f3d18f32a81cae170dec8f2cba32

Observation d1bd5683-8ad1-4bfa-a72d-2b274f026b7b · outbound

This paper cites Speech language models lack important brain-relevant semantics.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Speech language models lack important brain-relevant semantics

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:56.210466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:52.556342Z digest=sha256:cbfdd56dbfa331207d1ea1e966218cb9ff14e26c58e0288e8e830e235d19e6c8

Observation 32fc21fb-8b78-4d9e-9d17-1bafb46267b0 · outbound

This paper cites Tuning in to neural encoding: Linking human brain and artificial supervised representations of language.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Tuning in to neural encoding: Linking human brain and artificial supervised representations of language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.703135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.703135Z digest=sha256:6f2ff94be480e5ad60c49e2635c66d8556c6e94bacf4e32670cac5524b43e674

Observation 30d43f35-4313-4945-b3d5-643980756172 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.788312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.788312Z digest=sha256:ce1d5dab3e568e821ade09233f17fa5e8490e774b5f4f0ce4bda0b58550b4d0b

Observation 40ee46a3-fcc9-4991-9ee2-c181b98ac617 · outbound

This paper cites Natural language supervision with a large and diverse dataset builds better models of human high-level visual cor- tex.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Natural language supervision with a large and diverse dataset builds better models of human high-level visual cor- tex

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.788664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:52.944682Z digest=sha256:59c23c38a5381bb36f0fedfa1c985c0a61594a3c48bdf7d4d8af74b75e8069e0

Observation 0154a279-ab1c-4c37-a938-4adaf1defe66 · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Transformers: State-of-the-art natural language processing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.603814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.035379Z digest=sha256:37fd638de09cf185e44325de61ad331a7619e4b4fa7e2af002445c3a5e537942

Observation e5e9885c-13ff-4e8d-b582-05c08cc75bab · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:53.137091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:53.137091Z digest=sha256:20e3f6c705aa9ed167a22f10340a9dff050a203fe1966e2ab158a02cf5d76b9c

Observation b27bc93b-42a4-40a4-a7d5-a0b7defa4b0a · outbound

This paper cites mid temporal lobe bodies.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) mid temporal lobe bodies

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.391941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.218111Z digest=sha256:b0598e4e256973b6b17d831faca6444019c63c7f56643ca06b995fbfac2557bb

Observation 01491c0b-48a0-4842-ade4-c7e0d4003240 · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:55.116310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.385595Z digest=sha256:1d8ee89d212f17502c77ea374997bb24169e7333878d9b89dbfa282a1fc0d675

Observation 2d486ccb-729d-4194-bb47-984abd3823a9 · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:04:54.987578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.471257Z digest=sha256:b00649a70d8160900ad903656ef36a42cb0a263da20db88698f5b1e129dccd28

Observation fcfa8868-7021-4d47-8262-27de04c40461 · outbound

This paper cites The color bar highlights color codes for each instruction.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The color bar highlights color codes for each instruction

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:54.702164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.511816Z digest=sha256:c0f193d70add04c9589a1ca5edab86a02c0e0991c88205d305ec853e9de4ba66

Observation a993ba63-c9ec-4e8e-8890-80fabe7ef333 · outbound

This paper cites an unresolved cited work.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Unresolved cited work

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:04:54.163990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.630159Z digest=sha256:0f38299b76007835eed6c6f1b67bd5ab2a3292ebbc765aa5abce6cc450894fad

Observation 4a698fe1-db21-4942-8ed8-3644532d760d · outbound

This paper cites The color bar highlights color codes for each layer.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The color bar highlights color codes for each layer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:54.428394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:53.570042Z digest=sha256:75acd559df8e97861995ff91b40d4126e851c241d3dec31e853c7f0ca435557d

Observation 49f1b797-11f5-486f-bd89-5ac289dd2bd0 · outbound

This paper cites Shared computational principles for language processing in humans and deep language models.Nature Neuroscience, 25(3):369–380,.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Shared computational principles for language processing in humans and deep language models.Nature Neuroscience, 25(3):369–380,

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:56.870712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:52.053045Z digest=sha256:dd2c28017cab26cd230acb27ab96e6e35e5f111ddebcd1be8a8a375231f620f4

Observation 077b16f1-8bfc-4cd6-850a-d2ce1188c6a1 · outbound

This paper cites Mistral 7B.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Mistral 7B

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.192451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.192451Z digest=sha256:d0aa4e8664052f4121ab96e8fa6999217cf70013d66048d25490f09cc7b1349d

Observation ded41309-6046-4826-9a9a-4b67125ee344 · outbound

This paper cites Visual representations in the human brain are aligned with large language models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Visual representations in the human brain are aligned with large language models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.812008Z digest=sha256:749b131bbcc43e0dd6fccbeb84328d966feb74d825dfe9df423fce9e13a9d566

Observation 4031bc71-7040-47ef-b0ef-5a4a88e636d4 · outbound

This paper cites What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? bioRxiv, pp.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? bioRxiv, pp

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.690010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.690010Z digest=sha256:1cb68a6554f2d3566b381b26f4b0698f42c21724b4508a8cd98764a1fe76e3f6

Observation b3c81966-98d0-43c4-ab4e-1a65c5eec772 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Learning Transferable Visual Models From Natural Language Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:52.633416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:52.633416Z digest=sha256:b8f41310518b8bf1c7c36073a2bdd9cef4f5f63a6123706c82ff0d3db33243e9

Observation 2da6bc0d-5468-4a17-8ea0-d5438a48a744 · outbound

This paper cites Neural taskonomy: Inferring the similarity of task- derived representations from brain activity.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Neural taskonomy: Inferring the similarity of task- derived representations from brain activity

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:55.980113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:52.844290Z digest=sha256:2d6dcd2377bd73bc955b3ef5437f6e7356c1dc87db9b772d325883ba46277566

Observation 2b172c6a-0b48-41b3-9ff3-20874950eab3 · outbound

This paper cites Instruction-tuning Aligns LLMs to the Human Brain.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) Instruction-tuning Aligns LLMs to the Human Brain

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:51.556068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:51.556068Z digest=sha256:5071e062ec29aebf788dd5e465b92becb6fa43fa18c13fb5c920e74261154937

Observation 0960fd90-1235-4a58-b508-589fceb87291 · outbound

This paper cites The brain tells a story: Unveiling distinct representations of semantic content in speech, objects, and stories in the human brain with large language models.

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain) The brain tells a story: Unveiling distinct representations of semantic content in speech, objects, and stories in the human brain with large language models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:56.721573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:04:52.401068Z digest=sha256:b22ce6abd05a57c7144945f28c65f31cfb5fc29159f3f0b855c0f615cb39bac5

Pith citing papers

No inbound Pith citation observations are available.