Pith. sign in

Paper Citation Record · LEDGER

Cross-modal Information Flow in Multimodal Large Language Models

As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2411.18620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18620 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:06:10.147882Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:44:10.027628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T00:02:17.731733Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 94d4e1d1-55df-4aef-a79d-d5fdbb77f07b · outbound

This paper cites https : / / www.

Cross-modal Information Flow in Multimodal Large Language Models https : / / www

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.836713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:09.948464Z digest=sha256:7755c04502bf8bcb05de23d38a8347d89b1d6eedbc32a4040c642b7fdb8ccdae

Observation b93005f4-16e5-4816-8d8c-3f3374ffa350 · outbound

This paper cites https: //huggingface.co/lmms- lab/llama3- llava- next-8b, 2024.

Cross-modal Information Flow in Multimodal Large Language Models https: //huggingface.co/lmms- lab/llama3- llava- next-8b, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.823396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:09.953444Z digest=sha256:8043ce468607f4dc90b3f597fad913dd8389b754df19438e00b32c5fe090590b

Observation 656b0e57-6fb0-4515-80e9-c2ccddac5ba5 · outbound

This paper cites Vl-interpret: An interactive visualization tool for interpreting vision-language transformers.

Cross-modal Information Flow in Multimodal Large Language Models Vl-interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.957745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.957745Z digest=sha256:0f19159403a8c17b6c77b31c4a2b50c9a403c1bdcd4c8605d5f0d892da4a5600

Observation d663816d-610a-4413-b229-5146f384283c · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Cross-modal Information Flow in Multimodal Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.962225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.962225Z digest=sha256:fa0ca07b4ceb73720eba7e50ef64d58176f3179f73088204eedd648d640c2f8d

Observation 35af6b75-8de0-45d1-a119-2ea0f693bbd5 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond,.

Cross-modal Information Flow in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.801689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:09.966787Z digest=sha256:3f8add2535d3bf7e8446da8a59d43377c902ba59f07402b01d233483ad54dce0

Observation 8b5ca4c8-7f59-4eb3-a0c7-10259dca8f40 · outbound

This paper cites Understanding Information Storage and Transfer in Multi-modal Large Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Understanding Information Storage and Transfer in Multi-modal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.975089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.975089Z digest=sha256:bd27f26c1a4ec110c8a4f16a85da6a067eb14ae410203b583169d0bca7317fdb

Observation 5679eee8-5e35-46a9-8d91-13f3d1c77669 · outbound

This paper cites Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:06:10.458456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:09.979415Z digest=sha256:9b454dc6b8a96b558da0f741189fb0a920f709fe367ddb368ac3dd1e952c7df0

Observation a4c46df5-8b62-4585-9e3b-cf128bfcc9dc · outbound

This paper cites Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:06:10.438330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:09.983791Z digest=sha256:7ddd01cbc61ea27c600733014ea6765fd491014d3380ac32ccbf1b2f211efed0

Observation ac4c82cb-e752-4536-b798-a8e0067ceac1 · outbound

This paper cites Probing multimodal embeddings for linguistic properties: the visual-semantic case.

Cross-modal Information Flow in Multimodal Large Language Models Probing multimodal embeddings for linguistic properties: the visual-semantic case

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.788524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:09.988420Z digest=sha256:825f50844fd6567706c03a2ce99bfce326d79b2961dd5e9226b408d082469b00

Observation 4bfa2db5-488c-44fb-baad-05e8fa16563f · outbound

This paper cites Knowledge Neurons in Pretrained Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Knowledge Neurons in Pretrained Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.992993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.992993Z digest=sha256:a37ad7e5025ec3b4a3a689a3923ba791032c7f36b4c1fd3bcb0cca3de6b1f913

Observation c359de28-3689-433e-a264-2f747cfb7911 · outbound

This paper cites InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,.

Cross-modal Information Flow in Multimodal Large Language Models InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.997576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.997576Z digest=sha256:919ab91f12578b75329288643232d2ebfd0fc169c8ebe3e01fba42e618d4c503

Observation 8ddc5ead-d7ef-44f1-805f-914c96ea8c55 · outbound

This paper cites Analyzing Transformers in Embedding Space.

Cross-modal Information Flow in Multimodal Large Language Models Analyzing Transformers in Embedding Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.006430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.006430Z digest=sha256:edecd67f49ff6fa6f31585cd43d0b6eb1094425f883c6cdc185daa61d241eebe

Observation 73934708-6a89-4f42-919c-c82c3fb29a1e · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Cross-modal Information Flow in Multimodal Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.001598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.001598Z digest=sha256:c817a500cd59cb127a9a6a9d7d322257255646500b0f474c9f787518a4c8c428

Observation 506a9b9d-a8d7-4a18-986f-01af989092a7 · outbound

This paper cites The Llama 3 Herd of Models.

Cross-modal Information Flow in Multimodal Large Language Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.014632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.014632Z digest=sha256:f2e2ff89a825cbad08fea99bd986d22ccb5e59458c154642f7b3f22419a45825

Observation 9743b524-6887-4e46-b4df-1f50d94e4b79 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Cross-modal Information Flow in Multimodal Large Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.010600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.010600Z digest=sha256:c0554cba03f10630110b870c0328007b9373ef5bc09c7e03cd8c217a1f5689cf

Observation 9b3243e9-8b72-4c70-9301-003728575019 · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

Cross-modal Information Flow in Multimodal Large Language Models Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.752851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.022505Z digest=sha256:be1df65c4ab9cff58572914a1d7fcd5626bce32d3315959929248caad7b64dca

Observation fc5d2625-70ea-493e-9483-d3f5209f3c9b · outbound

This paper cites A mathematical framework for transformer circuits.

Cross-modal Information Flow in Multimodal Large Language Models A mathematical framework for transformer circuits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.765687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.018553Z digest=sha256:cdfea6610919714dee185d4261ec38b689c4492ef6077d5ef77c17ed5ec95455

Observation bc3efc44-c022-4ca5-82b5-ef1278bfe58c · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Cross-modal Information Flow in Multimodal Large Language Models Transformer feed-forward layers are key-value memories

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.739359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.030663Z digest=sha256:519e33d6962b83c77172e63a2c86c33a91fdb3ece120f728e484e1ef6c95954f

Observation d646a383-9468-4e36-af97-baa081b32688 · outbound

This paper cites Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.026636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.026636Z digest=sha256:a429db9db455ceac765b5e75a1ffcfc188ee2b603b2f78313def4b30d1d638af

Observation 1ceb7cdd-dd6e-4f22-884e-130efcc4e62c · outbound

This paper cites Probing image- language transformers for verb understanding.

Cross-modal Information Flow in Multimodal Large Language Models Probing image- language transformers for verb understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.726312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.038701Z digest=sha256:7b1ab97300ff2536ce2227dae65f317a7ad161c0ff93ef5b25b0b55b9d50af5a

Observation c0887d75-c0cd-4fce-ba23-d1fb115fd837 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.034653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.034653Z digest=sha256:7454a8ce1760b36ebba81050607026417e1b645e29223f846e88fc8f761e7ddc

Observation 24f55189-75cb-47ef-a182-ab2065133cca · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Cross-modal Information Flow in Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.698375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.046431Z digest=sha256:574e2620dba5e7de29c5d5b523629e11cdd47f4bdac9f776f20c4f82db458da6

Observation b529b4a9-d90d-4394-9e30-6cac4dd11304 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Cross-modal Information Flow in Multimodal Large Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.713166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.042574Z digest=sha256:60600abd400f7cb83e035177cc14e02a0408a3e85c37d4ceb4f4360cc367f5ce

Observation 095b51f7-cf03-4e9c-b171-196e1198c93f · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Cross-modal Information Flow in Multimodal Large Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.054165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.054165Z digest=sha256:40841be00c3945d0fcb8a87a33f13ad5f355478327fc1a3be75a01335aa79f9a

Observation 03c8f297-00e2-4a9e-80f2-7088517572ae · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation.

Cross-modal Information Flow in Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.685237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.050256Z digest=sha256:17960033529f42335046f0f234a27ef07d2585d03119bc09f7e61507c0dbf33a

Observation 3d512324-036e-4835-b169-691bff587c95 · outbound

This paper cites Visual Instruction Tuning.

Cross-modal Information Flow in Multimodal Large Language Models Visual Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.062531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.062531Z digest=sha256:8c5d30948fab9f174d424b41addfc226f1627843d34302bf8955eecd479f3a07

Observation 1f2f1f64-2e6c-44a6-9cc1-e04fdfc19dde · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models, 2024.

Cross-modal Information Flow in Multimodal Large Language Models Mini-gemini: Mining the potential of multi-modality vision language models, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.671695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.058671Z digest=sha256:25fc0e8bd3588e69edb4dbff3674936854600ad5be759fc21bcdbad3392e24a6

Observation 7b909667-9ad6-475c-a12b-1769e563eb92 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Cross-modal Information Flow in Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.642189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.070475Z digest=sha256:0a2ce1eaf1cf7f7066b6b564e7373de092ebf7b7466d52d6a7b9a6db462344a1

Observation 2554fcdb-e476-4b9a-9307-a460cbfcd171 · outbound

This paper cites Improved baselines with visual instruction tuning.

Cross-modal Information Flow in Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.656793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.066709Z digest=sha256:bca4cc6e5ce082ba87c97980e058b8ed27cc310b3489ce93018327d0a4ca3e94

Observation 43bfb556-49b8-4768-93a6-1771ed119ad4 · outbound

This paper cites Locating and editing factual associations in GPT.

Cross-modal Information Flow in Multimodal Large Language Models Locating and editing factual associations in GPT

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.614732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.078402Z digest=sha256:9e99ef7f57e8326905cdf24d6280192c5e0b055be3df4ef82ecb31818d0b2a35

Observation 897b8e6a-1902-476c-820a-b999ced49df6 · outbound

This paper cites Dime: Fine-grained inter- pretations of multimodal models via disentangled local ex- planations.

Cross-modal Information Flow in Multimodal Large Language Models Dime: Fine-grained inter- pretations of multimodal models via disentangled local ex- planations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.628724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.074445Z digest=sha256:f7577cb02197f816d0f4f534a72dad2a2f0c0a9e5a833f0fdcd56e1ef028705e

Observation df9fd07c-8bed-4973-9504-cd4d08f1ae97 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.086378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.086378Z digest=sha256:5ea906b6b8315b39a7d94e53b16aca9fa3e5c28cd823b35d99149f1d42c14017

Observation 583d1950-009e-4941-8ae9-59e0e38da9e1 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Cross-modal Information Flow in Multimodal Large Language Models Progress measures for grokking via mechanistic interpretability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.082279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.082279Z digest=sha256:a07dd84232b25c59ecd6cc0f9d1769454e4c74746c0655d66f79e441c077c8f5

Observation 131e2a97-032c-4d66-b3f7-5af3468968ef · outbound

This paper cites Towards Vision-Language Mechanistic Interpretabil- ity: A Causal Tracing Tool for BLIP.

Cross-modal Information Flow in Multimodal Large Language Models Towards Vision-Language Mechanistic Interpretabil- ity: A Causal Tracing Tool for BLIP

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.586304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.094722Z digest=sha256:69445fcf6730fa1d0e99c2cf4870ecaa1813d1c1dfe5615491c0dc9f70dfc7cf

Observation d8d59e14-b39e-467d-9856-5b30d0efe9d4 · outbound

This paper cites Mechanistic interpretability, variables, and the importance of interpretable bases.

Cross-modal Information Flow in Multimodal Large Language Models Mechanistic interpretability, variables, and the importance of interpretable bases

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.600490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.090735Z digest=sha256:c1c97df9aa5b16b05c96d61da5fdef6a5f01150a6adef3ef4dc33431c1d52f3a

Observation 4df14011-c962-45a4-8808-e5d70d67ecf7 · outbound

This paper cites Are Vision-Language Transform- ers Learning Multimodal Representations? A probing per- spective.

Cross-modal Information Flow in Multimodal Large Language Models Are Vision-Language Transform- ers Learning Multimodal Representations? A probing per- spective

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.564761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.102413Z digest=sha256:01a87659539955aef9214f314a6e397aa8bd92e2dd0959e8ffb42d4640ce2a34

Observation d62d5ceb-a28a-46be-be91-5003edef89a5 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Cross-modal Information Flow in Multimodal Large Language Models Learning transferable visual models from natural language supervi- sion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.098826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.098826Z digest=sha256:8f5ef88df9641a5dcee8ca2f450b7c161348f0cd477686aa3f5a6ecf4fb13af1

Observation 918cbea9-5af6-48d8-8ab0-361f6fec4e68 · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.110272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.110272Z digest=sha256:4d84becc8d937093189fd935535ea8ef3e6213e62a0ee761b9b5620fb280c87c

Observation 6e240497-9f54-4a25-acaf-0bc3eb56ca71 · outbound

This paper cites Multimodal Neurons in Pre- trained Text-Only Transformers.

Cross-modal Information Flow in Multimodal Large Language Models Multimodal Neurons in Pre- trained Text-Only Transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.550805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.106441Z digest=sha256:05dabc378314ee901fe76e56dca8aec35db502de2efaa7db7b702682e42bb47e

Observation 5c458eaa-781c-488f-b47a-84b191c24b86 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Cross-modal Information Flow in Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.118405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.118405Z digest=sha256:398768f3b409abd9dda8a56f68dd4b3d2e5291ca0798a8563bb5a32a0b026bc2

Observation d985c5aa-b104-416c-8262-bb2e933a553b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cross-modal Information Flow in Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.536182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.114376Z digest=sha256:e1f61e2e6796549a0f982af2c5d786f849d4162fbeac31f9f3a18a99f03113c8

Observation dbc8cd6a-4d42-44a7-b9ce-0ab0fcb4713f · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning.

Cross-modal Information Flow in Multimodal Large Language Models Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.126476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.126476Z digest=sha256:00972c5edf0db8cf8fb6e23e7f106be8839d9c899b88f7b882e88651cdadba0c

Observation eb6cd9b0-6eba-4555-9d87-cb8cf918d49d · outbound

This paper cites Attention is All you Need.

Cross-modal Information Flow in Multimodal Large Language Models Attention is All you Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.122548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.122548Z digest=sha256:01ef89030efdf41dcfc77cc77e71c1bb46d56ef6a75f36082adfbfbeb661e9cb

Observation 454bc649-dece-44a8-897c-b51279e05c9c · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Cross-modal Information Flow in Multimodal Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.134669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.134669Z digest=sha256:f4e9b3d6ef05371ebd915eb3e29f51ac02fae762dddae0c205c35beb98ae68d0

Observation 1b5e16aa-6ce4-453b-b9d6-fbc87cb9fb93 · outbound

This paper cites Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models.

Cross-modal Information Flow in Multimodal Large Language Models Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.130548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.130548Z digest=sha256:bb66050226c599b4ac4350fb7756b1cf9f1422fc034f7d9562772492c739f48d

Observation 9393027a-a1a3-4aaf-be1e-faff84f16657 · outbound

This paper cites The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?.

Cross-modal Information Flow in Multimodal Large Language Models The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.142953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.142953Z digest=sha256:6132d760f0b5879ce86b8b754c6bb3ffc4ed502c1d124bbadede87824ae5cb57

Observation c13a0267-c980-4bde-b20f-2cdeacd98c95 · outbound

This paper cites From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.

Cross-modal Information Flow in Multimodal Large Language Models From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:10.138708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:10.138708Z digest=sha256:9e6d14fa5d57db93ed212928ec74f0ae149d78ab78af05f65c5e9c9c0ea10d1e

Observation f7589b04-ade1-42b5-8c32-67636639326b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Cross-modal Information Flow in Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:06:10.513689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:06:10.147882Z digest=sha256:d3ded38360a494de68e764bda0dc29adafac9f6d330a4aa964e7b835bd7c970e

Observation 0cca1222-5a9e-48da-a720-c98c617057bc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Cross-modal Information Flow in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:09.970888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:09.970888Z digest=sha256:47520982577bd3556ae22aced313d5a04b877ceda97ec0a58c592f105687d86a

Pith citing papers

Observation fbc8c65d-34f7-4862-ad73-5e0c6f2c7509 · inbound

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models cites this paper.

Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:02:17.735126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T23:58:57.819555Z digest=sha256:ec281f17870acfcbab88d094fa836329eb459c454a26255c5f74268615b98bd8

Observation d595b4d8-664f-4f4c-a21c-947feaa1fcd5 · inbound

Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation cites this paper.

Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation Cross-modal Information Flow in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:10.027628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:10.027628Z digest=sha256:9e5ca0a06c887d1cdfd559cc5a7a5e9bf483da8db237b2ec513e2c8e12d27681

Observation e90e38c3-498d-4f73-a74f-53f56c143979 · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation Cross-modal Information Flow in Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:26.857624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:26.857624Z digest=sha256:af791e5beedadd60476b659bd32463106305ccc22bf532384dbba89dc55f7173

Observation 03c0985e-2934-4751-9bbb-32103dfcca6e · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.735239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.735239Z digest=sha256:5b6ff3e633f3dc66f8143b8f302820a883771d90346c5a95ea92098b5a5c79b3

Observation 75c6598e-4096-4d67-8213-229f86126856 · inbound

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models cites this paper.

Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:40:05.544824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:40:05.544824Z digest=sha256:dc60beb47496fe40344ead00cce07a6ca156fca0301f13a1e1edae8156b74963

Observation 37595e36-29d1-423e-8221-6cd810049313 · inbound

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models cites this paper.

Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:12.654347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:30:12.654347Z digest=sha256:5ebaaa14a86ef9c96e22f53446a25a71778d7d53dff75ac64e007a28f07db68f

Observation 9f9da4c6-c19e-49d8-91e9-b939a7f244c4 · inbound

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination cites this paper.

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination Cross-modal Information Flow in Multimodal Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:28:05.500293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T06:25:43.369455Z digest=sha256:2f5dee136da4efc2d00ad5a6a01eb65e36645b179f19dd5ef482013b241fd7b5