Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:06:10.147882Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2411.18620.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:06:10.147882Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:44:10.027628Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T00:02:17.731733Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 94d4e1d1-55df-4aef-a79d-d5fdbb77f07b · outbound
Cross-modal Information Flow in Multimodal Large Language Models https : / / www
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b93005f4-16e5-4816-8d8c-3f3374ffa350 · outbound
Cross-modal Information Flow in Multimodal Large Language Models https: //huggingface.co/lmms- lab/llama3- llava- next-8b, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 656b0e57-6fb0-4515-80e9-c2ccddac5ba5 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d663816d-610a-4413-b229-5146f384283c · outbound
Cross-modal Information Flow in Multimodal Large Language Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35af6b75-8de0-45d1-a119-2ea0f693bbd5 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b5ca4c8-7f59-4eb3-a0c7-10259dca8f40 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Understanding Information Storage and Transfer in Multi-modal Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5679eee8-5e35-46a9-8d91-13f3d1c77669 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a4c46df5-8b62-4585-9e3b-cf128bfcc9dc · outbound
Cross-modal Information Flow in Multimodal Large Language Models Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ac4c82cb-e752-4536-b798-a8e0067ceac1 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Probing multimodal embeddings for linguistic properties: the visual-semantic case
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4bfa2db5-488c-44fb-baad-05e8fa16563f · outbound
Cross-modal Information Flow in Multimodal Large Language Models Knowledge Neurons in Pretrained Transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c359de28-3689-433e-a264-2f747cfb7911 · outbound
Cross-modal Information Flow in Multimodal Large Language Models InstructBLIP: Towards General- purpose Vision-Language Models with Instruction Tuning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ddc5ead-d7ef-44f1-805f-914c96ea8c55 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Analyzing Transformers in Embedding Space
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73934708-6a89-4f42-919c-c82c3fb29a1e · outbound
Cross-modal Information Flow in Multimodal Large Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506a9b9d-a8d7-4a18-986f-01af989092a7 · outbound
Cross-modal Information Flow in Multimodal Large Language Models The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9743b524-6887-4e46-b4df-1f50d94e4b79 · outbound
Cross-modal Information Flow in Multimodal Large Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3243e9-8b72-4c70-9301-003728575019 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Eva: Exploring the limits of masked visual representa- tion learning at scale
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fc5d2625-70ea-493e-9483-d3f5209f3c9b · outbound
Cross-modal Information Flow in Multimodal Large Language Models A mathematical framework for transformer circuits
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc3efc44-c022-4ca5-82b5-ef1278bfe58c · outbound
Cross-modal Information Flow in Multimodal Large Language Models Transformer feed-forward layers are key-value memories
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d646a383-9468-4e36-af97-baa081b32688 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ceb7cdd-dd6e-4f22-884e-130efcc4e62c · outbound
Cross-modal Information Flow in Multimodal Large Language Models Probing image- language transformers for verb understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c0887d75-c0cd-4fce-ba23-d1fb115fd837 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f55189-75cb-47ef-a182-ab2065133cca · outbound
Cross-modal Information Flow in Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b529b4a9-d90d-4394-9e30-6cac4dd11304 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 095b51f7-cf03-4e9c-b171-196e1198c93f · outbound
Cross-modal Information Flow in Multimodal Large Language Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c8f297-00e2-4a9e-80f2-7088517572ae · outbound
Cross-modal Information Flow in Multimodal Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Uni- fied Vision-Language Understanding and Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3d512324-036e-4835-b169-691bff587c95 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Visual Instruction Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f2f1f64-2e6c-44a6-9cc1-e04fdfc19dde · outbound
Cross-modal Information Flow in Multimodal Large Language Models Mini-gemini: Mining the potential of multi-modality vision language models, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7b909667-9ad6-475c-a12b-1769e563eb92 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2554fcdb-e476-4b9a-9307-a460cbfcd171 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Improved baselines with visual instruction tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 43bfb556-49b8-4768-93a6-1771ed119ad4 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Locating and editing factual associations in GPT
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 897b8e6a-1902-476c-820a-b999ced49df6 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Dime: Fine-grained inter- pretations of multimodal models via disentangled local ex- planations
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation df9fd07c-8bed-4973-9504-cd4d08f1ae97 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583d1950-009e-4941-8ae9-59e0e38da9e1 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Progress measures for grokking via mechanistic interpretability
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131e2a97-032c-4d66-b3f7-5af3468968ef · outbound
Cross-modal Information Flow in Multimodal Large Language Models Towards Vision-Language Mechanistic Interpretabil- ity: A Causal Tracing Tool for BLIP
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d8d59e14-b39e-467d-9856-5b30d0efe9d4 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Mechanistic interpretability, variables, and the importance of interpretable bases
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4df14011-c962-45a4-8808-e5d70d67ecf7 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Are Vision-Language Transform- ers Learning Multimodal Representations? A probing per- spective
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d62d5ceb-a28a-46be-be91-5003edef89a5 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Learning transferable visual models from natural language supervi- sion
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918cbea9-5af6-48d8-8ab0-361f6fec4e68 · outbound
Cross-modal Information Flow in Multimodal Large Language Models LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e240497-9f54-4a25-acaf-0bc3eb56ca71 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Multimodal Neurons in Pre- trained Text-Only Transformers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c458eaa-781c-488f-b47a-84b191c24b86 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d985c5aa-b104-416c-8262-bb2e933a553b · outbound
Cross-modal Information Flow in Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dbc8cd6a-4d42-44a7-b9ce-0ab0fcb4713f · outbound
Cross-modal Information Flow in Multimodal Large Language Models Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb6cd9b0-6eba-4555-9d87-cb8cf918d49d · outbound
Cross-modal Information Flow in Multimodal Large Language Models Attention is All you Need
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454bc649-dece-44a8-897c-b51279e05c9c · outbound
Cross-modal Information Flow in Multimodal Large Language Models OPT: Open Pre-trained Transformer Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5e16aa-6ce4-453b-b9d6-fbc87cb9fb93 · outbound
Cross-modal Information Flow in Multimodal Large Language Models Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9393027a-a1a3-4aaf-be1e-faff84f16657 · outbound
Cross-modal Information Flow in Multimodal Large Language Models The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13a0267-c980-4bde-b20f-2cdeacd98c95 · outbound
Cross-modal Information Flow in Multimodal Large Language Models From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7589b04-ade1-42b5-8c32-67636639326b · outbound
Cross-modal Information Flow in Multimodal Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0cca1222-5a9e-48da-a720-c98c617057bc · outbound
Cross-modal Information Flow in Multimodal Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc8c65d-34f7-4862-ad73-5e0c6f2c7509 · inbound
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d595b4d8-664f-4f4c-a21c-947feaa1fcd5 · inbound
Interpreting Social Bias in LVLMs via Information Flow Analysis and Multi-Round Dialogue Evaluation Cross-modal Information Flow in Multimodal Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90e38c3-498d-4f73-a74f-53f56c143979 · inbound
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation Cross-modal Information Flow in Multimodal Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c0985e-2934-4751-9bbb-32103dfcca6e · inbound
GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Cross-modal Information Flow in Multimodal Large Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c6598e-4096-4d67-8213-229f86126856 · inbound
Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37595e36-29d1-423e-8221-6cd810049313 · inbound
Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models Cross-modal Information Flow in Multimodal Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9da4c6-c19e-49d8-91e9-b939a7f244c4 · inbound
Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination Cross-modal Information Flow in Multimodal Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.