Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:05:14.472756Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2502.01906.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:05:14.472756Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:36:56.248838Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T14:55:48.338931Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b031cd4-e97c-437e-8af5-2f2722b14244 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adceb6f2-2f75-4290-bab9-57dc011072c9 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e6294c-3685-4764-ba54-1a13bfcc929e · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d3230d-c0d5-4073-8b1f-2d8c5e271075 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e7d843-7f38-41b1-83b1-2eb4595de604 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa9ddf79-68dd-418a-90bc-41a8155328d9 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ff51b0-5a65-494f-881f-e2d71c255733 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 192e3792-14de-446e-9c37-4b189678d361 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Flashattention: Fast and memory-efficient exact at- tention with io-awareness
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cd7b00-ac84-40cc-88ef-97ac92ff8c1d · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e4642b-7869-4e3d-8473-003ef59eb356 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8464cba-c180-44c0-9735-26aad1a99455 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea370fc4-6ddb-4a39-9fae-a7d71f1b706e · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f001ea89-c505-4f05-8e3e-aa88ff44479f · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09fd5c60-7e62-45a1-a745-39ff9e9c0b50 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6589d66-e0a2-47d8-8bc5-26891dcdeee7 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Obelics: An open web-scale filtered dataset of interleaved image-text documents
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e9b0923e-e273-4b18-816d-c3bf85ffb1f5 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models What matters when building vision-language models?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173997d1-1f41-4af4-a811-6f1d730820b3 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e54a25f-f822-42c6-aa14-b949d6a8f2cb · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08265120-3b95-49cb-970a-de14a62c02b2 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223e6291-37bf-4480-abcf-566f856630a3 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417357ba-842f-48af-a660-a44b75d72d6d · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Vila: On pre-training for vi- sual language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d4add187-c6aa-4103-beff-e29fcc2406c8 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 433b69d9-cffa-421c-946e-a6b537f2c9ed · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Visual instruction tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 159aa463-374d-4d50-bf25-dc282f902de0 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models MMBench: Is Your Multi-modal Model an All-around Player?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454579ff-d616-4d8e-a4ae-2d84fc731b64 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2087160c-c19c-4a68-b326-69fbe4f4b50f · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Generation and comprehension of unambiguous object descriptions
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0a838e0a-a8ac-4e61-a576-cb261aadd4cf · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Im2text: Describing images using 1 million captioned pho- tographs
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 298d547f-edea-4791-9c8a-c3972b2f4acd · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Learning transferable visual models from natural language supervi- sion
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f977e5a-79b0-4a48-ba80-7f0ccac233ef · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02a2530b-e778-49ae-a070-33c3f99b6b96 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parame- ters
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee7fec9c-38a3-423e-bd96-273086dfc69c · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7d92c9-5105-4721-ad2a-f8374590bd50 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9960457e-7d6a-4517-b50a-bbd2e35a012f · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44540fb-c5fc-4b17-b7b3-8a751062a856 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba9cbdde-fc6f-4357-a5c5-a4c2309de7d8 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f04384-5289-4c8e-bff2-5b4a4ec89be1 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03869c0e-baec-44e3-8b76-a4125807fb12 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca2618b5-70a5-477b-9beb-1d596975b5e7 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3841db45-5b6e-4a84-b4cb-5a59cc3f7c75 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Attention is all you need
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9376278d-7200-45de-a595-6f5db885139e · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1abfd77-5bdb-42aa-813c-93b7497d5419 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models xgen-mm (blip-3): A family of open large multimodal models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd189c9-7df5-44e4-9779-fc85c463905f · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Sigmoid loss for language image pre-training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5354be34-96a8-4711-963a-93cfce087c8d · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Root mean square layer nor- malization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6dbfcef-af61-40a4-b18e-79bfe137ddfb · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a114ea8-42d5-4c90-a50d-934f99259a63 · outbound
D-Attn: Decomposed Attention for Large Vision-and-Language Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5edf4a9f-7223-4950-9f63-3642631244a3 · inbound
RAVE: Re-Allocating Visual Attention in Large Multimodal Models D-Attn: Decomposed Attention for Large Vision-and-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6161ec78-0fa6-4416-991c-7ef33d749c76 · inbound
RAVE: Re-Allocating Visual Attention in Large Multimodal Models D-Attn: Decomposed Attention for Large Vision-and-Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.