Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:20:02.961462Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2502.06814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:20:02.961462Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d338fb1d-b675-43e3-9579-d667c57762b9 · outbound
Diffusion Instruction Tuning Quantifying Attention Flow in Transformers
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006e4a96-c10a-43d1-a547-73b0e193cf74 · outbound
Diffusion Instruction Tuning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 753a6847-e0fb-4ca6-bff8-6e4114a0a820 · outbound
Diffusion Instruction Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e91e9b-6ceb-4084-807e-582139deaa96 · outbound
Diffusion Instruction Tuning Locality Alignment Improves Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97079387-3cba-4818-9ce4-1de22dc43b31 · outbound
Diffusion Instruction Tuning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd0b7c91-341b-4bd7-b841-e62f05ec2dd8 · outbound
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5dadf8-0c1c-4964-91ca-f7ea139c9adb · outbound
Diffusion Instruction Tuning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86928694-8f44-4b8e-8ac7-171e879ad086 · outbound
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da7329e-dcb0-4ae5-ab1c-9af495ef7fb5 · outbound
Diffusion Instruction Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec41dfa-e504-4b23-810b-c228404d5359 · outbound
Diffusion Instruction Tuning LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4cde6a4-3798-4820-b931-7571094cf985 · outbound
Diffusion Instruction Tuning An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce78c41b-9411-4a30-9193-69c056eb575f · outbound
Diffusion Instruction Tuning Yes I’m not able to provide a name for the person in this picture
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 29ba56bc-ad06-4e2f-9a4f-2abd4030c8c3 · outbound
Diffusion Instruction Tuning A diagram is worth a dozen images
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cef1745-d6ec-4c29-8da3-e693cbf9c072 · outbound
Diffusion Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea88340-2eda-4ee4-b965-2c0c86a1074c · outbound
Diffusion Instruction Tuning HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b1e9d4-cfdb-4221-b7d0-83d0c64e4274 · outbound
Diffusion Instruction Tuning WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a53c59-dd2e-4c85-a3f2-2498162596d4 · outbound
Diffusion Instruction Tuning K., and Chakraborty, A
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedd01ec-d94b-4eb1-82aa-2a9b642f4438 · outbound
Diffusion Instruction Tuning Null-text Inversion for Editing Real Images using Guided Diffusion Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533938b4-b2bb-4b3a-bd32-4cf798a355d5 · outbound
Diffusion Instruction Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee30dcd7-c7ae-46b0-88c6-2244f47e3098 · outbound
Diffusion Instruction Tuning LMFusion: Adapting Pretrained Language Models for Multimodal Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d4ebac-f61e-41de-8b99-dea333610fa5 · outbound
Diffusion Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb3cb22-1e42-4136-b9b9-fa8ed3b8d626 · outbound
Diffusion Instruction Tuning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831adb08-f292-4032-9ce2-52ce6f47f6c8 · outbound
Diffusion Instruction Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93323bd5-84d2-4465-a886-e599e250c5b5 · outbound
Diffusion Instruction Tuning Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcb8bb8-c620-4063-b910-8d2c30a36e4f · outbound
Diffusion Instruction Tuning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2dabd3-6c42-4bed-840f-89c5436cf7a0 · outbound
Diffusion Instruction Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 282b4b10-ff46-44ee-934c-395677431439 · outbound
Diffusion Instruction Tuning MoVA: Adapting Mixture of Vision Experts to Multimodal Context
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 083baf5f-7022-403b-8cd9-cfb287ed4f02 · outbound
Diffusion Instruction Tuning Lavender-Llama3.2-11B occasionally refuses to answer questions for privacy reasons, resulting in a FALSE score and reduced performance on MME as shown in Figure
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94208058-f678-4688-94a6-7e46de1a8519 · outbound
Diffusion Instruction Tuning Rlaif-v: Aligning mllms through open-source ai feed- back for super gpt-4v trustworthiness
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae997530-f339-4175-a6d3-29970b2e5456 · outbound
Diffusion Instruction Tuning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dcd946e-6b12-49be-819c-8c930699e662 · outbound
Diffusion Instruction Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2c2531-220b-4ead-ad40-fabf444c2d55 · outbound
Diffusion Instruction Tuning Generative Visual Instruction Tuning
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f27da9b-4f64-418e-8d25-7da08b19e408 · outbound
Diffusion Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44f9665-ad84-49ad-a442-55c9e4ea713b · outbound
Diffusion Instruction Tuning Unresolved cited work
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c7b5b60e-23e2-47cf-953a-30ef807b3ed0 · outbound
Diffusion Instruction Tuning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 887f7ad3-bf50-4b40-a7bc-c69fb48be241 · outbound
Diffusion Instruction Tuning Evaluating Object Hallucination in Large Vision-Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ca4871-c7c3-4e9d-bf23-56ed5f6ce608 · outbound
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02df6427-1317-4c79-bb21-8109ec65a7bd · outbound
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dae22de9-8599-466a-b913-24ff7e4a6a9d · outbound
Diffusion Instruction Tuning Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.