Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:36.710288Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2502.09925.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:36.710288Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:45:06.207778Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:27:37.032506Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4db4ffb9-9d64-44c3-a625-712967287eba · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b71ad4a1-42f5-4d43-a009-88110819570d · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Language models are few-shot learners
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d740592a-1470-40c0-b80b-0fc907a50796 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Additionally, the average performance across 15 benchmarks, excluding MME, increased by approx- imately 1.3 points with CoT
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb33dac4-92e9-4f04-b7cd-85f805478307 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f006b962-2118-4ea3-b807-24ca6e1b8ff0 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ce9b05-e205-4785-a849-6b65ef110902 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89c64930-2031-4d20-9fa2-62450ea5aefa · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types GPT-4o System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8440316-70c4-4b2c-aeea-5864f1099fa8 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2cff039-6b07-4082-a7f2-b173d14f87da · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eb6d2d8-0e6c-4974-bf75-4e4a18535df7 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MMBench: Is Your Multi-modal Model an All-around Player?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1d1e027-957c-42bf-94e0-f94604204e11 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad6874f-97cc-4519-9e1c-7f851a5ab7af · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Ocr-vqa: Visual question answering by reading text in images
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0ea9767-8369-471b-89d3-06807f362983 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Vicente Ordonez, Girish Kulkarni, and Tamara Berg
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36598b57-5f47-429d-8820-b1cf2a23c9b9 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Solving geometry problems: Combining text and diagram interpretation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4d15d3ad-3a5f-49ce-92f3-9463fdbcb74d · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c38e666-d464-42f6-b926-12bb72dfaba6 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f82cbf-2116-46a3-bd17-063b44d2e308 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbb7385-bbcb-425e-bbff-1cbe2f657cfc · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types LLaMA: Open and Efficient Foundation Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ce108d0-e454-49f4-87cd-7d8c01268d5d · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65a650b0-96d6-417c-aeb3-4f6ac7e412d7 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22adde2f-2960-441e-a9bb-fbb12cdc4497 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdab4b81-927f-4110-b384-2009ac551252 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Baichuan 2: Open Large-scale Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f161150-a1eb-41c7-aa20-5ac45c0e4fad · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7450c7-b950-4cda-bd3c-b69801092acd · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4454b98-c9c2-4fcd-857e-251e059622f2 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8edb9f04-add9-4cf5-9119-d2de3bb9b0e7 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Retrieval-Augmented Mixture of LoRA Experts for Uploadable Machine Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af849d5c-5d51-4666-955a-871ec3ad7d12 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c18cf347-567f-4610-8eda-40bcd69d31fb · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Model tailor: Mitigating catastrophic forgetting in multi-modal large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f224347c-01f6-43d6-802a-4f82d953065c · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types The approximate data sources and their corre- sponding sample sizes are presented in Table 1 of the main text
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b86f269b-e027-408c-81be-60da4a06a5c1 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types data fusion,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b21b3808-eca5-42b6-9b9d-285b43a72509 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types flooding inundates the marina and affects nearby buildings and facilities
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01ddcfb7-30c0-4aa1-b4a1-e2331c2a7d38 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6bb92ac-b941-4e67-b564-aa8923970de3 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Unresolved cited work
Reference 2000
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 703968e8-f301-40e6-98b0-f0cdc48b472e · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types A diagram is worth a dozen images
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 193ca38e-eb40-44fb-a523-8f2f0805e407 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e7c896-7c61-416e-86a3-ac736fff99f6 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Referitgame: Referring to objects in photographs of natural scenes
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165ab99b-7c93-46c3-b49b-2117d17e4627 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a67b69-b0f5-40f0-8346-5d8c6d2de46a · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MapQA: A Dataset for Question Answering on Choropleth Maps
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0ffdbe8-93f2-40c9-8df8-093b6ad9a183 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32db08a6-1578-42f1-88d7-7628405639c0 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ad1edd-eb65-4678-be2c-1be2619bde63 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f80a52c-1705-44dc-9737-1391e0f854d4 · outbound
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b4da9b-480f-4490-9f66-1bc60373c293 · inbound
Kwai Keye-VL Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97dacf39-386f-42ba-af0c-d2a9fb255764 · inbound
Kwai Keye-VL 1.5 Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526c5012-c317-4475-b0f7-ba0feb98be64 · inbound
Kwai Keye-VL-2.0 Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.