Pith. sign in

Paper Citation Record · LEDGER

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2502.09925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09925 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:36.710288Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:45:06.207778Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:37.032506Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4db4ffb9-9d64-44c3-a625-712967287eba · outbound

This paper cites GPT-4 Technical Report.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.587149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.587149Z digest=sha256:4b5adf604b2fac9e4feb6c5653f40bc6193599181618b17c7fa0c5a120ae15db

Observation b71ad4a1-42f5-4d43-a009-88110819570d · outbound

This paper cites Language models are few-shot learners.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.598507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.598507Z digest=sha256:104ebdf8835d191b21180b3a081dc679af10fc475da86b91c88141be6e490611

Observation d740592a-1470-40c0-b80b-0fc907a50796 · outbound

This paper cites Additionally, the average performance across 15 benchmarks, excluding MME, increased by approx- imately 1.3 points with CoT.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Additionally, the average performance across 15 benchmarks, excluding MME, increased by approx- imately 1.3 points with CoT

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.887931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.710288Z digest=sha256:faa247c98cddfaef1774bc0545873c996ea5b4e8c103911d1f0e2ab3a933d656

Observation bb33dac4-92e9-4f04-b7cd-85f805478307 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.609590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.609590Z digest=sha256:168a025ac04f64277e127ab2d9dd4ba93335a7176e39588c93ed831328a4e3d9

Observation f006b962-2118-4ea3-b807-24ca6e1b8ff0 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.612990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.612990Z digest=sha256:d676eebd84a8090ece406718710e0771955a8650ad26f12768f22ff10f673db1

Observation 00ce9b05-e205-4785-a849-6b65ef110902 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.984192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.615773Z digest=sha256:66d82078eda7f74c8c521011b20835a76b12d55a87010c290d66ce38d4251f26

Observation 89c64930-2031-4d20-9fa2-62450ea5aefa · outbound

This paper cites GPT-4o System Card.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.618283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.618283Z digest=sha256:b756aeaf6de90500ae47b518c98d64fc700361caea4f610fad19400e65aee753

Observation c8440316-70c4-4b2c-aeea-5864f1099fa8 · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.629857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.629857Z digest=sha256:4a2178b38da9de3363cf40dbef2160967f353726fb693d4ed965820db17cefe2

Observation a2cff039-6b07-4082-a7f2-b173d14f87da · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.633214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.633214Z digest=sha256:37215c3e4e19232cde2796badea7a043cbedba27b24deeaba1b4128faab906e9

Observation 4eb6d2d8-0e6c-4974-bf75-4e4a18535df7 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MMBench: Is Your Multi-modal Model an All-around Player?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.636293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.636293Z digest=sha256:df97794d7cecac8445c09f9f963f3f7c8b98435488f18da49fa2ad1da6a60b38

Observation a1d1e027-957c-42bf-94e0-f94604204e11 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.638524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.638524Z digest=sha256:1ac606fbc25e5cb8036eec40eadba72bbaec354c8c6cbc09f4a116dfe6d26d6f

Observation 8ad6874f-97cc-4519-9e1c-7f851a5ab7af · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Ocr-vqa: Visual question answering by reading text in images

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.969092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.643505Z digest=sha256:7e510bfebe38f2f00e6f38aede454d04a94d2dbf63a4239aa10e0d1882f19569

Observation d0ea9767-8369-471b-89d3-06807f362983 · outbound

This paper cites Vicente Ordonez, Girish Kulkarni, and Tamara Berg.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Vicente Ordonez, Girish Kulkarni, and Tamara Berg

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.961764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.645606Z digest=sha256:835c5108474d46db1e169df28ae93cfd5b357bdeafdcb730310ebe962df07394

Observation 36598b57-5f47-429d-8820-b1cf2a23c9b9 · outbound

This paper cites Solving geometry problems: Combining text and diagram interpretation.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Solving geometry problems: Combining text and diagram interpretation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.952198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.650079Z digest=sha256:44a3b9c012eb44c534687b37c8425dc37dec45c0b85a4f6a442937e04122fa13

Observation 4d15d3ad-3a5f-49ce-92f3-9463fdbcb74d · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.653295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.653295Z digest=sha256:13ce049b0c9fb341656c3278141081b007988cf53314f46636d48d47b252e0a5

Observation 4c38e666-d464-42f6-b926-12bb72dfaba6 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.656360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.656360Z digest=sha256:b5d4eea93f361925cc042d114af010b540d5cdff11de1bde294ab298aac83e3a

Observation 66f82cbf-2116-46a3-bd17-063b44d2e308 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.658902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.658902Z digest=sha256:28162e5f1f1a40c50a84f6a627a19096bc1529a29099ff6035015343e6cc69b6

Observation 8cbb7385-bbcb-425e-bbff-1cbe2f657cfc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.661803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.661803Z digest=sha256:9fa38e96f67eb6940230764ff5cd4fa09b00652c8886690ca91a35797f5e1c4a

Observation 5ce108d0-e454-49f4-87cd-7d8c01268d5d · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.665171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.665171Z digest=sha256:b0fe01e72c166c6cbe7381ceac4409b88b7bae28db3749c04e91712eb907d46e

Observation 65a650b0-96d6-417c-aeb3-4f6ac7e412d7 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.667965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.667965Z digest=sha256:f07acf5a6c4d7cfe7254ac36d0833e658e21d8c55ad0eb67498dc958a4620bd2

Observation 22adde2f-2960-441e-a9bb-fbb12cdc4497 · outbound

This paper cites Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.670967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.670967Z digest=sha256:da752044333fa7ec0b70a69eb159064b7fc31312f7b59332dbd2388d9ed63ec7

Observation fdab4b81-927f-4110-b384-2009ac551252 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Baichuan 2: Open Large-scale Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.673940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.673940Z digest=sha256:2a174f74a37e198d8c572b43bb5505c4af7425f12f272c4e859cc8514ae3721f

Observation 3f161150-a1eb-41c7-aa20-5ac45c0e4fad · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.676780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.676780Z digest=sha256:16d7f5e0c7d5672330f9d39c07c7bbc4a442e109acdfbccbb828208143dceb26

Observation 2c7450c7-b950-4cda-bd3c-b69801092acd · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.680130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.680130Z digest=sha256:33614164bbfe6de2ef81ce21c10a2628d2b4b2f116c815fbeec7336cd7ab45e3

Observation a4454b98-c9c2-4fcd-857e-251e059622f2 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.683745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.683745Z digest=sha256:8658ba63c3ff6390bfb5b728402678313d0eb3ab5ca9de52b0393f520ae7c9c5

Observation 8edb9f04-add9-4cf5-9119-d2de3bb9b0e7 · outbound

This paper cites Retrieval-Augmented Mixture of LoRA Experts for Uploadable Machine Learning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Retrieval-Augmented Mixture of LoRA Experts for Uploadable Machine Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.687453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.687453Z digest=sha256:aaed1619e48f884cce271c539a7df18c223236433e90be885735225364fbf41d

Observation af849d5c-5d51-4666-955a-871ec3ad7d12 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.690405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.690405Z digest=sha256:43ef8c6b8a7d0db461b2a98cc690454a3881cad2578e18096ec2709c0270b599

Observation c18cf347-567f-4610-8eda-40bcd69d31fb · outbound

This paper cites Model tailor: Mitigating catastrophic forgetting in multi-modal large language models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Model tailor: Mitigating catastrophic forgetting in multi-modal large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.941270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.693230Z digest=sha256:d98b40a77641c92d51a3ea8a7a1b102eddbb70e5871cd79fdafd1924ada55041

Observation f224347c-01f6-43d6-802a-4f82d953065c · outbound

This paper cites The approximate data sources and their corre- sponding sample sizes are presented in Table 1 of the main text.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types The approximate data sources and their corre- sponding sample sizes are presented in Table 1 of the main text

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.933258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.696223Z digest=sha256:74218b16fba1a5f524124046be62fdd55deaec88485d41cbf15cc88d18a6537b

Observation b86f269b-e027-408c-81be-60da4a06a5c1 · outbound

This paper cites data fusion,.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types data fusion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.923364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.699123Z digest=sha256:8406b5bfcc77d547e1a63a1eb22c6fa660a66692102828497141042438efe43c

Observation b21b3808-eca5-42b6-9b9d-285b43a72509 · outbound

This paper cites flooding inundates the marina and affects nearby buildings and facilities.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types flooding inundates the marina and affects nearby buildings and facilities

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:06:36.913966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.701887Z digest=sha256:a8f5732b97fdd58c86cbd47b8496c156cb062c74a1bc1962ca829427cb99c522

Observation 01ddcfb7-30c0-4aa1-b4a1-e2331c2a7d38 · outbound

This paper cites an unresolved cited work.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:06:36.896132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.707923Z digest=sha256:43a4a999dece344d56a33be90fd5ec07c2e2cb3937f062b72190e7c18c8d0f12

Observation a6bb92ac-b941-4e67-b564-aa8923970de3 · outbound

This paper cites an unresolved cited work.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Unresolved cited work

Reference 2000

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:06:36.905609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:06:36.705243Z digest=sha256:446bd13d33c94873a26b7e325c6c57d0d136cdaa4ada93a6c6e868d912c74553

Observation 703968e8-f301-40e6-98b0-f0cdc48b472e · outbound

This paper cites A diagram is worth a dozen images.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types A diagram is worth a dozen images

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.623857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.623857Z digest=sha256:a941d4c5f56d178c95fae50aeb7e5b3ccf95bc5ea7bd85d8b6b5ef6379800266

Observation 193ca38e-eb40-44fb-a523-8f2f0805e407 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.627225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.627225Z digest=sha256:22189771edc4ba14ed9040732eaeb0ee9cb7865fca1a2cb678e298967d6172e7

Observation 37e7c896-7c61-416e-86a3-ac736fff99f6 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Referitgame: Referring to objects in photographs of natural scenes

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.621120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.621120Z digest=sha256:b16142dcd881c79c772826146759553018a957108bd5db07716e64853ba72935

Observation 165ab99b-7c93-46c3-b49b-2117d17e4627 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.640824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.640824Z digest=sha256:d94a81bbb774e819ecb4c53f60809413b8a45f37b28ecfef7632116f083894ac

Observation 13a67b69-b0f5-40f0-8346-5d8c6d2de46a · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.601511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.601511Z digest=sha256:18f6c4914c69e01176e76d03899e2683a98dbd6e2bac7fda917c3a5f06cd31cb

Observation d0ffdbe8-93f2-40c9-8df8-093b6ad9a183 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.647480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.647480Z digest=sha256:ce2e58052450ba8e4933483677638f6012596114c812f80090891e7eff723a16

Observation 32db08a6-1578-42f1-88d7-7628405639c0 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.605834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.605834Z digest=sha256:623b37951fd1d2b75d5d26650cf5e0e4156eae6e492cd15b592a91912fc934dd

Observation 74ad1edd-eb65-4678-be2c-1be2619bde63 · outbound

This paper cites Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.595667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.595667Z digest=sha256:2835ff29dacd8cb86bec5d147a53ebacdc9acdbe1c4325fed399e66f4299cd04

Observation 8f80a52c-1705-44dc-9737-1391e0f854d4 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.592185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.592185Z digest=sha256:a13bca40d78786477577ad43d1673248e3a5b485be54fc0f355b0295e1a93bf9

Pith citing papers

Observation 87b4da9b-480f-4490-9f66-1bc60373c293 · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:06.207778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:06.207778Z digest=sha256:cbfe890fa5de9536102b7664a89ec6c9c20bce9886acbeadbd9e114e1b65c8da

Observation 97dacf39-386f-42ba-af0c-d2a9fb255764 · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.290584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.290584Z digest=sha256:fe1127f871d2231548c35891f828cf0db7b7cf256888372b6e9c9134fdf6ea0a

Observation 526c5012-c317-4475-b0f7-ba0feb98be64 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:37.033831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:f173ff08179e5d46269f0591e3b6f06ffb644241f3ff3decd4e55832032713a5