Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Instruction Tuning: A Review and Analysis

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2311.08172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.08172 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:26:28.791677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.807479Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8a249b44-2ff4-4a28-9c62-3e1c21cbc9be · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Vision-Language Instruction Tuning: A Review and Analysis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.211979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:bb0f109ced7e6a6c82358421f9d6555020916903162dcf91a3fbe3e657cafd2e

Observation 7971a455-4b7a-454b-8ed3-10a139425ea5 · inbound

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images cites this paper.

ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images Vision-Language Instruction Tuning: A Review and Analysis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:26:28.791677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:26:28.791677Z digest=sha256:7cd9670e98f8743a65a293c9f2b4486da1f5900c806cd7ec8d7dccc93ca676cd

Observation cb9a063c-f64c-4590-92f9-714188660197 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Vision-Language Instruction Tuning: A Review and Analysis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:39:40.581274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:0628e21f19d04bdaa04342ec8aa06b695de269fcb8f9fbe70b7660028bf9ba16

Observation 10edc680-a24f-46b0-aa44-b05557266594 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Vision-Language Instruction Tuning: A Review and Analysis

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.581264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:19:09.596950Z digest=sha256:e768b36fa06bcac668edbf62d05f48fd6c01e68712e8428ad2272d40f48b16f6

Observation 74733c52-4171-4b2c-8d07-b217c8e07e0b · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Vision-Language Instruction Tuning: A Review and Analysis

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.808826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:cf0946b2f2e03bff93f72686118a16f2818481054e50c2ef2449ce9aa76ec18d

Observation 43d818ca-0e90-49e2-9a28-0f5ba7eabe7e · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Vision-Language Instruction Tuning: A Review and Analysis

Reference 185

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:d84263c2de228b357a177a7a760edb822aaf275c02745602e3f92f026aa7fdc4