Pith. sign in

Paper Citation Record · LEDGER

Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2312.16602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16602 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:14.200042Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T17:41:53.208887Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab626ba4-9043-4e99-b595-da6c85bd7601 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.983223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.983223Z digest=sha256:24c181aade70732a85ea64f782e75e68f32e4d2d5741d5190ec39a5b59e9aa11

Observation 76c474ea-bf81-4d73-8188-d4bcdcaed8fa · inbound

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces cites this paper.

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:08.385867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:08.385867Z digest=sha256:7d07742c2547533e4e981b03bf2f3631eb0fd2450c26db0a112160247f8dca5a

Observation c881b9e8-5aee-47a1-bd0e-54d4f3633b19 · inbound

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model cites this paper.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.001774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.001774Z digest=sha256:b9ec9c89040df671d6017b3df5cf9e55b3493999fb8fea8406806cc700ed8beb

Observation eaa950c7-14d8-48dc-ac02-17e37d715c15 · inbound

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization cites this paper.

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:14.200042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:14.200042Z digest=sha256:2e5232732f75deede3735c226ce731c40e52483f4e4ce9b100e9efc109b2fffa

Observation 8cc1b633-816b-4697-8e7e-b6c08a0c6574 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 233

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:01.149070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:01.149070Z digest=sha256:f28e3f658a5a20fc324d343a564b61edeac7d005c6c3ac1922c929179e291858

Observation 258c1744-55b4-4d36-945f-0568c0653ae1 · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.212540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:2d16d235b1c85b24ad84a6f1a1e85dbecf00b62a704ca11705f5587893c8b2b7

Observation 29cbdcbe-db29-490d-80b9-d7b74546e00e · inbound

Mitigating Image Captioning Hallucinations in Vision-Language Models cites this paper.

Mitigating Image Captioning Hallucinations in Vision-Language Models Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:59.565756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:56:59.565756Z digest=sha256:15e63d3fec93ac38f0e3249e288e8c391afbd13363832f7caf8b91b18a435a17

Observation 211c8548-f824-4440-b349-4ea4f90927ae · inbound

Towards Unconstrained Human-Object Interaction cites this paper.

Towards Unconstrained Human-Object Interaction Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:30.381115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:17:06.952676Z digest=sha256:f75c0c71571e701e157fa2c0a5b7f597a500bea58cf9553f2f08cad5dea762f9

Observation e9fb204d-9419-4e09-8c04-cceeb7b8bc0c · inbound

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models cites this paper.

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:40.845800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T19:22:43.082083Z digest=sha256:ee70091c1a34b2e3ccca62594d85298fef5d7331cec35888a6d678b4b7f46d35

Observation 9319d652-9a93-454d-abae-1b934d7a731e · inbound

Learning in Deep Networks under Dale's Constraint cites this paper.

Learning in Deep Networks under Dale's Constraint Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 228

Resolution
unresolved
no resolver link, observed 2026-08-10T17:39:43.306915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:39:43.306915Z digest=sha256:a36f7fe67e3beef7e80618509aa1c8ebaa128754787140d3e8d14217bd3bdc7b