Pith. sign in

Paper Citation Record · LEDGER

SVIT: Scaling up Visual Instruction Tuning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2307.04087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.04087 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:32:42.920231Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:03.935337Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d649401e-77da-4d9e-89be-f84674c165fd · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning SVIT: Scaling up Visual Instruction Tuning

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.918523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:43ac7d8194f2128b58d04b482aeeccbdab510408f9d0a19fc32f01b51e763576

Observation 536d51be-36d6-4cb8-99dd-54a5b20de514 · inbound

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets cites this paper.

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets SVIT: Scaling up Visual Instruction Tuning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:14:10.262530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T08:13:18.191233Z digest=sha256:760d72c01417f578a9d0fe7cce877c54c89131056e5db5c253f7cf96021a62b6

Observation 10ba115d-d398-4bb6-99ff-8555551e5cff · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning SVIT: Scaling up Visual Instruction Tuning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.990145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:f64e7d04ddf89d5747361658faf4f34fbfab065a5e774d5426cdbf5c3de827e7

Observation 8a1aeacf-975c-4c5d-9467-ad30e2b5fedc · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration SVIT: Scaling up Visual Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.793866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:f64e7413c24a2af7b577fcea6ed191f945973d85aa951afc7900c3efc899eb29

Observation d5bd206d-b5b7-4509-850d-b019a7a13b6b · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI SVIT: Scaling up Visual Instruction Tuning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.761969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:19301f7bf7058ada8d6e2868061ee8454430b9e382dff0a4fd4e007400d756b6

Observation 819c52b1-606b-467a-9204-110350fc9fd3 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks SVIT: Scaling up Visual Instruction Tuning

Reference 184

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.262028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:8517f8f3ecc4d0abed4b22001e526eba4b3dacd5f1d09bea035e31417dc2c419

Observation 60743011-f516-4957-8578-263f3ccdb2dc · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model SVIT: Scaling up Visual Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.136023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:44b1378d19f092022f8024b1f2938174ae4140e70c48ced951cc7d8554a21f21

Observation bb44d84a-7d91-45f8-a19d-299aaddecf62 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SVIT: Scaling up Visual Instruction Tuning

Reference 132

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.413562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:b659615146d48253d9296d1db9cf8d9924caaf0e93830b99160885ccda93cc22

Observation 23f91ff9-793f-4653-aeb7-d5669ae8fd77 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites SVIT: Scaling up Visual Instruction Tuning

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.314247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:5d58a9afcaef07bfa6bf16572533a1d9a804da566593c30355f190691a960a9c

Observation 9a92866f-6f73-46a2-8710-fe58ee55fee2 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone SVIT: Scaling up Visual Instruction Tuning

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.159977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:a6d4fa1a1afcbbec5de5cb7eed9c039bd76e714d8c967e75023893914ca09f17

Observation d1a283ce-0177-4dc8-a8da-bf383082fd9a · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:20:36.368777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:3fc90a984e8f1dd36efb316c76cd9ce07bf633e3007b856858d4b8e166f72745

Observation 18ffe440-31d2-4533-a0b3-3d5af3f2de93 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark SVIT: Scaling up Visual Instruction Tuning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:51:48.483233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:789a9206ab4a3590bbabbb07fded7adce307b7c69087b1d55fc2a76334417ffa

Observation 9176bb75-69a3-44f5-b177-b97c2950872f · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.968159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:3f287241fea722f917ce1273a49372dc3ba33d39ba9760251dfbc8ff77856804

Observation 2c31d624-3712-4f64-bc7c-72ca9ff80519 · inbound

DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities cites this paper.

DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities SVIT: Scaling up Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:42.920231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:42.920231Z digest=sha256:6879cefd0cc9823c4569f6e5bc85cd0aa57972d68c15e4a20f4271f14508e040

Observation 13722c96-2eb2-41b1-a5c3-905c9dec1085 · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM SVIT: Scaling up Visual Instruction Tuning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:10:56.208834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:6589698e71fa3a10ea5ce9e49c189a3a8097bc713c230fd7744333b1c39a384c

Observation 05c6122e-8d8d-44a8-bd5b-e427e1374c22 · inbound

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion cites this paper.

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SVIT: Scaling up Visual Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:54.331667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:54.331667Z digest=sha256:f08fb880a3ad2753d1d19eaceb993b9be209963ccdb123d16a1ab6105cb4a1d8

Observation 1f62fce6-9aa9-49d0-be01-6a51329e294c · inbound

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models cites this paper.

EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:34.070554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:34.070554Z digest=sha256:455b39f148f92ac4e39edf08dd46a61be94f3a7d1fd417b9067149580b3394a0

Observation d3be2cd5-5769-47ea-b6ce-7c20742fe64a · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities SVIT: Scaling up Visual Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.746047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.746047Z digest=sha256:62f46ff63fec4e5c8bcebcaa30de8e0e102d08a210e4728b29fd75937aa69ed2

Observation d86c7702-3db9-44f1-ab53-65377a19e076 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:05.494156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:05.494156Z digest=sha256:8c2cad6626c8020495a4e551e47f5fb284b90b4e1b64df67b758e5a3fbaa95c2

Observation c04e7633-c62d-43a3-ba52-725bb69e0a6d · inbound

GLAD: Generalizable Tuning for Vision-Language Models cites this paper.

GLAD: Generalizable Tuning for Vision-Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:23.687949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:23.687949Z digest=sha256:e195b10576c02a54f9ff17ed439310ed6eb08183aa083d447f57199d839ffbea

Observation 5d3013a8-c5cd-43e0-bdfe-d4519226568b · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark SVIT: Scaling up Visual Instruction Tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:08.100077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:08.100077Z digest=sha256:b1a3e351d441a746839573458d78b2e82d26788e60b8b645e2ccb6f3412fc1f5

Observation bdb9d7e0-39e1-48e3-879a-fda3f07ec52d · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers SVIT: Scaling up Visual Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.703738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.703738Z digest=sha256:4a3557690665b7c195214c9a0048a0e50a93f11e6532a365c711e1f62339bf5b

Observation dc141185-86d2-427e-8b9e-36bf929837b3 · inbound

Representation learning from OCT images cites this paper.

Representation learning from OCT images SVIT: Scaling up Visual Instruction Tuning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:43.123677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:32:13.800433Z digest=sha256:05672a40d9d10bd97a8188e483c732862f251b90990971cea195d73639f62ce5

Observation a0308b5f-a2a2-417c-b789-1feb6cd8f6f3 · inbound

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models cites this paper.

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models SVIT: Scaling up Visual Instruction Tuning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:30.831974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T04:07:58.031100Z digest=sha256:94f55bffa3e3f104612345cb51eb7ed3fa69614001e690b3db4a5843ed73df76

Observation f63dd66d-a399-45fe-96fa-9ec78afd9d6b · inbound

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens cites this paper.

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens SVIT: Scaling up Visual Instruction Tuning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:09:38.498516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:07:12.965100Z digest=sha256:2cdf21bd03039864044389fa7df91927bd19014bf1b6b86f8cb60760ccacaf12

Observation 9d25ec62-a039-4f2b-8215-c6e822ac182c · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization SVIT: Scaling up Visual Instruction Tuning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.411770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:93e44a6af4ac815b75db90a3cb97563b012f3bc4da93b4a64700af14d90593db

Observation db848888-a092-4d1f-9fa6-ed77bb9ad570 · inbound

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning cites this paper.

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning SVIT: Scaling up Visual Instruction Tuning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:03.936899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T15:04:18.880016Z digest=sha256:2a36b2019af6568f015ad6996b71224ddf8d72c8e3c107880e92e49086583923

Observation fe6631cc-985d-41ee-a96b-da517e43bb84 · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design SVIT: Scaling up Visual Instruction Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.447932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.447932Z digest=sha256:193c192d379a41000bcbf2d85a9180f9237a824ea0cf7857fbafebb099e307aa