Pith. sign in

Paper Citation Record · LEDGER

VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2312.09251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.09251 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:01:12.829064Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T22:48:36.084743Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5928f9cd-bb60-4b94-a0f9-ed2f6486b02a · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.089738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:0d58f0e0debf44926e51e0eb5c3670fc63be7ceb1b2ccbfab0dc6dd31aa86938

Observation 517844bd-825e-4b7e-bd36-135d6f994670 · inbound

EventGPT: Event Stream Understanding with Multimodal Large Language Models cites this paper.

EventGPT: Event Stream Understanding with Multimodal Large Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T05:01:12.829064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:01:12.829064Z digest=sha256:fbb6d6d822f93f63908de2016fdf6703b1f9c7c1258e9554bb1a03f9748ed7c0

Observation 226b838e-729e-416e-b6dd-31d478af8707 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.634738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.634738Z digest=sha256:a1b4f9e4ce36a4f9bad67b47f3630599335e534084e5329887acebad23f1d10b

Observation a92aea54-cee4-49cf-a20e-bbf5f78850c7 · inbound

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models cites this paper.

Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:14:56.605430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:14:56.605430Z digest=sha256:6d690f932373d33dae3b620e0d07056eb80f2550270a7107906a420d144074c9

Observation ba0cdd7b-2591-4cd2-b95f-5bcce834c269 · inbound

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding cites this paper.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.118464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.118464Z digest=sha256:996952930d1f6dfac9d04f4baf648354425b754eb7e6f0eee1519fb17f56b8cb

Observation 2020f042-2401-46d7-aa31-dfdfa02ce3c7 · inbound

Olympus: A Universal Task Router for Computer Vision Tasks cites this paper.

Olympus: A Universal Task Router for Computer Vision Tasks VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.802539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.802539Z digest=sha256:7d495f8ee8f6e0b19560e34a6b7c105d0c2dc8e7b507842dc4d0896ea6eed597

Observation b9a7e13b-ff2b-46b4-aeb5-c499a095d2dd · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.771587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.771587Z digest=sha256:13ee446054fc4e6c0283d403e33da63b6e5a5b17608816346049303b63c6f79d

Observation e600f1ea-3796-4743-af71-30635ca88f4f · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.478614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.478614Z digest=sha256:a043a5e6be33ba4ff0e449c744f53969b35d71e4da62865f498f44cb064f94cd

Observation 0d456aa7-05ec-41fe-a08f-a3a0500c832c · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.848027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:e3d8877d860f143057ea7de45fa4810116aa86d10ebf9d068b85f35107706c94

Observation 9c091c59-76f3-4992-bda8-50607a9d87c7 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.154340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:12.154340Z digest=sha256:02f4d361df9b3b0da9c8c7e114992d75976d10d7146981d36fc9413565ea5a73

Observation f7f5f831-62fd-475b-bc0d-3474d9e72514 · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:20.743420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:20.743420Z digest=sha256:01ad3eb5981daa9f116da68bc51bf6a0c0b73600a1dbc8198bfe2126e7bdc0b2

Observation 88709946-dcf2-487f-870a-4537ffaccfc5 · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:57.735425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:57.735425Z digest=sha256:bb7f27f48d4c0954c0bcf18d66436dde0406fa06a91795dd272e595a37223791