Pith. sign in

Paper Citation Record · LEDGER

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

As of 21 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 22 inbound Pith citation observations for arXiv:2501.12327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12327 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:21:04.547305Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:48:03.500362Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.542063Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02ef5ae7-22ec-402d-8993-4ca00f815ea1 · outbound

This paper cites Albergo and Eric Vanden-Eijnden.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Albergo and Eric Vanden-Eijnden

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.055182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.055182Z digest=sha256:1af1513810f1303cfdac95343b65dfbe62ad04a9be58ab7ff5e049906e02e032

Observation a7091db3-b8a5-4832-b0ac-8719fe9ba569 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.060897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.060897Z digest=sha256:2d768992405b75617f4272d6fa553b3f1ddb58d4de24ab584f6300af79eed1ae

Observation 687c7d63-90b1-4a60-b1ea-791b5c9e0375 · outbound

This paper cites messages.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model messages

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.066545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.066545Z digest=sha256:f3b7bdf5dcd96724c9f57044b93265fdc3993d0aca50f194017221e6c0224a2d

Observation 78110544-edc4-4cf0-a5c4-ebec50a60318 · outbound

This paper cites Analytic- dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Analytic- dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.071756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.071756Z digest=sha256:87a27a8aed7c5029a3e0e5f1022c211be2dc29a51eac87d2e58539792626f6e9

Observation 1ec4130f-6976-4b60-ac31-064da350ec4b · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.076616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.076616Z digest=sha256:83b5d7eb01b4654474949834712849dcabfa323a2cd87174d9117e112fd16ed8

Observation c19ee446-7ef1-4332-9b85-34e2226e1571 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.081816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.081816Z digest=sha256:a1dae290a84e02be4a92e70488c89dba57415edb4fabf83650ad9a53998bbd38

Observation eaf80934-d9d7-4a17-8708-d5024abaf709 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.087133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.087133Z digest=sha256:7ebf64f46ee26bd8e17d8cfa2d44eccc640fc324d1481edbe9336ae21a80fc17

Observation e00e5a9b-bcb2-493e-b511-a2ae577baac3 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.092153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.092153Z digest=sha256:f50f0fe4d16a8f2fe6284911a9442989312b6d9787022c7bda48771bd1a6ab3c

Observation c819938f-f686-4df2-bff3-4272318f25c8 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.097084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.097084Z digest=sha256:122ccca223842330cacd6d3e4f562addd5fac0184ab9e7678728056feaa2248e

Observation 7e567638-7001-4bb0-a5a9-4c0278796200 · outbound

This paper cites Deepseek-v3 technical report, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Deepseek-v3 technical report, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.107444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.107444Z digest=sha256:b96111410aa199d58b6371a37ac27bfee6607b755c04079b164ee89aa28b7044

Observation cc5976e8-5f6c-4489-8476-096b526b951e · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.113109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.113109Z digest=sha256:e01e16346a970247ecdd2b2abe16576b08bdaab8f70b3c312c7c66396eca6d91

Observation ccaabffa-b2e6-4969-a60d-0cf3c5f7e961 · outbound

This paper cites Cogview: Mastering text-to- image generation via transformers, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Cogview: Mastering text-to- image generation via transformers, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.117969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.117969Z digest=sha256:59eb4ec95dfec2ba4b6f2de647c0f46afabb6db53f78f1663df3ca85417d0917

Observation e2ba1748-3173-4c43-9221-9c2e8c359b3c · outbound

This paper cites Dreamllm: Synergistic multimodal compre- hension and creation, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dreamllm: Synergistic multimodal compre- hension and creation, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.122884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.122884Z digest=sha256:e461d5a926b6837efdb890711a5bfe9dc746e681945570ccda0594e952bd29cf

Observation 8187f5d8-4eae-4649-b2f8-30c8dfab353a · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Taming transformers for high-resolution image synthesis, 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.127322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.127322Z digest=sha256:584c3fd5a33dbe27bea1335abe5a89253f037330c61ba512124c88d301b78c51

Observation 6b7f01ba-b9f6-4e40-a94a-3baf2bb28b9f · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.132171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.132171Z digest=sha256:f1e0f06583c78f743aa18902d10aea304cfcbd80d0169f72b541acbb1c4b6505

Observation 4b049a85-1fee-4821-b33f-9a53edb8b170 · outbound

This paper cites Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.137915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.137915Z digest=sha256:54ab4a9c78d1c41cffc221d8e3a2e807e7151006eb1e7a03e02bd87266481032

Observation 7db3e256-48c4-40de-8a5e-8aa20d838f68 · outbound

This paper cites Making llama see and draw with seed tokenizer, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making llama see and draw with seed tokenizer, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.142561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.142561Z digest=sha256:c437a49eaee410258f4e9980aa7c133505d6f4672124732a537a43368b71f7fa

Observation c7e95b52-af29-49cd-8716-f2adb530d859 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.147572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.147572Z digest=sha256:cc627b1172dffea670399a97ed7919496abf856374fab10d4a4a1c0986b6e90c

Observation 395f0c5e-97e6-4b70-9bcb-f6804c9c849e · outbound

This paper cites Making the v in vqa matter: Ele- vating the role of image understanding in visual question answering, 2017.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making the v in vqa matter: Ele- vating the role of image understanding in visual question answering, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.152825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.152825Z digest=sha256:cd358d9ed9b69c939ef43cc0030f5bcb4b98dc656964e4f9cff4cf12c0d7a42f

Observation ad7bb450-fe95-49f0-bc2c-078aa624bb27 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.157578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.157578Z digest=sha256:4fc8fffb1973b00f8c67cbb74fd6548c503661efb06b91a9d196e84b71d582ac

Observation de348b17-ce73-4b05-a90d-647049fd6fb4 · outbound

This paper cites Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.162915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.162915Z digest=sha256:b01dd98a4faaaf96c37561f44d3c687a021ea515e72a0afc61b0785f9d590065

Observation fd96de1b-a8e3-458e-94ae-45c05faaf459 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.167558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.167558Z digest=sha256:ebf068d5e1dbeea2cb7af10b34c323021ac415247e3d80f17adc897cf29bb72b

Observation 83199dc8-7144-4c07-b62f-6f31345dea6e · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling Laws for Autoregressive Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.172479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.172479Z digest=sha256:519e76b9deda33b5e4190ce5874b728df6a8ad43d388f0d69e27e7d9ffae8f9b

Observation 6a913b5d-76b8-4ff4-9ac7-5456d24a75dd · outbound

This paper cites Denoising diffu- sion probabilistic models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denoising diffu- sion probabilistic models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.177268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.177268Z digest=sha256:a38e5d237fed4c49286b698bef4533eaf90a582a423f6d90f4bd46333210a9df

Observation e14c0d7f-ef16-40b3-893b-5cb64c3118df · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denoising diffu- sion probabilistic models, 2020

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.181964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.181964Z digest=sha256:ea8e16ac8d7dc496c684f1c0b785a500fe2e914e9cb62eb370857913c1f0cafb

Observation c710cb36-5d00-4c49-99b7-f9ba42f719a9 · outbound

This paper cites Fleet, Mohammad Norouzi, and Tim Salimans.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Fleet, Mohammad Norouzi, and Tim Salimans

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.186599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.186599Z digest=sha256:9e01c708831ae887adee69b86f8bb9822da80766e51acb2fd5aa650c1fa63f83

Observation 05c48a21-3f0f-415d-b190-9e083acc0075 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.191427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.191427Z digest=sha256:92fe9729a87f369b637d4af003c53bedb373654b043873146be516292ea00c0b

Observation 077c2209-9771-4993-9318-d9fca4d48144 · outbound

This paper cites Hudson and Christopher D.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hudson and Christopher D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.196480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.196480Z digest=sha256:c2f1c12812cef1dde38c2201ac68dd0df07155e2a514c817f38aeaf2f09ca0ff

Observation 1473683f-a911-4071-a4b0-79ac04e25f79 · outbound

This paper cites Scaling Laws for Neural Language Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.201224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.201224Z digest=sha256:d83a8a50626fce7ec2670696187a189242e0d0393b977d997fba211d5f98d81c

Observation a59168d2-cd0d-417f-a551-f25d6a825290 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.206167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.206167Z digest=sha256:8809ba60dafe0c5074ee8a2f091fbd93881a3e978839695a4407268c8d5b8fe7

Observation edafc5d2-ee27-42f5-a95d-57ac40f5108d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.838318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.211116Z digest=sha256:f43e1b3c92d7e5dc79ab06629ea89b13cf9661cef3c036433660d431040a80fc

Observation 76c23acd-d042-4fa4-92df-e026d85caa07 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.823150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.216098Z digest=sha256:e088111a919593031cd97a21743f3ddf83e83208e1ca70810489c0b22ea9ae8b

Observation 6b1f51eb-97d4-4ad2-8abc-b868ad943db6 · outbound

This paper cites an unresolved cited work.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T17:21:05.807189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.220768Z digest=sha256:92c9a642d05a814ce72776bef646f1ca75ede095959217870acfabef586aa91d

Observation a2674656-240c-47d0-8726-57db70b158b6 · outbound

This paper cites Datasets: imagenet-1k-vl-enriched.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Datasets: imagenet-1k-vl-enriched

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.792194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.225733Z digest=sha256:c59d427aad4e32f743abecd783d9a6c83b11d7ef2edce6a7cf430e3c97ed9ce6

Observation b1af209a-5da4-4f7b-9133-5f254b5b1a4a · outbound

This paper cites Deep learning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Deep learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.777349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.230513Z digest=sha256:3ff9a4eb2b288dd3e013d3cf65450a90f05cdc5b21322c95a5bc5fca616d0f29

Observation ebbe9e8b-7c63-4326-914a-90bf4f6025bd · outbound

This paper cites Autoregressive image generation using residual quantization, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive image generation using residual quantization, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.761859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.235485Z digest=sha256:26a234a8cbbb23551c9fc873c6503b8915e2c127242041246d21e286b64cfeea

Observation c5d0fb22-0895-45e8-8094-17af05f79de6 · outbound

This paper cites Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.240341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.240341Z digest=sha256:6d43b4d1950c25f3451700eb89b3dc6063266b767bf40f7e1e99f455dd083880

Observation c690a647-abe8-4335-87f0-ebd032a52f8a · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.746758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.245178Z digest=sha256:a1708321b62fc7ca850a90a0a3fa7b690324f392232e1aa87ce56e19b9c82dd1

Observation 5840bd26-27e1-476a-acc6-65ef61c90cda · outbound

This paper cites Llava-next: What else influences visual instruction tuning beyond data?, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: What else influences visual instruction tuning beyond data?, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.731951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.250131Z digest=sha256:b92bae94e6a5e1885ea8e32522d889b5d7d77c4012cf6a2df8356c5d28a1d4d5

Observation a75f6018-7d76-4c07-888b-d6a32297b8e0 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capabilities in the wild, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Stronger llms supercharge multimodal capabilities in the wild, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.716817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.254643Z digest=sha256:6d218fe7d3aa6b0e4b7dce2e2638d67b7f7eb909335895ad553676a09b71ffad

Observation 94b0743b-231f-45f1-946c-d4c93240bdb2 · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-onevision: Easy visual task transfer, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.702250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.259415Z digest=sha256:0c726d762a9f5c466222fdbc1851c78b9c3ff58782fb60c3429f68a839517cd5

Observation 58bda4e4-2289-47ce-ad53-f14da0f948d4 · outbound

This paper cites Llava-next: Tackling multi-image, video, and 3d in large multimodal models,.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Tackling multi-image, video, and 3d in large multimodal models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.687367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.264372Z digest=sha256:203f1786993f81d882e174a97a2d0e0c9ec7b3e461d3d2df7c3bdbae9c43bbc9

Observation d0023b5b-6601-469d-a346-fd2e97af8db9 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.672425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.269031Z digest=sha256:57ee0a3a97af5213543d15e0f3cd46ee0fe69f433cb6f19cb3aee30548a35e6f

Observation 53ab4899-7181-43f8-a309-a5e25a2ed6b7 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.656693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.273858Z digest=sha256:1da20355d0562fcafdac26f01cb73ef4beff8e861c1b93eef90e43a41e2eaaf5

Observation 3b873459-04b5-4fb5-b0f4-0d25d8fa23a4 · outbound

This paper cites Evaluating object hallucination in large vision- language models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Evaluating object hallucination in large vision- language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.640969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.278654Z digest=sha256:271576cc4b175cd1f02004d2d2e11a957cf5ae320ac47b7549719ed5601b47db

Observation 2c4a206b-008a-4440-ad1f-e8d0e1bc9402 · outbound

This paper cites Dual diffusion for unified image generation and understanding, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dual diffusion for unified image generation and understanding, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.625246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.283580Z digest=sha256:600b1b6f7053e09ec7217aaf0e6ec86cd9b68cfc2e1a235c1d07b75344ae10d1

Observation 7f25221c-d942-4d8b-919a-cd855aff610c · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved Baselines with Visual Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.288423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.288423Z digest=sha256:c01acfb7f31c92fc9c21c9a2332b620bbd0a136769c85358fd49b2712ad77f86

Observation bfdfb097-5980-4de6-83ca-b0cf21c628b3 · outbound

This paper cites Visual Instruction Tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.293521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.293521Z digest=sha256:5e3e3bdf71948eee64331c3c95626f892f50ac1bf6e80b501ef5456b174b27d9

Observation 878d5f9d-8086-4274-92d0-89b11770d923 · outbound

This paper cites Visual instruction tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual instruction tuning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.609283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.298785Z digest=sha256:8cf00dcb70667d7ec0bc0e454a9f197462bfa278dae55adb2831233fa36af4aa

Observation 401fee1b-c74a-4080-9fdf-84cc84033a6c · outbound

This paper cites Improved baselines with visual instruction tuning.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved baselines with visual instruction tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.593616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.304374Z digest=sha256:db76108eab6779ee7a6133c2db0c46f718732a402b9874d26cc09344095d7cdd

Observation 1e21cf32-bc91-4ed1-bd31-3ecdcce2692c · outbound

This paper cites Improved baselines with visual instruction tuning, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved baselines with visual instruction tuning, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.578610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.309619Z digest=sha256:47a2f91af612b0f50479e22aa891b1fd3a51d59b62c6f39e20bb1c4f12cc982e

Observation 821f2ff4-a50b-4d46-86e5-e27b39403060 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.562380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.314855Z digest=sha256:636c973e16f41e4fbfdea8fc2791c767b24d4c2f92ffff47c2a7c94ed8793226

Observation 367b11ed-4739-47f0-852c-cdacab181abc · outbound

This paper cites World model on million-length video and language with ringattention.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model World model on million-length video and language with ringattention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.545290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.319618Z digest=sha256:a7c00e4a06b867ba58ee27396fefec027d87db426698fd796690b844804eb73e

Observation ee285e19-6127-4fdf-be4f-100c99da170e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.529584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.324310Z digest=sha256:8cedd520de9c32b8f8a15ed5c8a7f0dbc11b618c7e92fc8c012c82705a3db34c

Observation 93a6883f-d62b-40bd-b416-7dbb7e0339c3 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.511965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.328973Z digest=sha256:bccc704742619641a481471ced922222bf0dce621fd45f7c45a60e0a7496cab2

Observation 0232c3dd-4465-4a54-a616-ba7197d81649 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.495366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.333752Z digest=sha256:fe0c09681e988a1411c5bf5e4882264ade091e61f9936be9118bfa25f7275e66

Observation 84008c93-d13a-40c1-97d6-d1ef4d9b0602 · outbound

This paper cites Unified multi-modal latent diffusion for joint subject and text conditional image generation, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Unified multi-modal latent diffusion for joint subject and text conditional image generation, 2023

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.478887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.338929Z digest=sha256:01d9a0211253dd592bf090b2e3bf0575831850666064a773e80324b804d4f63c

Observation 4f9e81fe-ab55-47af-8ea5-e797451337d0 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Generation and comprehension of unambiguous object descriptions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.343869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.343869Z digest=sha256:8d964016328911179b5134103b61fa806bac91e3183498cadeda4815fc241fed

Observation ed88223d-20e4-4c11-bdf6-8c549e1d9350 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.452449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.348553Z digest=sha256:166d424171592a123793db0c6905656e17fa005615f2d577027bc90e378c0247

Observation 5f1c9815-c403-4bae-99fb-b01ac8f9e695 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.435954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.353053Z digest=sha256:be05ae870e673398a0df916e9fab7db3a840e7d8eb2adbbff1dd56f80a1bd5ca

Observation d3395c2d-6c74-44b1-aeea-bb2680354427 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Ocr-vqa: Visual question answering by reading text in images

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.357580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.357580Z digest=sha256:e7e4748cf7050ce5d9dafc64a00f9227992a6258fa64d77b1721fbbc43ab589b

Observation 2949aca0-29ff-4e41-ab65-d58979c19940 · outbound

This paper cites Improved denoising diffusion probabilistic models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improved denoising diffusion probabilistic models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.362267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.362267Z digest=sha256:3ab45aac4ed13d89a9f51e214267ff28b17a05726e6813862023ce30a8be6930

Observation 6adfc238-4cb5-4864-b2ab-f49da80dc63b · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.367146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.367146Z digest=sha256:dc5ea376c138e502ccac9da7821241f05be95c175cd049de3c5cccb59d11ff0d

Observation d4f27e65-f17b-4304-a135-edd8c59b979e · outbound

This paper cites Du, Zehuan Yuan, and Xin- glong Wu.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Du, Zehuan Yuan, and Xin- glong Wu

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.387857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.371864Z digest=sha256:96df5d2576ba129857b78a13ba609aaa939513aeda23742b9f05e28d30caf355

Observation 3cf0f681-042d-4c4c-9b7a-5ab69a9c262a · outbound

This paper cites Improving language understanding by generative pre-training.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Improving language understanding by generative pre-training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.376519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.376519Z digest=sha256:278ca506b854cd792b89c9ebaf8248d0b136b558d77d28fdd500eaefd0e95ac5

Observation 5e957856-07e5-46f2-9882-ad0d28dbbb64 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Learning transferable visual models from natural language supervision, 2021

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.381413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.381413Z digest=sha256:9878dffc2c4cc9615e771b6b2dc8ee9f6b9ab0be58ff23472e7fb4ba65e09079

Observation 5dedcf49-fcdc-44c6-a8a6-a2a7723e72b1 · outbound

This paper cites Zero-shot text-to-image generation, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Zero-shot text-to-image generation, 2021

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.386393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.386393Z digest=sha256:fa21219296a558129af7e9f489ad29e08a12e40b6bcbc2d755b1d8606b6abba2

Observation 672f3e67-c3be-4d75-a9e2-eabc8f8a6676 · outbound

This paper cites Hierarchical text-conditional image genera- tion with clip latents, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hierarchical text-conditional image genera- tion with clip latents, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.341071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.391482Z digest=sha256:8436fcf6043e752d7798c6853f266f834e9c0bd360c071050aae1caeda722db2

Observation 504fc21f-a55a-4ce4-8205-1b4882b92f90 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model High-resolution image synthesis with latent diffusion models, 2021

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.324461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.396305Z digest=sha256:0c49cb6d4874f88761e1f935910d35dade56a0fc8b3115eb8fd8bd8b98d38738

Observation 64eed9d4-bce1-4264-95ed-7368ce130d2f · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.308472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.401232Z digest=sha256:ab478d93208e1a72f0876e40cba3a45b76e8b48f3508bc0f0f15cbac311d27b2

Observation 3b30c682-79a9-4642-8646-a0cc7866efa5 · outbound

This paper cites https://sharegpt.com/, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model https://sharegpt.com/, 2023

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.292565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.406214Z digest=sha256:922db32637f0308f461d7c27e9ecbf7a44f53aa0da5ac41b833cc3bc6d6fe660

Observation 4f535246-2cb0-49a0-b2bc-4420abf03743 · outbound

This paper cites Textcaps: a dataset for image captioning with reading comprehension.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Textcaps: a dataset for image captioning with reading comprehension

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.276429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.411103Z digest=sha256:29928381ff05e1bd32088f3d0d3c446dac3d0fbf59d0509e0e50ad105d71b62e

Observation 78a7022f-34f2-440a-af68-a4048fe072ce · outbound

This paper cites Towards vqa models that can read, 2019.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Towards vqa models that can read, 2019

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.259803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.416070Z digest=sha256:aa785c14915d2785b6f6ddd9c3f6600c60801b2fadaaae1a29634d9a8ff2e11e

Observation c884e091-49c7-49de-bc2f-8a41d6eb0565 · outbound

This paper cites Denois- ing diffusion implicit models, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Denois- ing diffusion implicit models, 2022

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.421214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.421214Z digest=sha256:5d5142d30ccdda083025c48ee81334cc8aeda9cc1bb6e6af52f32493c945b35d

Observation f336a288-88b0-4b96-9958-bab1ceaad257 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution, 2020.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Generative modeling by estimating gradients of the data distribution, 2020

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.233418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.425900Z digest=sha256:7433873b80b255de1ca369140af6845f00c99097687bbbf526b35a0700021848

Observation f557c42d-75b3-414b-8d52-9bd3628b6b76 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.430791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.430791Z digest=sha256:adf9731cb0adb200046027915732d255bd0452a1f1d8e41a5fa599a15d87d8b4

Observation 9534376e-6416-4d9e-a1db-112c70e201eb · outbound

This paper cites Autoregressive model beats diffusion: Llama for scalable image generation, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Autoregressive model beats diffusion: Llama for scalable image generation, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.218037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.435876Z digest=sha256:7304bd9a6c09c3e5c144ef02469b9b208a610ccb5f86277020b4508285f4cce3

Observation f1b697cc-4b73-4fd7-a086-d4ba312325a1 · outbound

This paper cites Emu: Generative pretraining in multimodality, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Emu: Generative pretraining in multimodality, 2024

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.202213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.440563Z digest=sha256:28a06a065018c0cde58886446bd1890567303e73dd1368256efe8483ad73e403

Observation 4284d52c-fe79-427f-9813-d5d51dfcddac · outbound

This paper cites Hart: Efficient visual generation with hybrid autoregressive transformer.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Hart: Efficient visual generation with hybrid autoregressive transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.185718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.444991Z digest=sha256:1f490ef3af4f9f0f618e4a6c1f30fbb7e9d1afa47edcd73e75ea5feaa0de5118

Observation 33cd3e92-1192-4ad6-ac2f-e43e1264750d · outbound

This paper cites Any-to-any generation via composable diffusion, 2023.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Any-to-any generation via composable diffusion, 2023

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.169008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.449643Z digest=sha256:1100148c240fb242fea6995eee175bcdb13a829c68bbb23f3e221294bda75e22

Observation b7482b26-fc8c-480e-8e16-18dabf8f7f66 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.454118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.454118Z digest=sha256:43872af94e4ffffcfd62a706e17513c4eb19be0bf06b9ebd92b8006e94f6c247

Observation b09cea6d-cd21-4e3a-b502-604c4227c340 · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.153496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.459369Z digest=sha256:d1eca01de07bc3f6ccebb3e6df018dc5a3842b82fffb726e1b3a3d55e85bd20f

Observation 0eb9c253-a750-4dfe-8ce0-e714ac4f1155 · outbound

This paper cites Gemini: A family of highly capable multi- modal models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Gemini: A family of highly capable multi- modal models, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.136580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.464229Z digest=sha256:50d3be5c9067825bc7e72ed01e36bf34a1249f40dfaee5ad5c99f24ebdbcee90

Observation c46ec641-224c-4622-a994-0dbfbee052f9 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.117973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.468923Z digest=sha256:320b2bc410994f91b95c80ae765dee443676f15b9717426e9bcd34107946bda9

Observation afa5e832-8fdd-474a-b72b-2ffe528815ed · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model LLaMA: Open and Efficient Foundation Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.473439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.473439Z digest=sha256:5e3f090c088d574e5e82d9ba6caeaadaa8adde4ead9d2972be725f8c35e97025

Observation 575930a0-480b-4283-8db9-9c67ae203c2f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.477802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.477802Z digest=sha256:a2ee5df6e30707672b3299804b433063412f2aa1f0d6237869d760b7f0358280

Observation 223ddf3d-9771-4eec-aaca-9a5db3eccc20 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Emu3: Next-Token Prediction is All You Need

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.482440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.482440Z digest=sha256:e5b1b11fd9acba48175a2663290d30d3529ec354d55165533ac7402fc8d7f72e

Observation ccf4fa23-a0f7-4caf-98be-6f1a4a9cc6b0 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Janus: Decoupling visual encoding for unified multimodal understanding and genera- tion, 2024

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.099022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.487152Z digest=sha256:b2229718c76133b07e74a6bf358ca164ae986f2bf23dcf96b3f2ff6cb13dc643

Observation 07804346-6d60-4570-81a4-816bb00d4b59 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.491869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.491869Z digest=sha256:ad4a70d137831574d00c20df28da9e1d4345d843fdd26d897128ba489566723d

Observation 2919fe17-42dd-4e1d-a91c-f17cf84ad08b · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.497553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.497553Z digest=sha256:599d7d1d87fd132c2f92ba62b3be4368f1103e6806a5bf6401cf57a68d31079b

Observation 1b285d06-f445-495c-965d-fcd465fe69d6 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.503740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.503740Z digest=sha256:a84c42ba016cb4cdd1f0a6fb046c08f9b83fe7aa5eaceb4ac0683d46c190df90

Observation 4acae73c-0414-4deb-bcfe-5f87fece7fba · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Show-o: One single transformer to unify multimodal understanding and generation, 2024

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.081298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.508878Z digest=sha256:e9ed393b2461128ce78deedd4f527057e8b8f1c2f647f58a5a7cb342c8c7fb78

Observation 3192d077-b9ca-406b-99f0-7d193a3837d9 · outbound

This paper cites X-vila: Cross-modality align- ment for large language model, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model X-vila: Cross-modality align- ment for large language model, 2024

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.065607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.513698Z digest=sha256:bdcf042c6fa70adb35fb3e98327d115886b8e7ce9de6b80c9fb6563962cbad05

Observation 7f74507e-3009-4f7d-be6e-dc1309ffbd58 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.518425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.518425Z digest=sha256:7c2181971486e199ed6e540cd672058c9d862c8e6de38a3710ed8f382ee7f0c4

Observation 2c9753cf-45c7-4c39-bbb5-5c5c724de964 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.048940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.523483Z digest=sha256:95c3365d2e50de3a4be76bc9e12a5e3b97312b0f20db3e5f994d9bc4951f1836

Observation a86774a0-f579-4365-b430-b82d493d0519 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.528220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.528220Z digest=sha256:b1eecbdcc247d321d326bcabd941e9ea7f03b5edc25ada1216b1294cfb802bd0

Observation b0a5f092-939f-4f8f-904d-ecfc93be4057 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation, 2022.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Scaling autoregressive models for content-rich text-to-image generation, 2022

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.533079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.533079Z digest=sha256:1745dc680a388e0e515f2fd7ccb1ec9d83e923605acd93624554da3be0e56959

Observation 8e66f95f-5fb4-4782-b7b8-9f09cc37c952 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi, 2024

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:05.016487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.537979Z digest=sha256:bc290769f7d8dffb9b5aa19fff2a9883f0f943d34117f3056562009f16103969

Observation 86b1424d-1762-47fc-8a10-54433d2ceff1 · outbound

This paper cites Lmms- eval: Reality check on the evaluation of large multimodal models, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Lmms- eval: Reality check on the evaluation of large multimodal models, 2024

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T17:21:04.542716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:21:04.542716Z digest=sha256:ae79d54c6344b4652a8ab589d9e82a1c7ddaa66e202fef5a89bbdc3550dcb66e

Observation 016c92f5-1ba3-4d55-b361-2ff766f765d6 · outbound

This paper cites Var-clip: Text-to-image generator with visual auto-regressive modeling, 2024.

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model Var-clip: Text-to-image generator with visual auto-regressive modeling, 2024

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:21:04.983358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T17:21:04.547305Z digest=sha256:721d213c1b341e2994db6abfe0100ad2c8d16bb741d84e0f0dd1f424d83257c1

Pith citing papers

Observation ce6f86c3-8ae3-464c-859c-31cbe8c6bfc5 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.412504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:3f8087fad2b65116fb4ad594c7fb985bf0f68da414a2e61c58ad2f7d4ceef1a8

Observation eeeb5c56-95d2-4e74-ba1a-1502d5f79c9d · inbound

Do we really have to filter out random noise in pre-training data for language models? cites this paper.

Do we really have to filter out random noise in pre-training data for language models? VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.442501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.442501Z digest=sha256:b42db0b359af8c56f8b6b862fbb43ba102edc8b764f0773695953dc36dfed63d

Observation 18734500-4d57-4696-84bd-10f7449a3a0c · inbound

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens cites this paper.

Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-16T11:48:03.500362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:48:03.500362Z digest=sha256:87c7c6dcb81f955690ab78c553fe7c7b8447a8b16aa493e663be9ef5e73abad4

Observation 3678c49c-a694-4464-8743-21f60760af38 · inbound

Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation cites this paper.

Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:59.260983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:59.260983Z digest=sha256:7d2d5a257080753950729ab9a496f4eec90822fd84c2e3618974ec9d34ad64d4

Observation e6b68b22-0eb4-4371-9e7a-5af03bd4b398 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.753839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:1fec60404ab8f1ce66552fa8e3520b67e14f92ca71df752e5d7e75a781109f66

Observation 20f41297-a7fb-4c90-beca-1ca735b99dca · inbound

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation cites this paper.

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:35.569184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:35.569184Z digest=sha256:324cbf234b4e07dbb6ed53789e0177d99bdd8b2dd8740ef0fd0a8fdee49c4802

Observation 712f8acc-f2d9-4f5e-a1d2-08fc799d6f78 · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:26.930717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:26.930717Z digest=sha256:ce8c290eb9855ca9ab054f97674df115819d57d8aac6473f192db1eeec392b86

Observation 84a7f1c1-7cc2-4716-a4de-cae359319d48 · inbound

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies cites this paper.

ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:22.561926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:22.561926Z digest=sha256:6b142c86d82af775ec1d9d2ea1f0c2f44001582b47647537a8d629727991e6f3

Observation cd3cdd43-3777-40cd-aaf1-1e35af05d4e1 · inbound

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation cites this paper.

Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:04.966173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:04.966173Z digest=sha256:99c8eccbf236d154dea47d4f79f49f8e0eaf85fb39edf742eaca4f6b6c4a0ea4

Observation 226e02d8-a9c5-477b-b69c-0990c34ccad1 · inbound

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models cites this paper.

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:20:09.981818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T16:16:15.819622Z digest=sha256:8ad4cfd6695ca795d85b89bf333b70055d8702386104aa73c5c7b2e951c0b5c4

Observation 07a519a3-d3da-46d2-b5f8-00102b18b19a · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:6a1d9118cb09f27583148e26a1a2d2a2de2d47886740a30fed751094d20889ae

Observation e988bc26-b644-4d96-aa89-0e5d9a344453 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:01.603385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:01.603385Z digest=sha256:e8e6c2e6cc72feaf1ae94a47c2a2fa65ae204a0849d0a8a643a6144809c50571

Observation bf8003dd-c19c-4a18-adff-8a871acc3f23 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:08:15.816753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:5e05f0d227ec64fa0def4f44df6c5f59d122317735683d4c8cb8dde90b870cc3

Observation 996a1a94-c94d-40c6-8ce7-64ed0331975c · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.285652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:5e40f3a5a5d1ff2afd74d185e3d23c99af02ce7e5b8eb5aaf91cce708cb26c43

Observation 061a9d94-8747-45e3-a2a1-d89cc9b3bf21 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.311433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:9988b2fa446b4c0b3167eb264249c99831db94399f21251c72d90ba2706b32cb

Observation 39ee1ca7-63ea-49ae-acd4-15ccf867488d · inbound

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling cites this paper.

HACK++: Towards More Effective Head-Aware Key-Value Compression for Efficient Visual Autoregressive Modeling VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:27:24.385865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:46:43.514413Z digest=sha256:fb4d62a945fda37e00c485df8c3ca8cf17ec285839f0026a07ae2eb7c3d15ee6

Observation 8263a698-25c9-4f07-839f-f4a8d6190a4d · inbound

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards cites this paper.

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.492070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T04:58:15.891214Z digest=sha256:bf0523413780e95f4eb03c2c25049869da5dd28bc8ff68fee1e100cff5a84d8f

Observation e4e52c5d-0910-439e-b9d3-a27e11303402 · inbound

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts cites this paper.

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:17:07.191767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T15:14:36.946247Z digest=sha256:b78a292a87ea488ff6eab1ec0b03ea78005f0816df778cb81bd2191c9cd0d8e6

Observation 158619c6-8924-48c6-82e8-3e8a88ce896a · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 268

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.543433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:2ed464901ff6ebff30400ba57536f4e4d100736faa470ded453c6a00dd18943b

Observation c4f225b9-d26d-4fc4-bedc-b1cc4b7b8c79 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 268

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:880701396089cc497df169741fe6158bb9cebcad9788e9ccadf95e72c883d6d1

Observation 2483e1c5-8177-4b4e-be7f-ee2e17676c4b · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.526450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.526450Z digest=sha256:afea645eaef88842d0686c42798f77d9260564044c03fa0b65a4379a7870cdc7

Observation e391844b-f528-4ed1-bb93-78e823d1782e · inbound

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model cites this paper.

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:44:46.078357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:44:46.078357Z digest=sha256:9450e1bb2cd6f90746a42a88085ff5b24e393534d5483bc6d086576c53168623