Pith. sign in

Paper Citation Record · LEDGER

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 7 inbound Pith citation observations for arXiv:2505.23043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23043 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:00:59.644394Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:06.879405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.359705Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1633aeeb-f418-4f81-ad87-d005d7b4126e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.690899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.690899Z digest=sha256:18de13b6e01118d8103cb90cb5229911be5ce60c7f73456cb23fd954dba95c3f

Observation 82528c4f-51c5-4214-9f6a-e7557d20293d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.923032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.923032Z digest=sha256:dbeba957d527c06d10a952f6e752ce54582c2bfbdd19ea388e48bdd000120fb5

Observation cb0105ff-ebfd-4517-88d0-7522bd9265c6 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.044538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.044538Z digest=sha256:abe227c5beed84881f3ac2d3b09089c74380c1403d2027528e82ef69d1953aeb

Observation e97a0651-39e3-4f7a-bef6-fdb932b66777 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.375909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.375909Z digest=sha256:8d25a413330ba17415b40ccd0e3495d0fd46b887b999c7721f6b5ddfe6d98a4c

Observation 99aaf820-8c5e-4420-8422-7c36a7c5b90e · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.491677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.491677Z digest=sha256:292e422410391e30b861d9bb4ed949cfcaa78f870c48709bcfd391bdffe69432

Observation 88636981-0a87-40b5-98da-fadb43d3c281 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.581514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.581514Z digest=sha256:d8200806f4592e2a0e8e5c6a4a94e88bbc2d650fed9b5912a5f1d6c402705f69

Observation f5a6238d-09c5-477c-81d2-1735344b6635 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.779890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.779890Z digest=sha256:51676f2325ca73c684ba5b75221ff7ab1cb7de38dbeae0563a84ae213bdf9f09

Observation e9ed10ee-6de4-4cfc-bbad-9cbcf6387e45 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.920678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.920678Z digest=sha256:5dd76314449250563a10af8cc6e8fae030946d6c4d5adc13c957b025a853c661

Observation c8bc3102-07b3-4228-a66a-a07cca591a52 · outbound

This paper cites Instruction Tuning with GPT-4.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Instruction Tuning with GPT-4

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.061087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.061087Z digest=sha256:833feafe0bde5d9fd1a0c556d94fa670d3218bbc545a9b392a2ea04f84918984

Observation 485851a2-adb4-4330-affd-9aa9646c0757 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.135161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.135161Z digest=sha256:f095af3129f7f128efb447e23c2a81cf5a4f4602c9265e767f11df8851a65b3b

Observation de8b0986-ae8e-4247-9022-ae1ddb487e73 · outbound

This paper cites Lmms- eval: accelerating the development of large multimoal models (2024).URL: https://github.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Lmms- eval: accelerating the development of large multimoal models (2024).URL: https://github

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.992004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.224555Z digest=sha256:9436df0283331beca056d7d94ddd6e09d2ffaab9c404cff57dd1b497933ec308

Observation 94957f9f-b3b4-43be-a67e-93cdd9612117 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.314486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.314486Z digest=sha256:7b3d1cbb2dee4782530907fb7d6f3f14801cd40ce8295fc2623f0ac055fcc916

Observation d04638e1-b8f4-44e1-a21e-6c2e57cbda22 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.421149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.421149Z digest=sha256:ef22d8b0eeb4d05bd02a51c2978db234ae92d938d64569dcd313cd96025b8fcd

Observation 1136db49-8e14-48c3-8528-270948dfa0b4 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.532990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.532990Z digest=sha256:8256751609bdbd1fbb255f27b2776ec1ae5a9d4a88fc30f8fdad1d0f9cda10e3

Observation e4721438-4639-474b-b2f0-1b4b5fca92fc · outbound

This paper cites _u” refers to understanding-only; “_g.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation _u” refers to understanding-only; “_g

Reference 393

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.855360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:00:59.644394Z digest=sha256:218878155b88418cfc471b1b40f0eef284b716d7dd3f2a0325d429db69da9994

Observation bcf40dac-dd34-4641-9038-bc2f8d4cb9c7 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.975213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.975213Z digest=sha256:097470f4d1f646d98aad167c07dd282e6a0291a6349bcc3038718111ad441ff0

Observation f0c5e914-991b-4de0-87ea-f8c90b403eca · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.701131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.701131Z digest=sha256:65dcef1ec6c644fb90f8fae571731725e791bb33f8832be0fce9d61b9b8a4293

Observation a3ee0b67-963b-4bf2-8664-3c3fbb24b85e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.782523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.782523Z digest=sha256:b6645f457341ee363fddc51c161a947ef081414427cb1506182d4b0a9918ddf3

Observation 369ca1a0-1eff-4d6e-8cc4-9fbe475d1db6 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.276354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.276354Z digest=sha256:3510d9a2ad3801cb1143f253738512cd5f63e1dae4ebd82fae9d652213879c73

Observation 69654701-8958-473e-93a8-dc96e7cd432a · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.148282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.148282Z digest=sha256:c5e06f677d4aea55048f28a6ee0927dccb93d71ec22cc495128f4cc1434bfab3

Pith citing papers

Observation 844b3e6a-ea89-4fc1-b5cf-a92b9f6d2639 · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.086711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:1cd6e8123667c55f19a90c4001b14e97a081f224f5f74dfb2227c1a7bb949573

Observation bfbb7e93-76ba-4638-918e-68d65e3803b1 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:06.879405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:06.879405Z digest=sha256:8f45522e32d49a6e5a070edc53f9476fcb61c9beae502e9cc35554c654bf15fc

Observation 46b1327c-347f-44a8-ab5b-a159e9c18129 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:58.520524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:58.520524Z digest=sha256:31e3009b7acf46f5ac44417e3b7e223162a44cd4918bae1398013766c0c2e4cf

Observation 90b13ec3-a465-48dd-ba92-993348a6ce22 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:837aefdf7892bf6a7d99ebb868cba0fbb28ccd83dc8e3246a9e3db812908b67a

Observation 3739057f-d01c-45e9-a4a0-4a849778039b · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 261

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.360990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:fef34719c7408bfa8df59c53fccb79b33ea123b82c37f7f125fe04c3ecc7c3d7

Observation e2ee9f89-6e09-4582-958a-ab9830eeb666 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 261

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:7704b0f6a06c8f9ba063a5985d51cb6908ebe82af2ceb339cd6dea2563cdeba9

Observation 13e1738b-ddd8-4ed3-877e-099c2c1db321 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.215574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.215574Z digest=sha256:a47d983828a1c781fc61091d772d930be63c9590801ab7f1dfeade7274482ebb