Pith. sign in

Paper Citation Record · LEDGER

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

As of 19 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 8 inbound Pith citation observations for arXiv:2505.23043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23043 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:00:59.644394Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:56.973761Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.359705Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1633aeeb-f418-4f81-ad87-d005d7b4126e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.690899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.690899Z digest=sha256:fb698921c3870bee319d78a04fe95c9ef7b9d15011bf91a2d38fedc35d687271

Observation 82528c4f-51c5-4214-9f6a-e7557d20293d · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.923032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.923032Z digest=sha256:f8a2c78ec0b46b78e15be4ab53f9455c5b10820eed7f906ce0f4c221adc3e080

Observation cb0105ff-ebfd-4517-88d0-7522bd9265c6 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.044538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.044538Z digest=sha256:2ab3f342082729107be089eb3ce6f453de8007a0478b9c25143e1daf5329b1fb

Observation e97a0651-39e3-4f7a-bef6-fdb932b66777 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.375909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.375909Z digest=sha256:378d59ec8d8110f749fadfe2570fb0cbbec02423bbeb02b50385d6de1b542b3c

Observation 99aaf820-8c5e-4420-8422-7c36a7c5b90e · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.491677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.491677Z digest=sha256:cc53657d296ab3786cf396d954ac7ef51a2f7e684ac926d7f48be5ac00fac3c8

Observation 88636981-0a87-40b5-98da-fadb43d3c281 · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.581514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.581514Z digest=sha256:6c5f25e01513f4274162c54cc0b7d64894ad66ec0e2d3a725d7b9306a258d256

Observation f5a6238d-09c5-477c-81d2-1735344b6635 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.779890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.779890Z digest=sha256:e5b00d549af39c3ed50f68258e1975676954b93292f8b2eac9fea72b903d3753

Observation e9ed10ee-6de4-4cfc-bbad-9cbcf6387e45 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.920678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.920678Z digest=sha256:fe057603189a6d3bc163e04b4e062c2f8f2a6028c1ae74ae90876f392003aa4a

Observation c8bc3102-07b3-4228-a66a-a07cca591a52 · outbound

This paper cites Instruction Tuning with GPT-4.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Instruction Tuning with GPT-4

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.061087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.061087Z digest=sha256:1fa61c3926d65ce716eb7caeaae0ebbc6188271a759bddf9313a7b4d6385772a

Observation 485851a2-adb4-4330-affd-9aa9646c0757 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.135161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.135161Z digest=sha256:50b5569e50e197b69d35676c74cc26e69442476d81a7f2c645e2f76c93f89372

Observation de8b0986-ae8e-4247-9022-ae1ddb487e73 · outbound

This paper cites Lmms- eval: accelerating the development of large multimoal models (2024).URL: https://github.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Lmms- eval: accelerating the development of large multimoal models (2024).URL: https://github

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.992004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:00:59.224555Z digest=sha256:bae716d983f41a02ef014295fec73f33f2d47fa57ee6912a7b14ca977f58257f

Observation 94957f9f-b3b4-43be-a67e-93cdd9612117 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.314486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.314486Z digest=sha256:d669e7df5db968c55a74c37fdd9e3acc51d602bd6fe4e74421e498ccc816bc17

Observation d04638e1-b8f4-44e1-a21e-6c2e57cbda22 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.421149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.421149Z digest=sha256:6048d3844bf733124b74a29a522de84aed02de79d297887040725f3f084ce5ef

Observation 1136db49-8e14-48c3-8528-270948dfa0b4 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:59.532990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:59.532990Z digest=sha256:f2430a2d18cda4ea6b7bcb3400557aa2171255957ac396a226a7a26e90206fe8

Observation e4721438-4639-474b-b2f0-1b4b5fca92fc · outbound

This paper cites _u” refers to understanding-only; “_g.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation _u” refers to understanding-only; “_g

Reference 393

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:00:59.855360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T13:00:59.644394Z digest=sha256:f3ab811389f12ad1ecdb178da4c2839acb48fd76d555c74f3ca36797452e4839

Observation bcf40dac-dd34-4641-9038-bc2f8d4cb9c7 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.975213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.975213Z digest=sha256:c0f0ae0405ef45dee9ddbe419867db80d45abc5cdb4f0a0d8b70d0de2093c1ff

Observation f0c5e914-991b-4de0-87ea-f8c90b403eca · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.701131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.701131Z digest=sha256:490d231cded41463ce01ec64650195b05234c91f8ca59d3688b7720035951a87

Observation a3ee0b67-963b-4bf2-8664-3c3fbb24b85e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:57.782523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:57.782523Z digest=sha256:8b03c2e22da5f5314a896fb0222a1d220c09d528370354961a36c4637a50b021

Observation 369ca1a0-1eff-4d6e-8cc4-9fbe475d1db6 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.276354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.276354Z digest=sha256:3ecf9986fb279d5392e75c17a66b8fa9133b70a2e74b9fa58c3ce1e00c7fe1a3

Observation 69654701-8958-473e-93a8-dc96e7cd432a · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:58.148282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:58.148282Z digest=sha256:6f1a106faaceadc8f19d3658a0c4a24400fe070759bd56c34213f6a0e54bb87d

Pith citing papers

Observation 844b3e6a-ea89-4fc1-b5cf-a92b9f6d2639 · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.086711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:4f5b0afdde3e0d19dc95a14277e8b9036d92c8c59a2044f18f33dc715828ebf5

Observation bfbb7e93-76ba-4638-918e-68d65e3803b1 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:06.879405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:06.879405Z digest=sha256:f063a255a731bf525417cb3154924211dc049834cc7b4ec5abbbc0a84438f497

Observation 46b1327c-347f-44a8-ab5b-a159e9c18129 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:58.520524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:58.520524Z digest=sha256:47c47fbff974887afd07ac0303d7cd6a76998a945014e9c09307a9f8aae88059

Observation 90b13ec3-a465-48dd-ba92-993348a6ce22 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:5911acc8ad1d52d57e2c118bb9748e84857d7ff07fc319272637712e342f8b7b

Observation 3739057f-d01c-45e9-a4a0-4a849778039b · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 261

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.360990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:0bfb03502eeca18132eac96ee414b2ef11f7a3409bf2c7c57aa8f53ee8c6874f

Observation e2ee9f89-6e09-4582-958a-ab9830eeb666 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 261

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:76468aee55e965facaef7bf34ec09f879dc78bc565ae11e5e08a43c37d356706

Observation 13e1738b-ddd8-4ed3-877e-099c2c1db321 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.215574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.215574Z digest=sha256:58cde8596f41d60b9faf97848ef4d816c4e57306cda12a624019b6f7aee453f3

Observation 54ddf904-8025-4371-bf60-a34609ceaede · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.973761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.973761Z digest=sha256:52e772515628df634cb28dcfb8cb630c24b6030fbbda0f5a12caf57c93c67c92