Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:00:59.644394Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 7 inbound Pith citation observations for arXiv:2505.23043.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:00:59.644394Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:14:06.879405Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T00:04:22.359705Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1633aeeb-f418-4f81-ad87-d005d7b4126e · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82528c4f-51c5-4214-9f6a-e7557d20293d · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb0105ff-ebfd-4517-88d0-7522bd9265c6 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97a0651-39e3-4f7a-bef6-fdb932b66777 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Emu3: Next-Token Prediction is All You Need
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99aaf820-8c5e-4420-8422-7c36a7c5b90e · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88636981-0a87-40b5-98da-fadb43d3c281 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Liquid: Language Models are Scalable and Unified Multi-modal Generators
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a6238d-09c5-477c-81d2-1735344b6635 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ed10ee-6de4-4cfc-bbad-9cbcf6387e45 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Emu: Generative Pretraining in Multimodality
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8bc3102-07b3-4228-a66a-a07cca591a52 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Instruction Tuning with GPT-4
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485851a2-adb4-4330-affd-9aa9646c0757 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8b0986-ae8e-4247-9022-ae1ddb487e73 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Lmms- eval: accelerating the development of large multimoal models (2024).URL: https://github
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94957f9f-b3b4-43be-a67e-93cdd9612117 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04638e1-b8f4-44e1-a21e-6c2e57cbda22 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1136db49-8e14-48c3-8528-270948dfa0b4 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4721438-4639-474b-b2f0-1b4b5fca92fc · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation _u” refers to understanding-only; “_g
Reference 393
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bcf40dac-dd34-4641-9038-bc2f8d4cb9c7 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c5e914-991b-4de0-87ea-f8c90b403eca · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ee0b67-963b-4bf2-8664-3c3fbb24b85e · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 369ca1a0-1eff-4d6e-8cc4-9fbe475d1db6 · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69654701-8958-473e-93a8-dc96e7cd432a · outbound
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844b3e6a-ea89-4fc1-b5cf-a92b9f6d2639 · inbound
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfbb7e93-76ba-4638-918e-68d65e3803b1 · inbound
Skywork-R1V3 Technical Report Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b1327c-347f-44a8-ab5b-a159e9c18129 · inbound
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b13ec3-a465-48dd-ba92-993348a6ce22 · inbound
Transferability Between Understanding and Generation in Unified Multimodal Models Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3739057f-d01c-45e9-a4a0-4a849778039b · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 261
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2ee9f89-6e09-4582-958a-ab9830eeb666 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 261
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e1738b-ddd8-4ed3-877e-099c2c1db321 · inbound
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
Reference 156
Source-reported events for the cited work
Unavailable: canonical work link unavailable.