Pith. sign in

Paper Citation Record · LEDGER

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 7 inbound Pith citation observations for arXiv:2505.23380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23380 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:27.275816Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:36:08.089323Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.505020Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aec5fbc9-7aeb-4bbe-a4ec-b481b7096a02 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.247048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.247048Z digest=sha256:a4d6e34b73f8add154c96e0d80c9ca5ab5d5c888c4140fff26e66057db5f71b0

Observation 58bb2215-a757-4c2a-965b-9064c404f0fa · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.392275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.392275Z digest=sha256:8bda5f2b2201283eecf9d797bbeccb08e7c5ee22f796e79d866382f3ae1b7bfc

Observation b548974f-40f1-47cd-8543-216fa5d1fe4e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.498548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.498548Z digest=sha256:c20b3654f8abfa362e28a707778b6ed8c1cc091fb55e82cb648cff7154c4ffa1

Observation de27a906-4d36-4c4a-a3db-7a66dbc3c57d · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.647902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.647902Z digest=sha256:93c5c0daaa4d1c96301dd94832e55860e72c3ab88c7b147cd46e38b8988cfcae

Observation 5188efac-6b8d-4830-88d7-73269d8607d5 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.803824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.803824Z digest=sha256:9caa671056c3630cf4eddb66d2a9e5680ecff8011db644b7ac1921dec6a6d875

Observation 451f3035-a90e-470e-9aec-921385dd5ba2 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.903111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.903111Z digest=sha256:b55d9b149da2b631a9518329044230f96e508c19203e63736c950d314eb58af3

Observation 20b0f5b4-8f58-44e8-a842-3f91de0cca06 · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.932772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:21.990830Z digest=sha256:2f5c9d7c84fc204e3aa3016c6fc451664e385bc3402cc5b40e0e70f06f32b0b9

Observation dd42ec44-cf96-4338-868e-c364240aa56b · outbound

This paper cites Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.653559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:22.106661Z digest=sha256:4382431ee6b917e2e8ae71d83de6caff5f08b7a55c99a1904249fece11f090bc

Observation 9dc2eece-21b3-4a7d-a269-9cd6afc167cb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.203748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.203748Z digest=sha256:93b5c1ce9b9b3ac83afb76e94f0c885e33a005215af7612c4573060f27935f8f

Observation b9db14ca-f33c-4c9c-a4ac-1bcc3f660906 · outbound

This paper cites DreamLLM: Synergistic multimodal comprehension and creation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DreamLLM: Synergistic multimodal comprehension and creation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.295884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.295884Z digest=sha256:20025c5e810389239a01bca0e2188b4e1e42d43497692a5e761d50a089729c01

Observation c8db548b-d4fe-41f3-a841-c5bd44d5cf05 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.410352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:22.391653Z digest=sha256:8f2bbecba858640603caa5a2e808c8223659bd7779a24d13c50ce8297957390f

Observation ba29a64c-395e-4fe1-afd7-b2d004a890d3 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Taming transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.134836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:22.481332Z digest=sha256:a1d8337ab0b6b10856faedd3ff55b4a401e73775f0c353172eb0fa01595d2489

Observation 66ebf948-5046-4075-86b9-878955536ed3 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.622402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.622402Z digest=sha256:bca18da26ac7bbfb13772af429054eea9a90831fa24fd97e5bfa8f83354a466b

Observation 018ec1df-5998-4250-a5a9-2177b76fca32 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.801107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.801107Z digest=sha256:0e115ba410af4721b81c0eacca45d4e5303f59458912e100956ec2a0746857e9

Observation 3e8053ea-4c2d-445e-aebf-e7b2b1ee8a09 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.895103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.895103Z digest=sha256:8f7607887008ca7cf38f15c0f385f74b11cf18bffbae4cd9296a7adce6895a25

Observation 9b0265b6-958f-4093-95cd-96faa8adc758 · outbound

This paper cites OpenAI o1 System Card.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.010341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.010341Z digest=sha256:a57ada8fbb6274c7b787b612cd65106cf04af56be8b4f3fd8db3370ddab3b915

Observation 668be6c0-d07f-4c16-b854-b3f4c42bb6ef · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.148041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.148041Z digest=sha256:2e9e7a4463bf07a9b225eb8911e544d512ef49666da393ab409fb52d3327c929

Observation 2c634cf9-a721-40d5-8aac-613f2429c5f9 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.280824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.280824Z digest=sha256:9ed4c2bfbac5a75e46c5d163ecc3014ebfc47c874ac49a6b8dfc149c13fdfe2d

Observation d4fa1d2f-af8a-4a60-aea4-08ba69542773 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.378411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.378411Z digest=sha256:f4df0f9f778823e61af1f32098fbedcab50ca039db03833e3bce94c4231d31ca

Observation 43be3b75-d5fa-411d-a840-991ebc006b4b · outbound

This paper cites OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.481480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.481480Z digest=sha256:0ac591130ccd89c2633e84343d7f69b9c4e0ed013fae1d3d0e932e0ea4d7263e

Observation 2e9ad2e1-f710-4333-9696-fce27c100765 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Evaluating object hallucination in large vision-language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:29.928449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:23.546372Z digest=sha256:deac32c865fd184859d7d49111fe9673f8a088737e230172717382de0b99d8af

Observation fc666632-e083-4888-866b-1ae50628efdb · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Textbooks Are All You Need II: phi-1.5 technical report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.636560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.636560Z digest=sha256:941e4fc8b52428d4c14b5f7fa1e3763a761ba28b16cabe9ea744db1557ed5a8a

Observation c72fe73b-cab8-499d-8243-8f62cb7fce89 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.757213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.757213Z digest=sha256:69f57fd7261c95f22a0b9e1af2e9eb7abbd76e513607beadecd733a5897ce5fb

Observation 51de8089-3bda-4734-b1ce-cfd61af89862 · outbound

This paper cites World model on million-length video and language with ringattention.arXiv preprint, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning World model on million-length video and language with ringattention.arXiv preprint, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.884539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.884539Z digest=sha256:1d8f98754aeeb213a223f99ca2a612dcbb4815c7b5574a0c61e79b9c0ed9d4aa

Observation ef2e40da-f413-46e3-91e8-cafa6f0483eb · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.079598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.079598Z digest=sha256:5bdf30c2a621d39391d6853c8573deb09ff65087b4175f5fb1213801a94d5485

Observation cd8acce9-ac40-4322-a4be-1e017b220653 · outbound

This paper cites Visual instruction tuning.NeurIPS, 36, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Visual instruction tuning.NeurIPS, 36, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.253921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.253921Z digest=sha256:aee3c3e55c5fbbae706e3dfc88dfa740ca78799dbe2e44bd36fa4223563bd558

Observation f7a0385d-c4bf-4561-aeb5-f7d4665e6dd0 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.426795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.426795Z digest=sha256:d09e13289007d3d4ac3de218af9107766a2a04b9a32c76555b21300773b73cc2

Observation abbcd6f8-fb9e-4226-bd62-53b52b9d3f5c · outbound

This paper cites UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.588515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.588515Z digest=sha256:8145256ae85b65b6f4f10b2337e7b8244d81f674fc9240537c86999b6eab189e

Observation 8ed1468c-7cf9-4a0c-b1c7-8dcfafffb433 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.689189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.689189Z digest=sha256:4d34a0159b2cffe9d9d9a9213922e719c3136e986dc681c5f4dceb0fdc44c147

Observation f7ab2202-5522-41f1-b062-d5223aef3e5a · outbound

This paper cites Learning transferable visual models from natural language supervision.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Learning transferable visual models from natural language supervision

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:29.228511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:24.823422Z digest=sha256:c67221f816532eefac141ef1d4e563f95970865a2ef2a1a87ca75c803f06e3c3

Observation dc850bae-8438-49fc-b1eb-d0faf2e5271d · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.018029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.018029Z digest=sha256:901acae62d991617806a42f2827c58f3392a37ecdfc22e586e66bb84eaf86c72

Observation 971cb1b3-6d1f-49e9-bc2e-e50d243c1c7d · outbound

This paper cites Proximal Policy Optimization Algorithms.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.202348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.202348Z digest=sha256:1e0dfa87968cd076a49e32667a55ab15e2bcbbe59b4ed132f50daf5646f73191

Observation 62ac2f3a-ce0f-4374-8b79-36b0b4b1a2de · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.336876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.336876Z digest=sha256:8a2e781e5da92a648675881d6f8e78de3ededfafb9a354b241ba8acbd5a586d3

Observation 05dac12b-2ec1-4114-b49e-baa9bdda25cf · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.463377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.463377Z digest=sha256:233dced5990abb6aac34e05fe211ded96f92517db45a65b89802d9b2427ea342

Observation 2b70c008-595a-485d-b86a-263d5d806029 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Journeydb: A benchmark for generative image understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.595623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.595623Z digest=sha256:0aad324fc71ac8b9e110b6211e07a0245685974313db717afbc4f2c3caa7dbfe

Observation 16b7d5a9-ffd4-40e9-a283-5569a2b83d4b · outbound

This paper cites Emu: Generative pretraining in multimodality.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Emu: Generative pretraining in multimodality

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.723242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.723242Z digest=sha256:6d6aa7169bcc0f1060f1dd5c3164fb569fc4ce70801f1c675ffeb73e16bea46e

Observation 2a227863-0077-490b-81b3-7e8ae235e776 · outbound

This paper cites Any-to-any generation via composable diffusion.NeurIPS, 36, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Any-to-any generation via composable diffusion.NeurIPS, 36, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.798707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.798707Z digest=sha256:d46e18e096eb9055201b05a5f51d2557571cde0e800f94146b05ba45348bb2cb

Observation de4c8042-f473-4ec0-be7e-fa036b0c4a4a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.864139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.864139Z digest=sha256:74d88c06efed556418f5e1df561f42c81e8f50f2775c7bf499e7f9da3a777023

Observation bc8082fa-ff6c-40a4-86de-58673450e2ed · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.933048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.933048Z digest=sha256:1b80a551eedb00a93b0fcee364462d2e0d5c48ec071fde8d3333c27554df107e

Observation c9438d87-f325-472f-852a-a530122fe4ab · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.025472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.025472Z digest=sha256:910591e648271f0d65cdd6757d630e48db7508c218aa8d8194a178b09aeecb88

Observation cbde4961-f1a8-460f-a1a2-23d466d71dd6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.121878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.121878Z digest=sha256:0b6ed4c18635f5e6c6ef1ea1dbc3afe750fc6758714dfb2bd046c1c40a56c263

Observation 3b1f686e-b87c-4b9a-a8fa-b6dced2d1622 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.245919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.245919Z digest=sha256:c8a5e2d6e9ee2e306ddbbae929f7bd35ee677e9f19829f1d3323ed41b2fcd5cb

Observation e308dcb7-501d-44bf-9a0f-d4708432b45f · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.366317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.366317Z digest=sha256:ce192a5a7b014851e9a9d99b41a79609848d40e9a938f0a6016e806fec13755e

Observation 727c3c1d-52ce-4d27-bf43-8b46fbc2fb7f · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning NExT-GPT: Any-to-Any Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.460651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.460651Z digest=sha256:f607374e06e2843c221e6da5e2941686d248caf442ddde6f266ee37e6b39e6ae

Observation fd41ded9-b835-43de-a77e-0011efa915fc · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.558833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.558833Z digest=sha256:b0875638b57a57c049d1bbdf3dfb45ecf066d32a045e108b9a3426504d663168

Observation 4c66c372-c2ee-43d4-9d91-34869edf7d61 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.689185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.689185Z digest=sha256:2b329b62cf8cee1beb322f25a7e3a9497d360609bb1d5a9b59d0b92dcbe2053c

Observation 17d2601d-6d50-4659-876f-f6998775a666 · outbound

This paper cites Qwen2.5-1M Technical Report.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Qwen2.5-1M Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.847399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.847399Z digest=sha256:4816f9cca4daae804d8a84232dcffb7dd5ad9d6f75355df9b4a58b7e8ec2cb76

Observation 34296fd0-0dcc-4c8b-b655-9991d86c2266 · outbound

This paper cites Hermesflow: Seamlessly closing the gap in multimodal understanding and generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Hermesflow: Seamlessly closing the gap in multimodal understanding and generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.974393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.974393Z digest=sha256:47663ded504ee0580e365b030fe10ea89d97918aa7c2f077d75e4e7fe8fce39e

Observation 598112ee-2156-417e-9f68-063ebb26669d · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:27.100136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:27.100136Z digest=sha256:a201aa024902554215577bd85fb0291ae96a97544d317541b554a0d6a992e2f2

Observation 1266f9e7-3c5b-4c78-baf3-94b4c1da3891 · outbound

This paper cites DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:53:27.611095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:53:27.204685Z digest=sha256:a3956b1746b420ee2166cc3f8c2495d5f648551ce2a7c994727fc0bf99528619

Observation b7ac9371-df22-4361-af2c-b5731326aa0e · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:27.275816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:27.275816Z digest=sha256:188cf376d6f954683ed9027c6d767d10fead69adc3b626cd4c129975b1152f30

Pith citing papers

Observation c63afbc4-b1d7-43af-ba5c-dc92b302b53b · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.089323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.089323Z digest=sha256:2118b27645226b53ce61e19fab652b5c4e9811372c69349888185277d6ac27d7

Observation 1b38343b-63f0-4da6-9d9e-c2958e2738fb · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.172394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.172394Z digest=sha256:f648bd733ed12c690c0875282088a0065e856ffd07c7e85961974f7ca97f3724

Observation f65b3710-3e23-49d2-92cc-c271d97ff38f · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.438705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:5bad386826413a0925106b2bdd68535e3c5edf2df49465897b41e6f366e958f9

Observation d3fe3d45-ba2a-4eb6-98ba-c8dccb1f69d1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.711658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:2f2231cfe44649591470c52a71d5b9dec175ec26e6882517bdb8e7694548ebac

Observation a4bd0c27-f7ff-4223-9366-d8b551a7fab2 · inbound

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards cites this paper.

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.506849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T04:58:15.891214Z digest=sha256:6f6d9ea0ce884a480ac3a402c138ace5305e6fa43f7ddee77ef2373b00372a80

Observation 75fc4684-086d-4358-8cb7-ba8102369c62 · inbound

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning cites this paper.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:e700bf7f307298769896f99d621bd9226dc8fbfc56ac5e7bbcb616f309f8b066

Observation a2b02145-2d74-4af7-b2cf-57e6af75cced · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:58.185926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:58.185926Z digest=sha256:5c7f6dd248962b7a317865887b2d3be9ed3ecfd6475358cfd074b8f0414782d4