Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:58.409542Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 10 inbound Pith citation observations for arXiv:2505.13031.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:58.409542Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.990159Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T03:06:29.440180Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7431ca4d-00fc-4ac9-ba5b-c7a196d5b8f2 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a51e9c89-782f-4e8b-966d-59c6fb7bd70c · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389c266b-d5d5-4b08-8564-f988b4986220 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0cfad7-a596-43a5-9889-8a122e251c33 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO r1-v: Reinforcing super generalization ability in vision-language models with less than 3
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d82be19-ace8-4e34-a7fa-675d7a3ffa2d · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766af505-288d-4f51-95e5-2adbc908af8d · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Science China Information Sciences67(12), 220101 (2024) 10
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d9523808-2f9f-45e0-8845-8718a37b2ac6 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emerging Properties in Unified Multimodal Pretraining
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83349dd2-4b7d-4060-8d87-1c24a8796109 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71cab7d0-ea13-4b25-88de-9852fedef3cd · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee365db-c203-4bf0-a66f-19a157c902cb · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 257300d6-e990-4252-a71f-e6d7c8cfda74 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2edbfb5-4dce-4b66-9659-d94dd2214f9d · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems pp
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 25e2075e-38a0-460a-90fe-77141058ab01 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3fcd54f2-d329-46df-825b-d4fff51a8731 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7528f7ae-75b6-40a4-bb11-02e6e231d17d · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf4f96a-731d-42e2-8f39-e4b1da2c1424 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90a0824-2081-47b8-928f-21aa3cabc3d1 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f7ffe76-90bb-44b7-ad54-10d5e720510a · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ARGS: Alignment as Reward-Guided Search
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4882627-c6d9-4a15-b758-8c0764825b81 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Training Language Models to Self-Correct via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970508f6-7142-4887-a75f-fbc7eeecfd46 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18aca474-96d7-4523-ba18-750605542886 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-OneVision: Easy Visual Task Transfer
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a3dfc0b-2b32-446e-a4c0-85130968bbd0 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97f83c26-75db-462e-ab39-20c7d35cdd82 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7598a37a-d303-4a7a-b95d-38b2f95001f9 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://openai.com/o1 (2024)
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1a45b6d6-9c64-4ef2-be54-7642c52af2ef · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 291b28c4-8111-430b-9d8c-04cdec535f42 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfer between Modalities with MetaQueries
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85580cca-6308-4b8d-8487-64a66c2ff527 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919c26c5-2940-41b1-b90f-853ae5be2a05 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2050e8c-61c5-498d-b050-32f0d785b733 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec67038c-8d8a-4cdd-b29e-5fc2732a4d94 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a7daf4-7488-4264-a553-191935772155 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd21ddc3-1885-4e10-86b5-078be8500a1c · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO URL https://laion
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dc57ee99-1fd9-474c-8474-4f5352a5cc1e · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201000fc-ad77-4f4e-bc67-e3e3d979dec3 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c262794-8a07-4f9e-bede-c27c41c66ddf · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems36(2024)
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c803fed-c3c5-4f6d-8871-30da9ed25ac7 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4632963-3123-4104-b269-d512bd4d163d · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a15ab1-62d7-4d53-9d89-e417288b8476 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu: Generative Pretraining in Multimodality
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29421bb-db06-4785-8940-54aa1b3a4209 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b03bf8-fcd7-4704-bded-6c384a40c4e6 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://qwenlm.github.io/blog/qwen3/ (2025)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24f5c3f4-324c-4a48-9a70-8ebbefc79636 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73daabc-cea4-4e13-a09e-18dec93b933a · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3230d9c1-e31f-46e9-b892-1fe14743546a · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu3: Next-Token Prediction is All You Need
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4351010-997d-4a3b-badc-15655ea7642b · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5335c9a1-029e-479f-bd32-a2e70ac05448 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in neural information processing systems35, 24824–24837 (2022)
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd01c75-0378-4401-9204-80ee2baefb7c · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be61ba36-fb07-4958-b08c-b7afd65ebeba · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO OmniGen: Unified Image Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53886b6d-17f0-4ff6-9054-18d186fb1d29 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd3e111-3b5e-4f16-a112-97f3ec83ebfa · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems37, 75329–75354 (2024)
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8245d0-d674-457d-bc56-822780fdfde5 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b46e205-a7f2-4b77-bdad-2522772c0a0f · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d036f42-8e5b-4d3f-914f-31e1b6e806bf · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa0872c-e1fe-4ea6-a846-d49e2085b19a · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Forty-first International Conference on Machine Learning (2024)
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f0d08762-d7eb-44f3-9d89-59bb95314e85 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MMBench: Is Your Multi-modal Model an All-around Player?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343a32e8-aecf-439f-acaa-75b61d7c7461 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da049e6-d79a-4250-93a2-0ed83facdd20 · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aafa572c-8539-4cdd-bbf2-ea8757a1f41e · outbound
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bdc7cd4-ff7c-49a9-905d-1c637f54ce25 · inbound
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b66d188-aedf-42e4-9b63-f192e703d310 · inbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15fdf608-50a4-443a-b870-878a7cd630e2 · inbound
LoRA-Gen: Specializing Large Language Model via Online LoRA Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e7eaad-27fb-40a0-a9c9-acdb4d3b6e1d · inbound
Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0967b82-5fb2-430d-ae00-a9f2fe1aabca · inbound
Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c4f5ff-d6f3-4ac4-b784-185498c4a051 · inbound
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d7b75b5d-a3ce-440c-b68a-90040d96e677 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f35ec559-82e5-404d-9f3a-262d9319b7cb · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c0041f6-f342-4a24-b0d8-2bfd251871a0 · inbound
RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6585bc5-4f90-46ac-91cc-5b059183f3b2 · inbound
UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.