Pith. sign in

Paper Citation Record · LEDGER

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

As of 18 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 10 inbound Pith citation observations for arXiv:2505.13031.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13031 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:26:58.409542Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:10.990159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:29.440180Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7431ca4d-00fc-4ac9-ba5b-c7a196d5b8f2 · outbound

This paper cites Qwen2.5-VL Technical Report.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.092350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.092350Z digest=sha256:83af90e4d18a0dbf44ea1c84f97731a1cad9c280048ee171e91682f3e82d8f3a

Observation a51e9c89-782f-4e8b-966d-59c6fb7bd70c · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.099262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.099262Z digest=sha256:8acb311d4f4bb4eb035d456ac157607a5bb65c6987362865be9d0822a5249cbb

Observation 389c266b-d5d5-4b08-8564-f988b4986220 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.105088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.105088Z digest=sha256:56806cf29d853a40ddfffdf2e659a6d49fc22c5e28d450d90c1a48c2dc7cead2

Observation 1b0cfad7-a596-43a5-9889-8a122e251c33 · outbound

This paper cites r1-v: Reinforcing super generalization ability in vision-language models with less than 3.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO r1-v: Reinforcing super generalization ability in vision-language models with less than 3

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.518115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.110901Z digest=sha256:7c45f6d7bc854391f20527da05a479cb5a3bc278e26064df9c2a00fcf8ffc1ba

Observation 8d82be19-ace8-4e34-a7fa-675d7a3ffa2d · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.116559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.116559Z digest=sha256:e6c03452517a506ba3bd2fa3f2780292f56531d5212a72dc57dde7dbff6ddf6a

Observation 766af505-288d-4f51-95e5-2adbc908af8d · outbound

This paper cites Science China Information Sciences67(12), 220101 (2024) 10.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Science China Information Sciences67(12), 220101 (2024) 10

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.488898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.121570Z digest=sha256:c3e339f7eab2838d738a7d610069c792be4e528963fedfeeae5c760fc408c9c3

Observation d9523808-2f9f-45e0-8845-8718a37b2ac6 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.126897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.126897Z digest=sha256:17fcdcb07139a51cc6d9221dadb67401e2cea2e266fab7510f7be4801320e0e5

Observation 83349dd2-4b7d-4060-8d87-1c24a8796109 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.133735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.133735Z digest=sha256:0e83ff8333bbfb80702f97858849d90a9de75222fb7c72c33cfcfca894ff50f9

Observation 71cab7d0-ea13-4b25-88de-9852fedef3cd · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.139315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.139315Z digest=sha256:75c8f5e5a2861de4a535f2ef6e311a16df063ffbcfc9fdd1aa4be3c634d72a94

Observation eee365db-c203-4bf0-a66f-19a157c902cb · outbound

This paper cites Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Alphazero-like Tree-Search can Guide Large Language Model Decoding and Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.146231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.146231Z digest=sha256:5e3fcbbb17bc489c85ba489c5968be27ae9a8f7f4914382b0baff11e60e4cb94

Observation 257300d6-e990-4252-a71f-e6d7c8cfda74 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.153083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.153083Z digest=sha256:c1886f140c47630cdaf1e421d62783c78a85e18c45d8aa034cd20736a24697cc

Observation e2edbfb5-4dce-4b66-9659-d94dd2214f9d · outbound

This paper cites Advances in Neural Information Processing Systems pp.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems pp

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.470305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.158664Z digest=sha256:11ea1ff975046a9bc00494b154009486adcb15d42514b6e04e146ec2b3501033

Observation 25e2075e-38a0-460a-90fe-77141058ab01 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.450727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.163671Z digest=sha256:6184b0e0680680e6d2e11dad76cc878d148981373d927136f7233be8e8a80af1

Observation 3fcd54f2-d329-46df-825b-d4fff51a8731 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.170393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.170393Z digest=sha256:6d7bd0087ca4bb94a450ab56d335a71d6c69191c7440ab2ec7beb9123eef9e38

Observation 7528f7ae-75b6-40a4-bb11-02e6e231d17d · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.175229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.175229Z digest=sha256:4d0e11623593d17e6ff63714bf3942dd4896c57d6152d386cc0de6a6d14fc1b7

Observation daf4f96a-731d-42e2-8f39-e4b1da2c1424 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.180483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.180483Z digest=sha256:fc1356a0c8544a2c30954d673d25c777a7d5320e618adca7bda93023b6e35155

Observation a90a0824-2081-47b8-928f-21aa3cabc3d1 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.185469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.185469Z digest=sha256:544e18275dcd4e829edd9270ef3305bf3f6723d6ffad9b29d833d1bca6106cde

Observation 8f7ffe76-90bb-44b7-ad54-10d5e720510a · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO ARGS: Alignment as Reward-Guided Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.192292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.192292Z digest=sha256:4faeeee168119b35369aec242b7077a5288d1483b16efa2ba859c79ec3c34cd0

Observation a4882627-c6d9-4a15-b758-8c0764825b81 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Training Language Models to Self-Correct via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.197909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.197909Z digest=sha256:540f7d6d651172cdadb62fced4ad0bfc882eddb4b15ffabf83cbc8ab2ff087f1

Observation 970508f6-7142-4887-a75f-fbc7eeecfd46 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.203015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.203015Z digest=sha256:95f101a125e1a0c6ca5092aa1f4c6c4f117a34189ad82998f1edbe230e9ff844

Observation 18aca474-96d7-4523-ba18-750605542886 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-OneVision: Easy Visual Task Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.208127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.208127Z digest=sha256:dff25d812bbe07469c377382da32597cc61cc26ce0df1bab328f60c0f4b6abc0

Observation 6a3dfc0b-2b32-446e-a4c0-85130968bbd0 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.214608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.214608Z digest=sha256:f3ebf026a26e91ebafa91f4cc32a62fa170ea7a85d67b5ab41d12a2e37e8416d

Observation 97f83c26-75db-462e-ab39-20c7d35cdd82 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.220542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.220542Z digest=sha256:78b6ed6f9b350906bbafacd0b08c92a4ffac082ae10c850acd5091d80f67fe39

Observation 7598a37a-d303-4a7a-b95d-38b2f95001f9 · outbound

This paper cites https://openai.com/o1 (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://openai.com/o1 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.420022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.225798Z digest=sha256:925b42b12e010687033aedbe48fa601f3e965de3f83da7a798eaaa63fa765496

Observation 1a45b6d6-9c64-4ef2-be54-7642c52af2ef · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.401571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.231910Z digest=sha256:2bf703cdb46aac9715cf920dc951a045871cd625f29b04e7dee837dd832443eb

Observation 291b28c4-8111-430b-9d8c-04cdec535f42 · outbound

This paper cites Transfer between Modalities with MetaQueries.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfer between Modalities with MetaQueries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.237599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.237599Z digest=sha256:80bc092fd97f464d430a103e183a00d80c20c047fcb62ebf8a76a2bd3df02180

Observation 85580cca-6308-4b8d-8487-64a66c2ff527 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.243052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.243052Z digest=sha256:41fe9838ea595ca34a01fd8fa6389aa839989a5212b559c443ae4f1fd0d3fd47

Observation 919c26c5-2940-41b1-b90f-853ae5be2a05 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.249057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.249057Z digest=sha256:a76444203e5e6edee7140b868f63a5590f7cdf3ef6739800426d3d668bec899a

Observation c2050e8c-61c5-498d-b050-32f0d785b733 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.255579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.255579Z digest=sha256:d8d1dc72c50962c46439d93e393ef3e1835684ce709fb030182bdab6dd7cbf00

Observation ec67038c-8d8a-4cdd-b29e-5fc2732a4d94 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.261282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.261282Z digest=sha256:08641f6b2c8e2995bf8608180b0b7eb5691f882602032fbfcbdd4880c5bd2585

Observation 28a7daf4-7488-4264-a553-191935772155 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.267389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.267389Z digest=sha256:4f2532024dc276095b59781fe6001cfcc16d8e36a432627a5ca87e5aa2ea0dca

Observation fd21ddc3-1885-4e10-86b5-078be8500a1c · outbound

This paper cites URL https://laion.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO URL https://laion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.361428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.272917Z digest=sha256:4eac80bdca509790bf5d15e8b2e6daa359d6c83b9df4ccc129fb8cdaf6e34805

Observation dc57ee99-1fd9-474c-8474-4f5352a5cc1e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.278116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.278116Z digest=sha256:e9f8689fb8cf12986812abd1ad025ee88bb780155f6c0ebeb328ceb3e5048f3e

Observation 201000fc-ad77-4f4e-bc67-e3e3d979dec3 · outbound

This paper cites an unresolved cited work.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:26:59.340085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.283274Z digest=sha256:3320987dff11846575ce8d11dfe0bdbb6eb9392f14a5b708e6f66de173d7d99a

Observation 3c262794-8a07-4f9e-bede-c27c41c66ddf · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems36(2024)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.288057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.288057Z digest=sha256:5b267814a9bc49474282407d761072629caae1f74952c220540784472b3dfee4

Observation 4c803fed-c3c5-4f6d-8871-30da9ed25ac7 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.292902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.292902Z digest=sha256:f190f2a6decd3ebe1fb8a79995bbc270ed4aa993fcf2abf544ed3924ed50bb5c

Observation e4632963-3123-4104-b269-d512bd4d163d · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.298120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.298120Z digest=sha256:cc08357f5b93f526c3ec3997722e992fbc6f48876dbd0e9057ad7c52d4fc5c2c

Observation 98a15ab1-62d7-4d53-9d89-e417288b8476 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu: Generative Pretraining in Multimodality

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.303151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.303151Z digest=sha256:115dec75040eff817ede31e591efc2303dda35b181534b090ff868bb26783585

Observation c29421bb-db06-4785-8940-54aa1b3a4209 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.308208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.308208Z digest=sha256:dfdd77fe70c556dab8e1305c639ecca2180df360d8d0a61243caff189dd4e678

Observation 66b03bf8-fcd7-4704-bded-6c384a40c4e6 · outbound

This paper cites https://qwenlm.github.io/blog/qwen3/ (2025).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO https://qwenlm.github.io/blog/qwen3/ (2025)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.291607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.313105Z digest=sha256:7330639fdaa070298758fa1f658bad14d7ffdc7fc37c6a28d4ee4caab7372c58

Observation 24f5c3f4-324c-4a48-9a70-8ebbefc79636 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.318304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.318304Z digest=sha256:ac7b4129c9e235cb8101b3a15bf1abf5775abb236de9e5e332fdf6eff2c6744e

Observation a73daabc-cea4-4e13-a09e-18dec93b933a · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.325121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.325121Z digest=sha256:d7f5f3969fadf325f2bf11b69ff2a930099e0d89a05b28235405fb43656dea38

Observation 3230d9c1-e31f-46e9-b892-1fe14743546a · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Emu3: Next-Token Prediction is All You Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.330544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.330544Z digest=sha256:3b5a65f9e786ae038ef17b7c55281c2a9d54fb856201f4596c55f01148976a22

Observation c4351010-997d-4a3b-badc-15655ea7642b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.336002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.336002Z digest=sha256:283c9b7efa65e0b9dfb38c38e833c7cd78aab9f1815e18629c671412b94a4293

Observation 5335c9a1-029e-479f-bd32-a2e70ac05448 · outbound

This paper cites Advances in neural information processing systems35, 24824–24837 (2022).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in neural information processing systems35, 24824–24837 (2022)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.341793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.341793Z digest=sha256:a3545819779e2231a7d05b973527e13906542355d416ff5fc7712c3c4145553e

Observation 8fd01c75-0378-4401-9204-80ee2baefb7c · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.349765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.349765Z digest=sha256:4526f06025d74e22ebe9cd431c9670cb5f29116fb872f97b3d11eccf5d89c0e0

Observation be61ba36-fb07-4958-b08c-b7afd65ebeba · outbound

This paper cites OmniGen: Unified Image Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO OmniGen: Unified Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.355229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.355229Z digest=sha256:7bb2a7ac9163bd9d3e8803e59abad549cceb17d9c9ea02acce9f062d637b144b

Observation 53886b6d-17f0-4ff6-9054-18d186fb1d29 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.361913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.361913Z digest=sha256:34a6f6d8f97ed5f5776a88dddd749677e0810743a34aa93e449466143c654261

Observation ebd3e111-3b5e-4f16-a112-97f3ec83ebfa · outbound

This paper cites Advances in Neural Information Processing Systems37, 75329–75354 (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Advances in Neural Information Processing Systems37, 75329–75354 (2024)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.367010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.367010Z digest=sha256:7ca0ff5a8336dff0527d924fd7883f3148e05856a0a0ce59a9c8ec323987b658

Observation 1e8245d0-d674-457d-bc56-822780fdfde5 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.372378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.372378Z digest=sha256:2c801d4e68b318494af264fadccf8f984690201807f79597349e62beff846cfb

Observation 3b46e205-a7f2-4b77-bdad-2522772c0a0f · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.378298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.378298Z digest=sha256:49d03a11337fd0013aa993780712d560ccd0f4e8e1376e88696e1799fecaa673

Observation 0d036f42-8e5b-4d3f-914f-31e1b6e806bf · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.383514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.383514Z digest=sha256:473443540de5060624417063eeabbb71f949617421290d2a3152493c54c4cd24

Observation caa0872c-e1fe-4ea6-a846-d49e2085b19a · outbound

This paper cites In: Forty-first International Conference on Machine Learning (2024).

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Forty-first International Conference on Machine Learning (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:26:59.222106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T20:26:58.389387Z digest=sha256:10d641686a36fafa9716d4fe56e8257967b2e73eac4f5bda41c8c5b9994e798d

Observation f0d08762-d7eb-44f3-9d89-59bb95314e85 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO MMBench: Is Your Multi-modal Model an All-around Player?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.394455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.394455Z digest=sha256:3a19723311aed4008d7515f3300d660ac0373a7181cf5cc5456e01803a4e65ea

Observation 343a32e8-aecf-439f-acaa-75b61d7c7461 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.399332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.399332Z digest=sha256:25b5549e1ea9c2f1d56fc2736d1c46d0ab299363658216e8692bc5afa3ecdb27

Observation 2da049e6-d79a-4250-93a2-0ed83facdd20 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.404285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.404285Z digest=sha256:77e4de704abe865047d754d742135ecb815a22dec251d3abf89c9d59ab9ca4d2

Observation aafa572c-8539-4cdd-bbf2-ea8757a1f41e · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:58.409542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:58.409542Z digest=sha256:55b578aa6e8f964db1bf3bd4b6df1d4764a686f69552953ea24611d70cd2cd40

Pith citing papers

Observation 6bdc7cd4-ff7c-49a9-905d-1c637f54ce25 · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:10.990159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:10.990159Z digest=sha256:05252a03471a7f27df79dc278721891992e91074e2f3d8b823bc0376336ad6a3

Observation 1b66d188-aedf-42e4-9b63-f192e703d310 · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.624099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.624099Z digest=sha256:2e5e6b7fb5c195aaa852240a748b04a5745261657efd780100d32e7a5952ee6c

Observation 15fdf608-50a4-443a-b870-878a7cd630e2 · inbound

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation cites this paper.

LoRA-Gen: Specializing Large Language Model via Online LoRA Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:26.605018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:26.605018Z digest=sha256:f2e0e63f776c33106ff9d0680459ec4fab77d5f064aaaa456cd581a49cc13f5e

Observation 21e7eaad-27fb-40a0-a9c9-acdb4d3b6e1d · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:d07cb39f91e1f9ce59198c9e1a2d600053b3278dae8187130d427725fa02f8e5

Observation a0967b82-5fb2-430d-ae00-a9f2fe1aabca · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T02:34:00.137676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:34:00.137676Z digest=sha256:159b80125462ceb057adb6f39d026bd3eed158cda378dad703976d17b15a32d1

Observation 17c4f5ff-d6f3-4ac4-b784-185498c4a051 · inbound

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cites this paper.

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:00:50.371078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:23:50.614589Z digest=sha256:d541da70d63d03e2a4f8aab4320cd052786dfe5615f3f05b52cc0bfbb0234a09

Observation d7b75b5d-a3ce-440c-b68a-90040d96e677 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:43.694718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:74d132b10141c292190ce95ca8b3124eaffc23afdfbf751562c645b4f3c57435

Observation f35ec559-82e5-404d-9f3a-262d9319b7cb · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.751849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:b3fdac7667da5832ad332159030026a145e52f387433d9636cb7d83b38c90106

Observation 3c0041f6-f342-4a24-b0d8-2bfd251871a0 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.889877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:017a9c29dc345523e84ac1dc9c0f9c4b8f1fe8c356d1ffc2263596f67c2b19da

Observation a6585bc5-4f90-46ac-91cc-5b059183f3b2 · inbound

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation cites this paper.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.442237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:e284fcf5bdf3465c92f8d77e77f1f4b82ffd458246fed68daa31d429fec62a5b