Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:29:58.706228Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:2506.21655.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:29:58.706228Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-02T15:06:38.215137Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T15:07:03.800004Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c96e234-31ec-458c-8f5d-8bb24e67d18e · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Qwen2.5-vl technical report, 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 327722d8-5644-4639-9b27-3bed28a57ba1 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization An augmented benchmark dataset for geometric question answering through dual parallel text encoding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d09779fb-616a-4f88-9d59-238be952b6ae · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e81456dd-7cec-4c91-9ef0-abfb61a3f24f · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016fa379-0fcf-4893-ad9f-fe665d6a571c · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbd749a9-6891-4608-86d9-ae3e8ea20915 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fdd68c9-9932-4818-b4ee-3ff99ee3e845 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e818cc55-f350-4991-8670-0a8a355a27a5 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbc04aea-902b-41ec-8568-3dd0883b44c9 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80de8bee-487e-4ee6-9186-82df041b8cbc · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OpenAI o1 System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ec7c38b-6828-4b97-88fc-9fa899b7ad1f · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Figureqa: An annotated figure dataset for visual reasoning, 2018
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation acf61fb4-74fe-426d-b3ee-ff087bf9ce08 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12dfd25f-c2a2-40f5-a76f-5eb6dfb29b1f · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f5e4e9e-3093-48d0-a5bb-64c12784dc1d · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56df06e3-2475-4ddf-ad5e-8574e58b8645 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206b3455-9ff0-4cf2-bd4a-da9688c92399 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3aac8ad-e753-4ebc-b15c-ffb64e941b4f · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d23385-d3e8-4bdb-8bf6-8e49c2346561 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71e362d7-4cd3-4a1f-a580-8f5c901f3d14 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f952b1c0-258f-48bb-aeb8-4b3a4d6ae0ca · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization s1: Simple test-time scaling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86fc2708-ff5b-4708-8c98-7bd8c78a9eb8 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4v(ision) system card
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc33ca82-101c-41c5-aaa3-3e8c96efbe77 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4o system card, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396e1984-423c-4197-943a-369eb298496f · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa884deb-08a2-42ee-8cf7-3c4eb220165b · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3222508-3b10-4c05-9fcb-1cb1ca70d0b6 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ebabd7b-508f-4a76-a38e-41f1264898a9 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39c9e3c3-3bf3-4f1f-a155-c8be6a65e26a · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Direct preference optimization: Your language model is secretly a reward model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c8ddea-9165-4496-a0d0-ea487f9a260a · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09744d3f-4f29-4648-8b1d-fa749b4b2543 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fb3d1a6-01cd-4bbb-aa48-a8eb1b7bc719 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df3d18f0-709c-4db1-a285-a23af471f3f5 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Chain-of-thought prompting elicits reasoning in large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6696a32d-4608-4ef0-8de1-58e824fae9d1 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe87a75a-7355-42d1-8540-f8e5d4d0c008 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee98e613-7690-4797-9504-94d2f83e0d75 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7352591-2e0a-4fc6-99aa-2ddd81b2977a · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32addc3-6660-4748-8b35-0f3b99ab08f1 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f559c9d4-cc7e-4725-96c6-91dcc2de487d · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f21f1cc-1561-4ae4-b7f2-662c6d7cd173 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d0598b-3945-4573-bf7d-2516c8da6e38 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb4bd50-cb27-4851-be18-111a1b7de6e1 · outbound
APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42811ffa-b70c-4621-bfea-976bd4e8bfe9 · inbound
Latent Visual Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · inbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a858a48e-99e8-4637-bbd6-e9b8dc097144 · inbound
PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16209c2b-dcd3-46ce-befd-a980d05c49f3 · inbound
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.