Pith. sign in

Paper Citation Record · LEDGER

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:2506.21655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21655 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:29:58.706228Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:06:38.215137Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:03.800004Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c96e234-31ec-458c-8f5d-8bb24e67d18e · outbound

This paper cites Qwen2.5-vl technical report, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Qwen2.5-vl technical report, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.781638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.888597Z digest=sha256:7cebc1d6fe2347d813874515d2f4b54fb5ea3d428112ed6a4cbb35ad438c14f4

Observation 327722d8-5644-4639-9b27-3bed28a57ba1 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.735122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.904495Z digest=sha256:a98987b9f8727e1441b9949a7b80bfe147a9765b800edc29216e609265377cbe

Observation d09779fb-616a-4f88-9d59-238be952b6ae · outbound

This paper cites Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.924914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.924914Z digest=sha256:ac6cfa8e11c185ca6435bc189e9c089040a6d8a8163a125446a9237d5e4f8599

Observation e81456dd-7cec-4c91-9ef0-abfb61a3f24f · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.932417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.932417Z digest=sha256:7272f1ee6db9dff0dc1f6a12f710c3a74f983fb1a30e8bd968951954663696fd

Observation 016fa379-0fcf-4893-ad9f-fe665d6a571c · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.652681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.955233Z digest=sha256:ffd569186645c78ec1807ca40a199218f747a8964fb753c5b6dcfc370fa95a23

Observation bbd749a9-6891-4608-86d9-ae3e8ea20915 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.974697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.974697Z digest=sha256:5c27b90f1d529037448e5bcc91287aab576bcd0b4c061fdde60d9148bbc8d515

Observation 5fdd68c9-9932-4818-b4ee-3ff99ee3e845 · outbound

This paper cites Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.988706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.988706Z digest=sha256:9d4d4ed386ed8c9a3921d3b1e159b8a69a90f793f74e50bb408114f39632d6cb

Observation e818cc55-f350-4991-8670-0a8a355a27a5 · outbound

This paper cites Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.614767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:27.994310Z digest=sha256:ec7907f5fb8e3b6df3fa89d2cd4028469389c5eb9e22d1dcb9e013bb70e6a35d

Observation dbc04aea-902b-41ec-8568-3dd0883b44c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.009078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.009078Z digest=sha256:2f798a8c014cc388b516dedfbdfe9e0713e540ef5bf54b5005c496a3e47b9f85

Observation 80de8bee-487e-4ee6-9186-82df041b8cbc · outbound

This paper cites OpenAI o1 System Card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.028456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.028456Z digest=sha256:949ec426be2d10eb0ef71008d6854fcdc6fd72b6879c335ce18ad52e240db132

Observation 3ec7c38b-6828-4b97-88fc-9fa899b7ad1f · outbound

This paper cites Figureqa: An annotated figure dataset for visual reasoning, 2018.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Figureqa: An annotated figure dataset for visual reasoning, 2018

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.544330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.044738Z digest=sha256:e01dc91c462bd9f63b77dc28c2f9a6b9c36746ede14a4e7592476cbbcf72380c

Observation acf61fb4-74fe-426d-b3ee-ff087bf9ce08 · outbound

This paper cites Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.480854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.050376Z digest=sha256:e9eaf0d1dfa1a6d819da7cf2aba3daaf9b3d14e682e9d6b2b0a5e947e36b7451

Observation 12dfd25f-c2a2-40f5-a76f-5eb6dfb29b1f · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.429329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.060256Z digest=sha256:cf77a8543f353140b4b0586f72a9040042c6ee4e6e4614affb55e69f845a81ad

Observation 0f5e4e9e-3093-48d0-a5bb-64c12784dc1d · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.079295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.079295Z digest=sha256:86848fff8fbe165496b1031a3d20c319d9823d3e8caa92e66c18fa327c4bac7d

Observation 56df06e3-2475-4ddf-ad5e-8574e58b8645 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.095636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.095636Z digest=sha256:84915840b7f227e54582fd6ed27800c4664793367040da78684eebd2575fdd1b

Observation 206b3455-9ff0-4cf2-bd4a-da9688c92399 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.113236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.113236Z digest=sha256:76c681ade90944eb310d7464b97571ca495e9226574912389dd1804025a5bc2d

Observation f3aac8ad-e753-4ebc-b15c-ffb64e941b4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.132413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.132413Z digest=sha256:31a78704cadfad5c5477919174cbb933d617cc55c5b2223d0bfb3f3aabc9253e

Observation 03d23385-d3e8-4bdb-8bf6-8e49c2346561 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.361633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.149495Z digest=sha256:aa316a429777a13a81fe068f93cbc785307a25fad439e174b8dc27ed1a3cd062

Observation 71e362d7-4cd3-4a1f-a580-8f5c901f3d14 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.158314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.158314Z digest=sha256:4ebec84c03c40071bfbc2ae31da6680a6bb3d8a0387a2a48720f8bb403d77712

Observation f952b1c0-258f-48bb-aeb8-4b3a4d6ae0ca · outbound

This paper cites s1: Simple test-time scaling.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.186779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.186779Z digest=sha256:8517e0ddac51fe58682cf32df62785558446487f333b14e9a62d6e0e14c2250b

Observation 86fc2708-ff5b-4708-8c98-7bd8c78a9eb8 · outbound

This paper cites Gpt-4v(ision) system card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.211930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.211930Z digest=sha256:5f21f610b85cd5265c731b2142ec26b9fc12b2901ba35ae1cdbf4df43f0ade15

Observation dc33ca82-101c-41c5-aaa3-3e8c96efbe77 · outbound

This paper cites Gpt-4o system card, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4o system card, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.226495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.226495Z digest=sha256:4c35e704470f08fe3063d0315e5066ea6c5a7a0f42782a2668bad85b399266d6

Observation 396e1984-423c-4197-943a-369eb298496f · outbound

This paper cites Training language models to follow instructions with human feedback.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.258630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.258630Z digest=sha256:e802ee4f117a684e8698085a9a44357876d2a79882d32fa5a4cfbc6623d4bbde

Observation aa884deb-08a2-42ee-8cf7-3c4eb220165b · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.285695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.285695Z digest=sha256:40452151a814c61c7b43b42fdbdb2ef49a2cb0e177700ef4640230da5b956db1

Observation d3222508-3b10-4c05-9fcb-1cb1ca70d0b6 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.314845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.314845Z digest=sha256:c332ab305ef22acfd33cf2594c1008c37e4d80eab49dcbd07a05bf2c23235ec4

Observation 9ebabd7b-508f-4a76-a38e-41f1264898a9 · outbound

This paper cites We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.270019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.332879Z digest=sha256:b1326a0470d4824949274d877fab384b1f1f11b7644663711eda6f2be3a6f642

Observation 39c9e3c3-3bf3-4f1f-a155-c8be6a65e26a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.350191Z digest=sha256:8f4b09c4fbecf80fe1aa53d498ade4171dbf02dc4eaadf11b0352170cd7fb78b

Observation 28c8ddea-9165-4496-a0d0-ea487f9a260a · outbound

This paper cites Proximal Policy Optimization Algorithms.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.369712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.369712Z digest=sha256:39865b29c707ebcd5ab2778693e7a51bfed6450c7ef457754f2e40a06febfd27

Observation 09744d3f-4f29-4648-8b1d-fa749b4b2543 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.391054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.391054Z digest=sha256:934ca45012ea1db7eab24c57d9a47b815a63d4b342e5d9d56790eb45ca7c7609

Observation 3fb3d1a6-01cd-4bbb-aa48-a8eb1b7bc719 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.205422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:28.405993Z digest=sha256:7b2143f4f72122acb81487588a0310f6cd2b6c0da8fad6e9289322b64cb193c5

Observation df3d18f0-709c-4db1-a285-a23af471f3f5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.422267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.422267Z digest=sha256:2dd39c8235bce9a3f86772980cc74ee3592500b343fe80118991ae9e37d7ecb2

Observation 6696a32d-4608-4ef0-8de1-58e824fae9d1 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.444374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.444374Z digest=sha256:ee6cd4cf41857cfe6889af3d2449f9c7771ec1f69550fa864dddc3d6f0bf66d3

Observation fe87a75a-7355-42d1-8540-f8e5d4d0c008 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.463737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.463737Z digest=sha256:b444d06f5e772c0b4f780358e43aefe14298ff71598c323c0e712435ea27f297

Observation ee98e613-7690-4797-9504-94d2f83e0d75 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.485026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.485026Z digest=sha256:ed66e339138917cccdfd60825f869a4a64fe3b3fb03514fab4506b6da8213904

Observation f7352591-2e0a-4fc6-99aa-2ddd81b2977a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.505793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.505793Z digest=sha256:6024b5b22b4cc1f5134e6e05a713231b019da876259c62e56f798a3995d72aef

Observation b32addc3-6660-4748-8b35-0f3b99ab08f1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.514846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.514846Z digest=sha256:e0b885a5a49c8e8cd9eb8cbf6f6ba605ea7fd675c3777240ee83b0d4f5140395

Observation f559c9d4-cc7e-4725-96c6-91dcc2de487d · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.645747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.645747Z digest=sha256:c21685e831a98c136eb4aa5692a354a747b26065c18a32e5edd1753125f351d8

Observation 4f21f1cc-1561-4ae4-b7f2-662c6d7cd173 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.670138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.670138Z digest=sha256:6d86ff5c77d7b16349f30e667fff09acd2af8a4adc12d92137525e7879f4c601

Observation b9d0598b-3945-4573-bf7d-2516c8da6e38 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.699089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.699089Z digest=sha256:2880f0816e9c740bd8e4b38ddb5ca3a7612dd3cf316f10f06ea4e13f41d87fbd

Observation 9cb4bd50-cb27-4851-be18-111a1b7de6e1 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.144612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:29:58.706228Z digest=sha256:8bdeb0ec33a8ec54dd17cf1409a463d27b1dddadcd54d6abbcfe66e879511b08

Pith citing papers

Observation 42811ffa-b70c-4621-bfea-976bd4e8bfe9 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.389824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:3471e8eb02d3a749cde1e391ab255f86c78a9a5b0a40bc2c7e2f32ab2d38abc6

Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.261864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:d4a4d8bb2559b8ed6063496a7f872de60ca200d35442d8fbf2a70919b5bb5344

Observation a858a48e-99e8-4637-bbd6-e9b8dc097144 · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:29.016519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:c02c44a2a1b71feb76618e7ec510774290d3551df93c79c2c397529394f99603

Observation 16209c2b-dcd3-46ce-befd-a980d05c49f3 · inbound

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning cites this paper.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:03.801892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T15:06:38.215137Z digest=sha256:8c15b770e93d1b82ee2b72f3a52f200e235a8758c297e4fc271d0e80bdc1681a