Pith. sign in

Paper Citation Record · LEDGER

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 4 inbound Pith citation observations for arXiv:2506.21655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21655 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:29:58.706228Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T15:06:38.215137Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:03.800004Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c96e234-31ec-458c-8f5d-8bb24e67d18e · outbound

This paper cites Qwen2.5-vl technical report, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Qwen2.5-vl technical report, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.781638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:27.888597Z digest=sha256:e0f5195316d9a60eab8285fb6849dde8ad0436291d6c87d3d69f8db1739341e1

Observation 327722d8-5644-4639-9b27-3bed28a57ba1 · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.735122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:27.904495Z digest=sha256:e39f0bf4c1cb8537fada97fc33f5b620a569d67a466b6aa2adb07f80cfb64ce2

Observation d09779fb-616a-4f88-9d59-238be952b6ae · outbound

This paper cites Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.924914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.924914Z digest=sha256:a2839c63b9fe3f3bc1da92eb83ec769d64230b6ec38077571218fea41f9ba73a

Observation e81456dd-7cec-4c91-9ef0-abfb61a3f24f · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.932417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.932417Z digest=sha256:774d3c4e6b43ba02f0327d6650c63885abfb00e97c9cf3aa10d85cdc35bec491

Observation 016fa379-0fcf-4893-ad9f-fe665d6a571c · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.652681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:27.955233Z digest=sha256:1d752202e8b2027df7de8de511ec2c2b088368b8210454e34f76e884719f54ef

Observation bbd749a9-6891-4608-86d9-ae3e8ea20915 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.974697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.974697Z digest=sha256:06b6d81a5d49dfbb995a067271cc3a7656aced41285b35764ba165cd945710b8

Observation 5fdd68c9-9932-4818-b4ee-3ff99ee3e845 · outbound

This paper cites Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:27.988706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:27.988706Z digest=sha256:6a1e2551208802d01f90e402521f3b417ce0b7bc8ae4e2a0720ea544f6fddc4a

Observation e818cc55-f350-4991-8670-0a8a355a27a5 · outbound

This paper cites Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Polymath: A challenging multi-modal mathematical reasoning benchmark, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.614767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:27.994310Z digest=sha256:5b8d900e0c68ed4c6fc22533febc6ff13b215eb534f0f3ebd12dc5e20e9571ca

Observation dbc04aea-902b-41ec-8568-3dd0883b44c9 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.009078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.009078Z digest=sha256:9fcdb9daa7ce6e6034163c51b6355c3fc6a6c804118fb8d067f33a8998cf1967

Observation 80de8bee-487e-4ee6-9186-82df041b8cbc · outbound

This paper cites OpenAI o1 System Card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.028456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.028456Z digest=sha256:b1a3db4f199f976f6b8e353cfd470610c666fb9da04b2f7577e956e1cd18c9f5

Observation 3ec7c38b-6828-4b97-88fc-9fa899b7ad1f · outbound

This paper cites Figureqa: An annotated figure dataset for visual reasoning, 2018.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Figureqa: An annotated figure dataset for visual reasoning, 2018

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.544330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:28.044738Z digest=sha256:6b0cdd2c5f6e6002b19c022c536418c2eae6c565929ef34cd335600de4e373b7

Observation acf61fb4-74fe-426d-b3ee-ff087bf9ce08 · outbound

This paper cites Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Scemqa: A scientific college entrance level multimodal question answering benchmark, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.480854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:28.050376Z digest=sha256:351252b0f7444ce31e38f728f324cae690edeacc57cf48464cc24b8b3f1228bd

Observation 12dfd25f-c2a2-40f5-a76f-5eb6dfb29b1f · outbound

This paper cites Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Clevr-math: A dataset for compositional language, visual and mathematical reasoning, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.429329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:28.060256Z digest=sha256:37916b9e987da8fe1f08693741a5905e1d8a9794832e86c6826af0d0b60f9aba

Observation 0f5e4e9e-3093-48d0-a5bb-64c12784dc1d · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.079295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.079295Z digest=sha256:e7bc9afcccfb6de10c8759e09f89d42a312d7f9d45364ab4aae348f4f2185f3a

Observation 56df06e3-2475-4ddf-ad5e-8574e58b8645 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.095636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.095636Z digest=sha256:552336b15cdbfc2cd129f8353b6c04c8f843ec2b62c087d69a75674efb326f46

Observation 206b3455-9ff0-4cf2-bd4a-da9688c92399 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.113236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.113236Z digest=sha256:8028852053738a5cf28852c91d1fbd900ae4d2fe0aa9aa28ffed98b95cb5b2c7

Observation f3aac8ad-e753-4ebc-b15c-ffb64e941b4f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.132413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.132413Z digest=sha256:143928c294f2c5ebf10d8a9b904fb5a2ebb1de956e3924139a1060ee7228818d

Observation 03d23385-d3e8-4bdb-8bf6-8e49c2346561 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.361633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:28.149495Z digest=sha256:bbac7657c991ca1cf26043216b0eecea68dbd6d4808a387052aa44b344509686

Observation 71e362d7-4cd3-4a1f-a580-8f5c901f3d14 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.158314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.158314Z digest=sha256:0d614344ff09cb68727832ab749a48ae131bddcae5d7aa3b1719f584b09d76dd

Observation f952b1c0-258f-48bb-aeb8-4b3a4d6ae0ca · outbound

This paper cites s1: Simple test-time scaling.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.186779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.186779Z digest=sha256:2dd47eabfb33455e9dd1bc1580fef63aad0c6048c5f7e3f828a2cfa49651412b

Observation 86fc2708-ff5b-4708-8c98-7bd8c78a9eb8 · outbound

This paper cites Gpt-4v(ision) system card.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4v(ision) system card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.211930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.211930Z digest=sha256:e46af38b96c7c89adc0bc5785a4a9166fb5eb9032a28e6e80d215fdfa14a283e

Observation dc33ca82-101c-41c5-aaa3-3e8c96efbe77 · outbound

This paper cites Gpt-4o system card, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Gpt-4o system card, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.226495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.226495Z digest=sha256:ba850eeb567f4cae43f307bddc8021f2958358ca0ba4fc370a6f1f103a4a840d

Observation 396e1984-423c-4197-943a-369eb298496f · outbound

This paper cites Training language models to follow instructions with human feedback.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.258630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.258630Z digest=sha256:74a7bb3fb7e4fa2fb616e9c73b4e5c877a6d9fca2c9372a6ff138e9ddec675de

Observation aa884deb-08a2-42ee-8cf7-3c4eb220165b · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.285695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.285695Z digest=sha256:5abc8b7d7f75a7d794b4d1c329b42fef61dd2c8fc74d92c5dac4c51f9ae805fa

Observation d3222508-3b10-4c05-9fcb-1cb1ca70d0b6 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.314845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.314845Z digest=sha256:925d729279c6b01ca35f953f024b6569dcc8a2db4f5e3141f065bcb86fd2ec58

Observation 9ebabd7b-508f-4a76-a38e-41f1264898a9 · outbound

This paper cites We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.270019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:28.332879Z digest=sha256:a1d2d886f183a48ebf2993894fd0c51735e20414a6cb1dedfa3806da1c545b62

Observation 39c9e3c3-3bf3-4f1f-a155-c8be6a65e26a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.350191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.350191Z digest=sha256:cbf8f77590586d0583f0eadb902b1b050aae44483408facb1a246c02aa28b18d

Observation 28c8ddea-9165-4496-a0d0-ea487f9a260a · outbound

This paper cites Proximal Policy Optimization Algorithms.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.369712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.369712Z digest=sha256:d131b8964009778b11bf0b9dcff0d37960342386f90ef721b937ddc0042de2eb

Observation 09744d3f-4f29-4648-8b1d-fa749b4b2543 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.391054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.391054Z digest=sha256:23e81e23b62d37bc461ead0e6bc52ea2645e5fc0772579d75cd2039dae3b796d

Observation 3fb3d1a6-01cd-4bbb-aa48-a8eb1b7bc719 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.205422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:28.405993Z digest=sha256:0e46e283cebf763795d60e73dc0416ffc548d68c58136f2e6da4e8ad9463af8f

Observation df3d18f0-709c-4db1-a285-a23af471f3f5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.422267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.422267Z digest=sha256:052e11aa83c8f1b7fed22afa3a3d61cdb3937d72ec20df6ebfc479e082a69fbf

Observation 6696a32d-4608-4ef0-8de1-58e824fae9d1 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.444374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.444374Z digest=sha256:b59cbdf8b0b420e41045226ee0b4d1196794de1719fbda604225e2c2920c8712

Observation fe87a75a-7355-42d1-8540-f8e5d4d0c008 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.463737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.463737Z digest=sha256:a1fc068fa1ca9c95ccd70dd9a93d5af09a211d42d05a5aca97bea6138ef34e67

Observation ee98e613-7690-4797-9504-94d2f83e0d75 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.485026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.485026Z digest=sha256:a4126bd76020adab9a158ee4adaf392d31a82bdb2a9e34585e95b6426f8231bf

Observation f7352591-2e0a-4fc6-99aa-2ddd81b2977a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.505793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.505793Z digest=sha256:14e5deee667a0c7220310626ca46110e14a554d550b3ccac729f1631633b884e

Observation b32addc3-6660-4748-8b35-0f3b99ab08f1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.514846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.514846Z digest=sha256:feb926e02be139ef0ca02f786ae7e77c14aa6d0d8dcc7142888657453e79deb7

Observation f559c9d4-cc7e-4725-96c6-91dcc2de487d · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.645747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.645747Z digest=sha256:7315157a4862cd91d436fb76d84cc843f06a1ac5180321ee352c56272e4cb856

Observation 4f21f1cc-1561-4ae4-b7f2-662c6d7cd173 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.670138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.670138Z digest=sha256:bc07efa30eb69e6c797066966f3cb46bcf6ba0907ff0899f0d00e45e287e0378

Observation b9d0598b-3945-4573-bf7d-2516c8da6e38 · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:58.699089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:58.699089Z digest=sha256:62e8cad5e85e4585d26a7e524ea6a0154f206de9121c10f1eb4310d38d7ba41a

Observation 9cb4bd50-cb27-4851-be18-111a1b7de6e1 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:29:59.144612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:29:58.706228Z digest=sha256:2cdbeba5bb69baa3286aa0757b0d635ca6e6b099a0527dd04a649de08aa4de81

Pith citing papers

Observation 42811ffa-b70c-4621-bfea-976bd4e8bfe9 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.389824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:3c5eb20ba552248d589d1d2031b83827b799e5486b684d8efde6ae06d99f3767

Observation bc77d68d-c518-4809-9aa7-c3db5c2fec50 · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.261864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:828d23a2a05d93d56298b0448d0e948f4b9662a0e3f8336813d0773667318674

Observation a858a48e-99e8-4637-bbd6-e9b8dc097144 · inbound

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF cites this paper.

PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:29.016519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T07:36:39.835616Z digest=sha256:9dab50276d5039002dd787a61c46dcb471b848eccbf5c8b0f5bd7f988a254bd3

Observation 16209c2b-dcd3-46ce-befd-a980d05c49f3 · inbound

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning cites this paper.

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:03.801892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T15:06:38.215137Z digest=sha256:bf43d61c49b3570079497357059609e419141fc18c561b7edb88d517f927700c