Pith. sign in

Paper Citation Record · LEDGER

Multi-Branch Policy Optimization for Multimodal Large Language Models

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.07581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07581 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:01.084911Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcc858e9-2c85-480e-b2b0-86ffc1b3ffec · outbound

This paper cites GPT-4 Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.968747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.968747Z digest=sha256:9a487f2d0e9430e6732fd90ac5d3df7835d584388e0314fb50d4d0718b351825

Observation dfd166a4-ba88-4690-96f9-c1abb37e5db3 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.557255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:00.972707Z digest=sha256:1bb439a11926334b125681b31d8380dc5546fc945df96c3da08abf484640a75b

Observation afdffb96-1086-4ee0-ac21-d9691563c3d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.975962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.975962Z digest=sha256:c4705b9e2e0a28049434d50d055c5f4d5a1654496500abcd0ed3eb1141a3df8d

Observation 42ebee1d-af1e-4d55-962d-4173f2797270 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.979284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.979284Z digest=sha256:5b8b7e6c7a39f394a27ab7f0c3142ede757419489ad0aa8777dc1af35dc9edd9

Observation b48126ed-e5b0-4c55-9e1c-282650858c1e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Multi-Branch Policy Optimization for Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.982550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.982550Z digest=sha256:a95b1f65d500c0ec8636aa2583002691f334ea75e17d60752267d0d054de77ae

Observation 253aa727-6c76-42a7-95bc-4e52d27df8c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-Branch Policy Optimization for Multimodal Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.985579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.985579Z digest=sha256:21ae0fff2894f53054035c65b329a51b865489f3226d8c11bfe003e8e52c465f

Observation f7d123c6-4c8c-45ee-a00a-f1511d78138a · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.989085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.989085Z digest=sha256:1608a6ff98136b24a34321000c42b155dac9706d610f1be039d1991926bf3ea4

Observation 267c2edd-0536-43d5-abcb-1d548312f3af · outbound

This paper cites Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward.

Multi-Branch Policy Optimization for Multimodal Large Language Models Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.994456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.994456Z digest=sha256:76546423ff14f01c2238ea909d18adcc1e2163982cd603d7084f509740d8061a

Observation f6e568da-3383-4219-aed2-934e79408e10 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.545458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:00.997566Z digest=sha256:22fcaad2b1601ec1cf43d96b70bb8d5445b04a9b8ac7e1272d5314050fd47e38

Observation 4a4c56c3-cc2e-4222-986b-f8c8e49df15d · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.537651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.000236Z digest=sha256:81b3ae8b466e7d949c1358f7d67cf0e8713dd9210e7f830169fbd0a29f35fdb7

Observation 113237e0-7488-4f53-ba33-eab608327b9f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.003086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.003086Z digest=sha256:913b79c65f39357c4eebf69470d5f81de891911f45c53bcc848ca9b31acc0e71

Observation 97492cea-090f-4037-a82d-ea53535adefd · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.529428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.005837Z digest=sha256:9e7213c7648535aff5e31a84d415f23f3bfa575f77a61c80e1a9b53b02d5bb6e

Observation 23605e2c-8ddd-4081-9b79-0d8fc15a111f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.008576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.008576Z digest=sha256:6eee3f315e69b72422021a629781974eb3b9dc4d2debb744e6496f4da27ba059

Observation b7060bc9-2abc-4d04-9b92-81cc05cc0c4a · outbound

This paper cites OpenAI o1 System Card.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.011369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.011369Z digest=sha256:2963688ea831c0ba5ec8b49f725680f9037b830064e8e8a395a9d4447498785d

Observation aaad2af5-b7df-4a2a-bb56-63acccfec38b · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.521879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.014447Z digest=sha256:3661743944ffe9aa23356b894b6131eedb3a7a29bd20d9f39e9d190641dd5c7b

Observation 3c826a07-a80f-46a4-b267-3abea1d0865e · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Multi-Branch Policy Optimization for Multimodal Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.017092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.017092Z digest=sha256:d1bcc86bec35928aab636988cc5f489279ee9ee54eda44fce022364190ac3fb8

Observation e9d5b877-3a00-45f1-be99-124b6b875cb9 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.019846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.019846Z digest=sha256:a255f7d4cf084821820acf6ed750548666e66c009508e4e4ebd24e18ac6cdb06

Observation ba49155b-49b8-4ecf-87a5-ee1ca0475d31 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Branch Policy Optimization for Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.022292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.022292Z digest=sha256:566ae12c9b39609f88f18ec6ba7834f4bb740eade2227e271721df544803eb31

Observation 09a0a754-b403-4527-be31-829b8138caa7 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.025134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.025134Z digest=sha256:00d520063b74abb27888bdf7b290acd70960e06b863dcc85384b0a5e4beb53fe

Observation 21633ddb-1b6a-4adc-bebf-3f93f69a07f2 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:36:01.513081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.028118Z digest=sha256:450bf1b8afa8f458b6d5ddbc54d06aea00f03232bae2b7b56da1f396ae750619

Observation 8b4c6a54-e0d3-4daa-86c6-0a10897c94ef · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.030834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.030834Z digest=sha256:90b68457617a652ed0cbae794b3792c00b61020d19731f2cbe2bf57362adf070

Observation dd4ccb93-5a99-4f34-81cd-c8dca86021e4 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.033818Z digest=sha256:043e0caddd543a34c21817418b5da075756d605fe4ae1a5a0526c2a1253a0681

Observation 1fd55159-d00a-4060-885b-2c879ece720d · outbound

This paper cites When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation.

Multi-Branch Policy Optimization for Multimodal Large Language Models When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:36:01.312293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.036574Z digest=sha256:796a115b58ccf13a88c8d89157ec64ba6a419bac66f12c351e9d6eb062fc18b5

Observation bf870af3-7ec5-4309-bf4f-d650250b6261 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Multi-Branch Policy Optimization for Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.039435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.039435Z digest=sha256:4661ffed33c2680f7f1b61b01ebb46c629756fa64d3d97e99cff3d58632d8b06

Observation d30159e5-8b9e-49f6-b35e-03a37328e9b4 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.505262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.042321Z digest=sha256:74c6c6d4419755f94881ac5a123c758f21b4ca38fe009fec5e7ba75e12f0c832

Observation 37162f79-8e37-45d3-90bc-e8e0baec79db · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.497435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.044931Z digest=sha256:6f972db3fbe46b5984ff4daeba32b443c03bfe1e91a0f9da6fa0f08361c02bb6

Observation dd410201-fb9e-493f-82c3-dd805918cdae · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.047616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.047616Z digest=sha256:73513c365d7a6a9442fbbb0cb73b9098b6e3163c46c709333a079a450f952a0f

Observation 6b83e7a8-4513-484c-9150-04db192292b0 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.050562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.050562Z digest=sha256:69702373ff6c9a45b7382a22ddfa5e38b35e126113ff2540fc7eedde37913fed

Observation 21dea8d1-a349-4d73-afc9-6085a711eb7e · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Multi-Branch Policy Optimization for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.053166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.053166Z digest=sha256:700e6be2fd00c4f41499ad5add5aeb7405120628e3f0446a778519428ce14296

Observation 8403012f-6a3b-44ed-aa5b-92cdd066c3d7 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.056098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.056098Z digest=sha256:b909881a319fefd5112c5ccdbd96cbb4aeebe29afc1ba72cad8c11dc2c0fcfd3

Observation 4cbd1c43-68bd-457b-9c40-68ad9b565561 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.058958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.058958Z digest=sha256:c54304694d35f21efb770ec203b6cd1011f7a3036fbdaebbf92b990b958fac33

Observation 7c5df10b-470b-4ac5-98be-8c1270150d93 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.061794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.061794Z digest=sha256:4332a56215684920c61700d5dace797461d5e77a5c35eebd6ffa454086d0cffb

Observation 74ec1b88-93e0-49d3-af6a-f1c2b3b09c3a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multi-Branch Policy Optimization for Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.064518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.064518Z digest=sha256:af5173c0e670a37bad8865d71039ce192f6a7e892a62d9b96f2bc734fd4cc30c

Observation 3f1e23d8-8ba2-4c1b-8af7-a4fded0586a9 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.067584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.067584Z digest=sha256:18b513940accc1e46bab7dfcf0f59f7a2d624fbeeacd9b13bc41624072cda63b

Observation 4032bd1c-5083-46a1-a163-72ec5eaf19d6 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.485564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T00:36:01.070552Z digest=sha256:0a189baf79eb554dee847bfd9e2f3d0ea02cbec040f76c582abf149f8a5aa466

Observation abf73bd4-8eb0-4be9-bdda-dc5ad1175fe1 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.073147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.073147Z digest=sha256:7175cbae4b966463428469d03627c61d274f9ab1cd230ca4c77390048e0e1c96

Observation 170634ff-68b2-45db-8775-ec9d71c43d34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.079000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.079000Z digest=sha256:ea11e9ae44f8ef7deeadb7a122f4e1b017e6a96c58ab82ee74ac97ea27ff0f5d

Observation 9d5a108e-0298-4566-8b53-50a1771d1f31 · outbound

This paper cites Group Sequence Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models Group Sequence Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.075961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.075961Z digest=sha256:8f1d8a44fd76ad1379b4a6fd66f385c2a054e42cf1411b8d5f4fee643059b970

Observation f6ac86ed-8a9f-4448-be1f-c4ba65ce1736 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.084911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.084911Z digest=sha256:3050fb89465dfdcfd3519cab639d3ab1f8414e45ef0dd2faee4edaf0a0ae950f

Observation 9b9b55b2-539a-4808-97f4-9198ae0094d1 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.081871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.081871Z digest=sha256:774a09ef5b0dd12ea0e1343c02503d61cd2e1ffad8e47fcf47e14e6ec9db167b

Observation 6179ed9e-bd58-467e-a60d-6adb24ec082f · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.991604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.991604Z digest=sha256:4b8af67d431e917b73e7cbd58f5b833accd8089b7aba4623a2960232874cc549

Pith citing papers

No inbound Pith citation observations are available.