Pith. sign in

Paper Citation Record · LEDGER

Multi-Branch Policy Optimization for Multimodal Large Language Models

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.07581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07581 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:01.084911Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcc858e9-2c85-480e-b2b0-86ffc1b3ffec · outbound

This paper cites GPT-4 Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.968747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.968747Z digest=sha256:321245ebc9aff741e7c417c186add6765d7610f50eb832851522e4d5f13a3ce8

Observation dfd166a4-ba88-4690-96f9-c1abb37e5db3 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.557255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:00.972707Z digest=sha256:e0205030b9bef107510717d9d7a8ca24bd9fb45cbe9a17db0379b8f47ff76f99

Observation afdffb96-1086-4ee0-ac21-d9691563c3d1 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Branch Policy Optimization for Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.975962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.975962Z digest=sha256:4ebd4fe01f32ee04834ec07a65b6a01e4572420e692d43db1f6fe50056369fe3

Observation 42ebee1d-af1e-4d55-962d-4173f2797270 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.979284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.979284Z digest=sha256:ab85f1822c81ae6e64f40369f5d56c6f7eace0933b3c2f406ef8945adb6c5dea

Observation b48126ed-e5b0-4c55-9e1c-282650858c1e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Multi-Branch Policy Optimization for Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.982550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.982550Z digest=sha256:6b08fac5d9db19c5819493d4347f4a7f8bca427db4df424b70c90361244cb2c3

Observation 253aa727-6c76-42a7-95bc-4e52d27df8c3 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-Branch Policy Optimization for Multimodal Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.985579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.985579Z digest=sha256:0b7935c92cd02ebf48209174f85197c84e872b99e107391f8a58a39f0acf630a

Observation f7d123c6-4c8c-45ee-a00a-f1511d78138a · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.989085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.989085Z digest=sha256:b5bcdb836f62fcf2572ab2bfd50fa80e6cb2af4a6d22fcafe7cc8e8ffba5780f

Observation 267c2edd-0536-43d5-abcb-1d548312f3af · outbound

This paper cites Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward.

Multi-Branch Policy Optimization for Multimodal Large Language Models Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.994456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.994456Z digest=sha256:32996621866d6f77c9c876584e5655ffc66448f6abe2174f967e8d40e49c6075

Observation f6e568da-3383-4219-aed2-934e79408e10 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.545458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:00.997566Z digest=sha256:dae8b6bf336949f0c40c404ff99ec818ad34a48be0f536bf400fd0d254e395fa

Observation 4a4c56c3-cc2e-4222-986b-f8c8e49df15d · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.537651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.000236Z digest=sha256:e1fd6a22c7748f6dc8f738f30152ea3aab0d5be96b11c8d66958450d44f1c926

Observation 113237e0-7488-4f53-ba33-eab608327b9f · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.003086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.003086Z digest=sha256:f81624f2a2a1ca4952ab23e7328f674b2d950af2ead4e5d4382f57f491984f89

Observation 97492cea-090f-4037-a82d-ea53535adefd · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.529428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.005837Z digest=sha256:d230f1ece60006fbaa994821a8ede8b74db48a0bd5dc8333209b9b155d3285a7

Observation 23605e2c-8ddd-4081-9b79-0d8fc15a111f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.008576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.008576Z digest=sha256:c93a87855c58240f86c72f1f48c8aa04c2f3362e1f73c42616e6fc5f8f028b5b

Observation b7060bc9-2abc-4d04-9b92-81cc05cc0c4a · outbound

This paper cites OpenAI o1 System Card.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.011369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.011369Z digest=sha256:7fb5f71a3e0017bf25df7a6fe35936c0bd016b057aca2541655bb8153a429b05

Observation aaad2af5-b7df-4a2a-bb56-63acccfec38b · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.521879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.014447Z digest=sha256:6157d4179b75c31573ed16cfc7e32a23062985dab1343e5478033f0f28284935

Observation 3c826a07-a80f-46a4-b267-3abea1d0865e · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Multi-Branch Policy Optimization for Multimodal Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.017092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.017092Z digest=sha256:190437c87cd306a41d6166d5b3c3a24386066f72e70f9361baa21c9b7e83060e

Observation e9d5b877-3a00-45f1-be99-124b6b875cb9 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.019846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.019846Z digest=sha256:b4b36391ba5181b5836ca5753cbbbde8a0fc69a1937496d6841494b4ab5be5c7

Observation ba49155b-49b8-4ecf-87a5-ee1ca0475d31 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Branch Policy Optimization for Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.022292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.022292Z digest=sha256:cb363986c21c2a721aae7a9b515eab3430e93827707676687003d76168c5f58e

Observation 09a0a754-b403-4527-be31-829b8138caa7 · outbound

This paper cites Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.025134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.025134Z digest=sha256:46003993128bf292ba217b7b4c30e276c311785d333cb75e797093e0a35d9f0f

Observation 21633ddb-1b6a-4adc-bebf-3f93f69a07f2 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T00:36:01.513081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.028118Z digest=sha256:b2f9c5e7149d005d9164709cb4cc5f19892dc33661b0d72f47807da0f2f1216c

Observation 8b4c6a54-e0d3-4daa-86c6-0a10897c94ef · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Branch Policy Optimization for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.030834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.030834Z digest=sha256:2aea34e57f66cde3034122def634b11119744b7df238a3155560a36404df0bdb

Observation dd4ccb93-5a99-4f34-81cd-c8dca86021e4 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.033818Z digest=sha256:2c9e2c950d7e1008908c8159969f0c74d934cc968c50deb2444de7f918c32acb

Observation 1fd55159-d00a-4060-885b-2c879ece720d · outbound

This paper cites When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation.

Multi-Branch Policy Optimization for Multimodal Large Language Models When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:36:01.312293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.036574Z digest=sha256:0849e61bef7bc98fd1ade5fe1c4af21956d3215cde75f376e4b042326a7da0bf

Observation bf870af3-7ec5-4309-bf4f-d650250b6261 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Multi-Branch Policy Optimization for Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.039435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.039435Z digest=sha256:867af95cf81105e269805261ad1a6c736f14ebdd8dd7cc4dfb4de5faeced2eba

Observation d30159e5-8b9e-49f6-b35e-03a37328e9b4 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.505262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.042321Z digest=sha256:2ec1bba916c5096903e1deb51e01e7bafb3bb9b5a56530e3e4376437dd960ee9

Observation 37162f79-8e37-45d3-90bc-e8e0baec79db · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.497435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.044931Z digest=sha256:62a35245cfafaf4c752b388921d48262ecdc0e24e20169b150622ba0bbda8e6a

Observation dd410201-fb9e-493f-82c3-dd805918cdae · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Multi-Branch Policy Optimization for Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.047616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.047616Z digest=sha256:9ec82c76134fde16fde14a8a2b38a60e7e45281e84646124a3b8cb9d89dd574f

Observation 6b83e7a8-4513-484c-9150-04db192292b0 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.050562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.050562Z digest=sha256:3595570809eb1022f6410de3cc3f79f8fefd48a240f6aa9b75eee9ffb637d707

Observation 21dea8d1-a349-4d73-afc9-6085a711eb7e · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Multi-Branch Policy Optimization for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.053166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.053166Z digest=sha256:5e233c396ab79a1493c5d3fcdc2e4d5c5abdcaa130f45585675f63c7dbda3be7

Observation 8403012f-6a3b-44ed-aa5b-92cdd066c3d7 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.056098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.056098Z digest=sha256:12cb4206909f66c81a9ab4803e59fbe76d2a18bf0b89b07297437848fb0590fc

Observation 4cbd1c43-68bd-457b-9c40-68ad9b565561 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.058958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.058958Z digest=sha256:4a50add6f0d0c1658e5c7d10008db48d3b6cce3b8019e700ce4b43306de9d31d

Observation 7c5df10b-470b-4ac5-98be-8c1270150d93 · outbound

This paper cites R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.061794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.061794Z digest=sha256:34d70c15ee7ffbadb3768ff03af6240c34c7414836178aeef1372c1b7c8d8233

Observation 74ec1b88-93e0-49d3-af6a-f1c2b3b09c3a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multi-Branch Policy Optimization for Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.064518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.064518Z digest=sha256:73b0763d27399d8c6bca0ade6fa759ee7eb1eebbbfc698909a419da66cec6732

Observation 3f1e23d8-8ba2-4c1b-8af7-a4fded0586a9 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.067584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.067584Z digest=sha256:9fea839847707d0076696c724e43a7a9661c2f7ea890aea72df8def2c057dca5

Observation 4032bd1c-5083-46a1-a163-72ec5eaf19d6 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:36:01.485564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T00:36:01.070552Z digest=sha256:44a9ea259c77f9dadc9e1a2a422788d01b7ae9ed2b71b44b06626f8ceae9c011

Observation abf73bd4-8eb0-4be9-bdda-dc5ad1175fe1 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.073147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.073147Z digest=sha256:ccf8c83c3903c8fa961210a404ad283fff90887ae7b789591a3c0e818e07f5a9

Observation 170634ff-68b2-45db-8775-ec9d71c43d34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.079000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.079000Z digest=sha256:25d1df0d0f35dded86bbf2b26af43c8c45d3097885b602156f2be66bbb998d5e

Observation 9d5a108e-0298-4566-8b53-50a1771d1f31 · outbound

This paper cites Group Sequence Policy Optimization.

Multi-Branch Policy Optimization for Multimodal Large Language Models Group Sequence Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.075961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.075961Z digest=sha256:7653c4f13097a3409dc68998aff9221289d25ed94e037a70cc019bf0c76dd262

Observation f6ac86ed-8a9f-4448-be1f-c4ba65ce1736 · outbound

This paper cites an unresolved cited work.

Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.084911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.084911Z digest=sha256:84ada55e9cc24f1ad85f1a4b2b5f68da8977eedb2cc390a2e2b5a77892700435

Observation 9b9b55b2-539a-4808-97f4-9198ae0094d1 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-Branch Policy Optimization for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:01.081871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:01.081871Z digest=sha256:d19f58d39376e0003037332a10798342a050f66f155fe68072a34e7601d01aa8

Observation 6179ed9e-bd58-467e-a60d-6adb24ec082f · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Multi-Branch Policy Optimization for Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:00.991604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:36:00.991604Z digest=sha256:a0e7e1cfd569779546b8d0e31b358c5ade1f2aa5c39d5e595343915a3e5b3b8f

Pith citing papers

No inbound Pith citation observations are available.