Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:01.084911Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.07581.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:01.084911Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bcc858e9-2c85-480e-b2b0-86ffc1b3ffec · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd166a4-ba88-4690-96f9-c1abb37e5db3 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afdffb96-1086-4ee0-ac21-d9691563c3d1 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ebee1d-af1e-4d55-962d-4173f2797270 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b48126ed-e5b0-4c55-9e1c-282650858c1e · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 253aa727-6c76-42a7-95bc-4e52d27df8c3 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d123c6-4c8c-45ee-a00a-f1511d78138a · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267c2edd-0536-43d5-abcb-1d548312f3af · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e568da-3383-4219-aed2-934e79408e10 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4a4c56c3-cc2e-4222-986b-f8c8e49df15d · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 113237e0-7488-4f53-ba33-eab608327b9f · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97492cea-090f-4037-a82d-ea53535adefd · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 23605e2c-8ddd-4081-9b79-0d8fc15a111f · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7060bc9-2abc-4d04-9b92-81cc05cc0c4a · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models OpenAI o1 System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaad2af5-b7df-4a2a-bb56-63acccfec38b · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c826a07-a80f-46a4-b267-3abea1d0865e · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d5b877-3a00-45f1-be99-124b6b875cb9 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba49155b-49b8-4ecf-87a5-ee1ca0475d31 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a0a754-b403-4527-be31-829b8138caa7 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21633ddb-1b6a-4adc-bebf-3f93f69a07f2 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8b4c6a54-e0d3-4daa-86c6-0a10897c94ef · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4ccb93-5a99-4f34-81cd-c8dca86021e4 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fd55159-d00a-4060-885b-2c879ece720d · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bf870af3-7ec5-4309-bf4f-d650250b6261 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30159e5-8b9e-49f6-b35e-03a37328e9b4 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37162f79-8e37-45d3-90bc-e8e0baec79db · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd410201-fb9e-493f-82c3-dd805918cdae · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b83e7a8-4513-484c-9150-04db192292b0 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21dea8d1-a349-4d73-afc9-6085a711eb7e · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8403012f-6a3b-44ed-aa5b-92cdd066c3d7 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cbd1c43-68bd-457b-9c40-68ad9b565561 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c5df10b-470b-4ac5-98be-8c1270150d93 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ec1b88-93e0-49d3-af6a-f1c2b3b09c3a · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1e23d8-8ba2-4c1b-8af7-a4fded0586a9 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4032bd1c-5083-46a1-a163-72ec5eaf19d6 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation abf73bd4-8eb0-4be9-bdda-dc5ad1175fe1 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 170634ff-68b2-45db-8775-ec9d71c43d34 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5a108e-0298-4566-8b53-50a1771d1f31 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Group Sequence Policy Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ac86ed-8a9f-4448-be1f-c4ba65ce1736 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b9b55b2-539a-4808-97f4-9198ae0094d1 · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6179ed9e-bd58-467e-a60d-6adb24ec082f · outbound
Multi-Branch Policy Optimization for Multimodal Large Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.