Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2503.01491.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:19.906869Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 753f2d48-65e1-43e0-b6fa-b36884c6ae38 · inbound
DAPO: An Open-Source LLM Reinforcement Learning System at Scale What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c60ecf9-6bdd-416f-ad2c-0f8b063fc3fe · inbound
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b87ec276-4947-4f84-ab2a-7014f9912e0b · inbound
Reinforcement Learning from Human Feedback What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e4056bbd-5d2a-4514-81dd-aafef7c0700f · inbound
Reinforcement Learning from Human Feedback What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 149
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac2e47b-4999-44bc-8193-6a25d9126a7b · inbound
Reinforcement Learning for Reasoning in Large Language Models with One Training Example What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc970383-16f6-4198-8620-b50c5f323077 · inbound
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 137
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a248983-7160-41f9-bc5d-cbb54a625ba1 · inbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 153
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdbef30b-f367-4bef-a478-724a06886ebe · inbound
Seed1.5-VL Technical Report What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 30caf4b3-60a0-444d-8f18-cbce39e87a96 · inbound
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a89e05c-8fec-4938-b143-5e4d83dab744 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bac26cf-a7c9-457e-b4af-6384c8746ff6 · inbound
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41185ef1-a4a3-4ab5-8244-c3389629e608 · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0359e4e7-fb51-4003-a55b-23be9a60d1b9 · inbound
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · inbound
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65780489-24fc-4086-8b40-c06b37a4b90a · inbound
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · inbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a039030-0780-4570-af85-a240679e1ace · inbound
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa64f27-2935-4dae-ba5a-1f772792bbb3 · inbound
ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ac9027f1-458d-4505-a4d9-19450ba8c1d2 · inbound
The Art of Scaling Reinforcement Learning Compute for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 04809f1f-fea8-4b58-a2c3-8888db5b6d5a · inbound
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 20b51108-69a6-4faf-9107-233b10865420 · inbound
Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31ef4ebc-f0b0-47b0-985e-68fa4d6f064b · inbound
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7951a39a-8e29-4105-9e88-6df39b8ef26c · inbound
Stabilizing Policy Optimization via Logits Convexity What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66db204c-ed65-47a9-92ef-c982aa62aa79 · inbound
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bb9ead0-a0e3-4ebe-9f8d-7d3751944c49 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 94558f63-35ae-47fd-8d26-779ff6493a39 · inbound
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc66f473-1bf1-4c82-93a1-2f496ac34c44 · inbound
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2228a7e-ff07-46f9-9d1f-8490b81e1d2f · inbound
Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe4635e-d5d1-4b5f-99aa-f632a0fb50c4 · inbound
Segment-Aligned Policy Optimization for Multi-Modal Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e12b29cd-60a4-4113-aa62-cab3f6236696 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b3fcf1df-66e4-4b00-a07b-b03b4210df44 · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 26dab740-55b7-4829-91d6-748d3a6f615c · inbound
Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1394ac26-6e52-457a-a261-5452c7f99f99 · inbound
Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b9cbef82-41a3-4267-abd8-159305b64d18 · inbound
Extreme Region Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3e59a4e0-5516-4789-bd87-45ea1d5982c1 · inbound
Trust Region On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 207
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation efc2402c-5688-441b-b8ca-debbd664132b · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bf00ac9e-5b34-4693-8441-de58a793eace · inbound
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 04f6656b-bf9f-42e5-9081-5c27fb4c390a · inbound
VIMPO: Value-Implicit Policy Optimization for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 906b2229-b53a-49da-92d2-0d6d5f83132c · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 254
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6bbf5e80-5a18-4823-a0c3-2e3b10152ac7 · inbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b0504468-7755-4dd9-a45f-4f2ae2ccb475 · inbound
CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eef9deb5-9289-4a84-8567-56f20e255c95 · inbound
Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b2cca1-ef7f-429d-a4b9-b4f81fd44a93 · inbound
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af1ee5bb-ac15-47a5-a867-87075d555589 · inbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.