Pith. sign in

Paper Citation Record · LEDGER

What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2503.01491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.01491 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:19.906869Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 753f2d48-65e1-43e0-b6fa-b36884c6ae38 · inbound

DAPO: An Open-Source LLM Reinforcement Learning System at Scale cites this paper.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:35:13.471029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:2e19d49307a2c8c798d324f435fd66b5d4e40d858d01705d002935a8e1d53e26

Observation 5c60ecf9-6bdd-416f-ad2c-0f8b063fc3fe · inbound

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks cites this paper.

VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:36:04.792060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T09:36:04.735688Z digest=sha256:250028e0fce525e5072bee70979d8a67457ff10ce20cf6975b987c5731ec2c0b

Observation b87ec276-4947-4f84-ab2a-7014f9912e0b · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.129490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:195b493e07ccc456f62a5c7a8c14742b0338580e407a47274bb97c9a04c52650

Observation e4056bbd-5d2a-4514-81dd-aafef7c0700f · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 149

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:19.906869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:33:19.906869Z digest=sha256:055a5ff64c0ab1b9e90223a6a897e1eafdab557a20378eb9979fe1f1f2b1f0f7

Observation fac2e47b-4999-44bc-8193-6a25d9126a7b · inbound

Reinforcement Learning for Reasoning in Large Language Models with One Training Example cites this paper.

Reinforcement Learning for Reasoning in Large Language Models with One Training Example What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:51:04.900537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T19:51:04.779597Z digest=sha256:a674906d5f570698531fc4e5e9d261011c7379d921f3a14c6de432571b600eb4

Observation bc970383-16f6-4198-8620-b50c5f323077 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.680956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.680956Z digest=sha256:4aa81830ed1708ffe3d4d24d0d5cfa565d62e78022f70dc02f25d2d1cff65725

Observation 8a248983-7160-41f9-bc5d-cbb54a625ba1 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.692497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.692497Z digest=sha256:a9e0b4687de57fbde1ce945c2d9d8ed5df098084456f32ed7d12ef9d127f556f

Observation cdbef30b-f367-4bef-a478-724a06886ebe · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.841675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:21e21f98d98f4fd514d7cf464f99aa601f0d894483535f4e384efe7815aa6236

Observation 30caf4b3-60a0-444d-8f18-cbce39e87a96 · inbound

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning cites this paper.

QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.582127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.582127Z digest=sha256:17f2a15531dd5aac631aec0dad66ae303d63edbcfe183891e79a9994b0a3eb7d

Observation 8a89e05c-8fec-4938-b143-5e4d83dab744 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:12.361455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:12.361455Z digest=sha256:2b40ba8518ebc647fd8a916283fc794cbdf2c5613ab929209b265bc1a01bffff

Observation 7bac26cf-a7c9-457e-b4af-6384c8746ff6 · inbound

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles cites this paper.

Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:08.843880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:08.843880Z digest=sha256:75fc5e572cfecb85b973bb1b6c41cfa3898903a72a0d0008058249b031d537af

Observation 41185ef1-a4a3-4ab5-8244-c3389629e608 · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:38.437982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:38.437982Z digest=sha256:c73f8330c8dde72b1bf212d311e8c7b3135c8baf961e1c79c5359fc6d7e304e2

Observation 0359e4e7-fb51-4003-a55b-23be9a60d1b9 · inbound

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling cites this paper.

Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:40:46.430743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T23:39:39.018498Z digest=sha256:c5ac900451d1f9b7b28884cd760ee15135a27869f8246e49c7668473e1d3bc0a

Observation d3e2fe4e-2353-4b39-b05f-cf96aee9919c · inbound

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization cites this paper.

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:55:42.604562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:55:42.604562Z digest=sha256:d3d6c24d27ff3484e77c2e98a434ad0f248e69914611e49c4684acbd5ccc2efd

Observation 65780489-24fc-4086-8b40-c06b37a4b90a · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:58.979808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:cb182519b18a3be94001e820227ef977c2842b5c3d0f32e2f3e7af80d61309f5

Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.938067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.938067Z digest=sha256:2f9c42f670f1c27c4ea0d8244ced08daf750aa62ed92520e70394a637b5d2fd6

Observation 9a039030-0780-4570-af85-a240679e1ace · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:04.288900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:04.288900Z digest=sha256:274289a92681ea144822bdf5a8eab58b887b2f45a851e2a0e65ad25c45d1670a

Observation 0fa64f27-2935-4dae-ba5a-1f772792bbb3 · inbound

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards cites this paper.

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:21:18.679910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T11:20:02.118579Z digest=sha256:6bb59b0dab44fcea3f36801fc09fbf055b0d530b7c21e9216847d3f37ec193a5

Observation ac9027f1-458d-4505-a4d9-19450ba8c1d2 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:29:14.050525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:eaa251125db9b488ad162e7b59077a8fb7e8bc3466c93c5c6dcdf9f060a81303

Observation 04809f1f-fea8-4b58-a2c3-8888db5b6d5a · inbound

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation cites this paper.

MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:15:26.747000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-17T23:13:43.754235Z digest=sha256:4328125775074d92105a6c0a434679419e3153ba6551a8400dc45a0d1f435496

Observation 20b51108-69a6-4faf-9107-233b10865420 · inbound

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention cites this paper.

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:15:17.712961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:15:17.712961Z digest=sha256:41821ee23b4ddfded81ba471f9776bec1ecd6f8e57179f8bf47a9f9925950d1c

Observation 31ef4ebc-f0b0-47b0-985e-68fa4d6f064b · inbound

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models cites this paper.

Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.030895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T14:09:26.842696Z digest=sha256:91911a53f37daf8b110a948cdc55ea7ecf0b66df9f78a1fbcece3e1b04207d0c

Observation 7951a39a-8e29-4105-9e88-6df39b8ef26c · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.739032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.739032Z digest=sha256:9eb990318af4b1a05d8c0918dcfbc310cbfa3cd54bd1565a6b94507c0950d2e9

Observation 66db204c-ed65-47a9-92ef-c982aa62aa79 · inbound

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning cites this paper.

Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.345214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:09:29.657563Z digest=sha256:0188bd62ca5ee5f6b00336016e37f818460abea88e41cf998d4403eb91a920ce

Observation 3bb9ead0-a0e3-4ebe-9f8d-7d3751944c49 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:08.079209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:00a2fd4e8a925487501de05f960aa8a91c77ec7b6f0d02a4458972910dffff78

Observation 94558f63-35ae-47fd-8d26-779ff6493a39 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:54.566752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:7347f048ffbe858a1b24253bf1001144c81564dbd9db5c0a777ee3b0e4a270b6

Observation bc66f473-1bf1-4c82-93a1-2f496ac34c44 · inbound

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization cites this paper.

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:56:01.737351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T15:45:15.962965Z digest=sha256:4336414ddae8b2066857fd1520d474b21788cb2cdfa13fbce075f76694e0a8fd

Observation f2228a7e-ff07-46f9-9d1f-8490b81e1d2f · inbound

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization cites this paper.

Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T20:58:19.104482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:58:19.104482Z digest=sha256:2f819bd87695fb0624f6b5bcdf8d28b9649028a6bdd933d6d6f53b6eb72166ae

Observation 2fe4635e-d5d1-4b5f-99aa-f632a0fb50c4 · inbound

Segment-Aligned Policy Optimization for Multi-Modal Reasoning cites this paper.

Segment-Aligned Policy Optimization for Multi-Modal Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.895095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T14:44:31.160543Z digest=sha256:4b21117fa749ce611ecf08567fe4be474f77787e3284a4fc9f37cdbe117735e1

Observation e12b29cd-60a4-4113-aa62-cab3f6236696 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.031255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:6e00010a24441d763009ecbbf0067470f7c18298e9677ca7a602c9d9181c4309

Observation b3fcf1df-66e4-4b00-a07b-b03b4210df44 · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.933292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:d73b237ad220b765c6a28f74fa42d7f55318487e87167b6b6c0093f952cfbe87

Observation 26dab740-55b7-4829-91d6-748d3a6f615c · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:12:22.747759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:67aa02b31b05c7a4f338e9a294ef6cc8940d7a44532331e23e5e25308f94c1f0

Observation 1394ac26-6e52-457a-a261-5452c7f99f99 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T10:01:22.994765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:b0043e32f8a6d016dfb2325c73650979a3fbbd36f76a57697c09e2f9c18b5627

Observation b9cbef82-41a3-4267-abd8-159305b64d18 · inbound

Extreme Region Policy Distillation cites this paper.

Extreme Region Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:02.202182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T23:14:56.223606Z digest=sha256:e5b84d78e7ed648233d657e881d9efeb7c670673cc84fc59795bed3c89364ce5

Observation 3e59a4e0-5516-4789-bd87-45ea1d5982c1 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 207

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.277692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:d1a3b7816efc97b714b8cd33cc15b03c66d34b45b8ce6db3f47be06923cd7d5c

Observation efc2402c-5688-441b-b8ca-debbd664132b · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.368996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:20935c62b898ee127804c8f74d9cbaa7c3a63a31299f24ce58ec9bb57c96d59e

Observation bf00ac9e-5b34-4693-8441-de58a793eace · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:30.041691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:43ea58418b117f8634fd49d5040327bbc621cde5adedcf935bb524e980073fc3

Observation 04f6656b-bf9f-42e5-9081-5c27fb4c390a · inbound

VIMPO: Value-Implicit Policy Optimization for LLMs cites this paper.

VIMPO: Value-Implicit Policy Optimization for LLMs What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.690053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T17:59:29.066879Z digest=sha256:533725b0a6de24ce45d6ce880d4b54fb87de50cf2cd8dcd81e732ad8b9e6c8da

Observation 906b2229-b53a-49da-92d2-0d6d5f83132c · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 254

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.705178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:c4d493e7d077475c7e9135a8c10bfbfaa2b4c3fac551789a6ddadb9061b8a1b6

Observation 6bbf5e80-5a18-4823-a0c3-2e3b10152ac7 · inbound

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents cites this paper.

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:03:48.691291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-07T13:57:33.821970Z digest=sha256:1cb73776bb04328680d6daa3e32dc45471e3be27cff7f0404b161baa82e5be53

Observation b0504468-7755-4dd9-a45f-4f2ae2ccb475 · inbound

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents cites this paper.

CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T14:03:48.463323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-07T13:57:33.821970Z digest=sha256:14221f3b3d4a5c7e1c7c9eafcbc0c4c810cbb0676bfa6d87ba172b7c6b8d269f

Observation eef9deb5-9289-4a84-8567-56f20e255c95 · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.706606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.706606Z digest=sha256:1ef8b24591eea003278f4ff3a0c69a13b3777911e7f30731a760973f811b5b65

Observation 47b2cca1-ef7f-429d-a4b9-b4f81fd44a93 · inbound

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning cites this paper.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.895805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.895805Z digest=sha256:280cd3a049b2dcc3eadf1c026b23361e29b83f0cffc1252a40dca8c235f41445

Observation af1ee5bb-ac15-47a5-a867-87075d555589 · inbound

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning cites this paper.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.992560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.992560Z digest=sha256:60e19907b4bf1bcfeca6414ac19320acb8379740ed0ee48181f8b17095b90761