Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:28:57.123789Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2607.27787.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:28:57.123789Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0f0a37e7-5dd7-44d0-9cc5-833ec769b190 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Tina: Tiny Reasoning Models via LoRA
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e50381b-8495-4fee-8f2e-e3fd4db4bd9d · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2506.07527 , year =
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db1f480-06ca-44b2-9dda-d0bf5756f576 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed2134da-3e3b-48b4-a765-437228c091dd · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2603.23871 , year =
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dfaeff7-2e76-4d4a-aa44-f287a72a3801 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d577706-0b84-40fa-a279-107c03bb98fe · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8666ee-9c78-461b-8061-8a6766171add · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a609134-d9db-4a73-af55-5bd7091cdc5c · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts LoRA: Low-Rank Adaptation of Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86296fed-8a99-430f-88c2-a5cd2006e074 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fd49f8-e1e1-40ae-961e-20159b36095b · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab03436-8a35-4566-ac31-d16a11b3dfbd · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Neural Information Processing Systems , year =
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2ab55e9-f415-4fd1-b962-59834df789e2 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts International Conference on Learning Representations , year =
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ec18f94-a102-46f5-b48b-f740d5d70b10 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts International Conference on Learning Representations , year =
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a322e6-43e8-4c23-82bf-bb4f8775ecf0 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Evaluating Large Language Models Trained on Code
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4495b5d9-cd5e-4933-b98b-7ca4606f0f15 · outbound
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994c961a-99d7-4a9e-8378-21d7cd9ceb46 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec73280-de7c-4484-9790-8ab6114ce01c · outbound
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a28dfa1e-ec7f-4daf-9161-39c79fab2afe · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2602.03143 , year =
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44774df2-16d2-449b-ad90-d338b1fb4e2f · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv preprint arXiv:2604.00698 , year =
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36266788-8891-4384-989c-d055085520ef · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Nudging the Boundaries of
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96aa2ad7-4baf-4a65-8b3c-8389cd68567b · outbound
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44efb7f-7866-4efa-bc4f-8b03875d0489 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8002cd4-72a9-4660-8c2a-48c472e5bda0 · outbound
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation febe30bb-ebf8-499c-800f-b25b943b293e · outbound
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57629963-d642-40e3-bd50-702dc651009b · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts and Jeon, Myeongho and Vu, Kim and Lai, Viet and Yang, Eunho , booktitle =
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2383483f-77f7-4825-baa7-7a6489d929fa · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Don't Waste Mistakes: Leveraging Negative
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8371c59e-caf1-4dd7-8099-7f61e6cc8e03 · outbound
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867c3a41-5dea-48b4-90af-b8473fb3ac32 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Learning-Zone Energy: Online Data Selection for Efficient
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2ddb01-de73-45dc-9642-0555e779ac66 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Advances in Neural Information Processing Systems (
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b379e5-4de0-40a8-8c4a-8349696534d3 · outbound
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9fe803-993c-41df-8b11-df54279391a9 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3af5fb-81d3-4b49-8cb9-37789a57780c · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Reinforced Self-Training (
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08879f88-e5ab-4680-8ffd-5309f37c4f48 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Transactions on Machine Learning Research (
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce414c73-cf32-4c85-863a-2c5b02663067 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Learn Hard Problems During
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e738eace-afee-47af-b3bc-b827e9956532 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Robotics: Science and Systems (
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08a5c976-f6cc-4178-aab0-586be606dc7a · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Advances in Neural Information Processing Systems (
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f6980c-6371-495c-8c23-170743bfbe5d · outbound
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5355ae5-d164-4197-8444-a5fd2b447f52 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Process Reinforcement through Implicit Rewards
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c9b497d-6b11-4188-b807-6a05c317ccd3 · outbound
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0005115-698a-4cb8-bfa5-8e2ecc2e6567 · outbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts International Conference on Machine Learning (ICML) , year =
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.