Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:39:41.431601Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.17247.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:39:41.431601Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d3639001-7ee8-425c-8e05-8ff8b063389a · outbound
Distilled Reinforcement Learning for LLM Post-training On-policy distillation of language models: Learning from self- generated mistakes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66aa2a9b-e845-4840-8ca0-79cba8c1e59a · outbound
Distilled Reinforcement Learning for LLM Post-training Reasoning with Exploration: An Entropy Perspective
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f747b952-059c-4908-b455-ed0d5199652c · outbound
Distilled Reinforcement Learning for LLM Post-training Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f308f626-b3e0-4513-abfa-ec7e4c1916f1 · outbound
Distilled Reinforcement Learning for LLM Post-training Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c158ab-115d-45df-a6df-efabde58cab0 · outbound
Distilled Reinforcement Learning for LLM Post-training ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a82ab3-edb9-4610-a60c-e8a8955ea2aa · outbound
Distilled Reinforcement Learning for LLM Post-training Minillm: Knowledge distillation of large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a80cd6-60be-4fa5-8f4f-2fd9e34166d0 · outbound
Distilled Reinforcement Learning for LLM Post-training DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4270e065-4d31-421d-975b-a6971cb08ef9 · outbound
Distilled Reinforcement Learning for LLM Post-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e52171a-0e32-40c1-80d6-66c8516014d0 · outbound
Distilled Reinforcement Learning for LLM Post-training Sequence-level knowledge distillation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2a8488-18fd-4831-9eac-02fb84f50a6f · outbound
Distilled Reinforcement Learning for LLM Post-training TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d66b7e15-bd33-47c8-b307-b7f66b3528e6 · outbound
Distilled Reinforcement Learning for LLM Post-training Let's Verify Step by Step
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e39112-4a92-4e04-897e-81a633265030 · outbound
Distilled Reinforcement Learning for LLM Post-training Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a52b37e-3c9d-4f12-9be6-b2f196349785 · outbound
Distilled Reinforcement Learning for LLM Post-training Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d209c0-ab3f-41e1-81fa-19c0cc639450 · outbound
Distilled Reinforcement Learning for LLM Post-training Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35476f98-1e81-4179-a447-6f7e5039932d · outbound
Distilled Reinforcement Learning for LLM Post-training Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4da1c71-9642-46c4-8a6e-806cf355410e · outbound
Distilled Reinforcement Learning for LLM Post-training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac733956-c1a5-465e-a790-f3fbbf16fb9a · outbound
Distilled Reinforcement Learning for LLM Post-training SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf09c03-2927-4a36-9243-d9819033cdab · outbound
Distilled Reinforcement Learning for LLM Post-training Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 823550af-5e8d-4066-89dc-f6741deba2b2 · outbound
Distilled Reinforcement Learning for LLM Post-training Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7b7c17c-c775-4d1a-b2e9-a3a38fbc143d · outbound
Distilled Reinforcement Learning for LLM Post-training On the generalization of sft: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025b
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723986cb-e5e4-4411-94a7-bf89bd911262 · outbound
Distilled Reinforcement Learning for LLM Post-training Qwen2.5 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a83a283-4a92-4c3b-82db-f4231eb2b2c1 · outbound
Distilled Reinforcement Learning for LLM Post-training Qwen3 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8acea5ca-4c80-4681-bac4-47206eb3b3dd · outbound
Distilled Reinforcement Learning for LLM Post-training Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb87e227-e96b-448c-96c4-342c584747ee · outbound
Distilled Reinforcement Learning for LLM Post-training Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9216b88-e480-4cc5-add2-b25348642673 · outbound
Distilled Reinforcement Learning for LLM Post-training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 866ed5c2-5541-4e61-b98d-380337bf4304 · outbound
Distilled Reinforcement Learning for LLM Post-training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee81150d-a934-48b1-8599-f78542a401ea · outbound
Distilled Reinforcement Learning for LLM Post-training DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419e3584-e959-4a1f-b782-c0a758894060 · outbound
Distilled Reinforcement Learning for LLM Post-training Consequently, the training procedure cannot directly determine whether the teacher is capable of solving a given problem
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74263fe0-6d6e-4c38-9125-3c5ac7ae9bf6 · outbound
Distilled Reinforcement Learning for LLM Post-training Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1d7b461-44bc-4f7a-b38b-99ed442eba31 · outbound
Distilled Reinforcement Learning for LLM Post-training Entropy-Aware On-Policy Distillation of Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b2c343-f33b-4bcb-8a45-fe710ceef832 · outbound
Distilled Reinforcement Learning for LLM Post-training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a53cbf-ad2a-43d3-a965-8e0fa4be83f1 · outbound
Distilled Reinforcement Learning for LLM Post-training Process Reinforcement through Implicit Rewards
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406fe001-ec82-45d6-8d51-a99bcb72e033 · outbound
Distilled Reinforcement Learning for LLM Post-training Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdb8a70-309e-4690-9fbb-1dfee0c572ca · outbound
Distilled Reinforcement Learning for LLM Post-training DeepSeek-V3 Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a927e4a9-c254-4971-943a-61e6a14053fb · outbound
Distilled Reinforcement Learning for LLM Post-training MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b401463d-7b58-4d3d-9f8f-8c43135e25df · outbound
Distilled Reinforcement Learning for LLM Post-training MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b7d37f-60a2-4deb-9aa7-6557f3e84074 · outbound
Distilled Reinforcement Learning for LLM Post-training Training Verifiers to Solve Math Word Problems
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.