Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:53:39.197075Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 14 inbound Pith citation observations for arXiv:2506.09477.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:53:39.197075Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:17.105396Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:40:07.862245Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 47d404ba-0f6a-45dc-9ac5-4b27c77f5616 · outbound
On a few pitfalls in KL divergence gradient estimation for RL OpenAI o1 System Card
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12fc0dd7-de8b-408e-bd3c-efb6ce4cf6b6 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Understanding R1-Zero-Like Training: A Critical Perspective
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbded47c-c979-4d50-b3c1-d4d2a82b2f49 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Equivalence Between Policy Gradients and Soft Q-Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0367503-491f-4ba0-b81e-918e5a46db32 · outbound
On a few pitfalls in KL divergence gradient estimation for RL RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2518048-a2e4-4371-aa52-f76b644542a7 · outbound
On a few pitfalls in KL divergence gradient estimation for RL d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7fb5e4e-5916-4210-8cf7-b16f66c20057 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Details on theoretical results of variance reduced gradient estimate We provide some more details on the theoretical results of the variance reduced gradient estimate
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bf10f9a3-a2d2-429f-a2f3-747c42b92411 · outbound
On a few pitfalls in KL divergence gradient estimation for RL The reward maximization and on-policy distillation experiments are both carried out on the 7500 MATH training prompts (Hendrycks et al., 2021)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5240904b-dc35-4e2b-924c-87670e4cfafd · outbound
On a few pitfalls in KL divergence gradient estimation for RL On the design of kl- regularized policy gradient algorithms for llm reasoning
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc43729-2309-44a1-b104-6dd6dd3a7ff4 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47349bc3-35d0-4808-a62a-78d5df0ee905 · outbound
On a few pitfalls in KL divergence gradient estimation for RL The Llama 3 Herd of Models
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fff336c-f0c6-49eb-ae02-0ed116cc2c0c · outbound
On a few pitfalls in KL divergence gradient estimation for RL Fine-Tuning Language Models from Human Preferences
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1165ba8d-9c0c-44ee-84a0-a7fdd4724392 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Soft Policy Optimization: Online Off-Policy RL for Sequence Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f1f486-ded2-49d0-ad00-e8831d0d6165 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Measuring Mathematical Problem Solving With the MATH Dataset
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b08c8b-4762-4328-8fc2-8a8ff5760855 · outbound
On a few pitfalls in KL divergence gradient estimation for RL PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b92a84-cfb6-44a7-8318-60dc0e22414d · outbound
On a few pitfalls in KL divergence gradient estimation for RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261636fd-99f3-48c9-865a-8eac72226e6d · outbound
On a few pitfalls in KL divergence gradient estimation for RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a893ba-dd81-4779-8f18-8dd96e71abbc · outbound
On a few pitfalls in KL divergence gradient estimation for RL On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d1c640-c154-40cd-b440-65865129c217 · outbound
On a few pitfalls in KL divergence gradient estimation for RL Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fb93e9e2-8fda-416f-af68-8a2d11cd188d · inbound
Outcome-based Exploration for LLM Reasoning On a few pitfalls in KL divergence gradient estimation for RL
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d111ca8-cc59-4718-882c-f33a1c380929 · inbound
Learning to Discover at Test Time On a few pitfalls in KL divergence gradient estimation for RL
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2ddf49ea-7940-4f64-adfa-cd484f4b793b · inbound
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents On a few pitfalls in KL divergence gradient estimation for RL
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f7149b28-9cc6-495b-b573-031bda18baad · inbound
Beyond Distribution Sharpening: The Importance of Task Rewards On a few pitfalls in KL divergence gradient estimation for RL
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8be8f200-8671-4822-9aa1-4abcfff1152f · inbound
KL for a KL: On-Policy Distillation with Control Variate Baseline On a few pitfalls in KL divergence gradient estimation for RL
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 21fb574d-adc8-4751-984e-03958d073072 · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability On a few pitfalls in KL divergence gradient estimation for RL
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6dedc3fd-0a6b-47c4-acca-3d86a6e5e0b1 · inbound
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation On a few pitfalls in KL divergence gradient estimation for RL
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec98e39a-e769-499a-81cf-4147de42a357 · inbound
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation On a few pitfalls in KL divergence gradient estimation for RL
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2041392a-b305-4125-865c-1ab986934b85 · inbound
OPD+: Rethinking the Advantage Design for On-Policy Distillation On a few pitfalls in KL divergence gradient estimation for RL
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8a1953ca-efd8-4d3e-a81f-03c8585f6e50 · inbound
Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions On a few pitfalls in KL divergence gradient estimation for RL
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cf6f60d5-6cde-45e3-8460-3ab7a060b84b · inbound
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents On a few pitfalls in KL divergence gradient estimation for RL
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e19c15c6-40ff-48eb-a0e2-3d112c0596d3 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d4cd12-08db-4a46-9185-101a8416eb73 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d6b363-5e05-47cc-8ecf-0ae3bcc016e7 · inbound
Interpreting Language Model Hidden States at Scale On a few pitfalls in KL divergence gradient estimation for RL
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.