Pith. sign in

Paper Citation Record · LEDGER

On a few pitfalls in KL divergence gradient estimation for RL

As of 17 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 14 inbound Pith citation observations for arXiv:2506.09477.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09477 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:53:39.197075Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:17.105396Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:40:07.862245Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47d404ba-0f6a-45dc-9ac5-4b27c77f5616 · outbound

This paper cites OpenAI o1 System Card.

On a few pitfalls in KL divergence gradient estimation for RL OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.063638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.063638Z digest=sha256:e8dbbdc54629102184244db1b1daadd6656c29cedcb8d22379bcafee35c445e4

Observation 12fc0dd7-de8b-408e-bd3c-efb6ce4cf6b6 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On a few pitfalls in KL divergence gradient estimation for RL Understanding R1-Zero-Like Training: A Critical Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.168837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.168837Z digest=sha256:cc7196829320ac4de0d46343e5126ef04e8413ee522c810c2526ed77329816fc

Observation fbded47c-c979-4d50-b3c1-d4d2a82b2f49 · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

On a few pitfalls in KL divergence gradient estimation for RL Equivalence Between Policy Gradients and Soft Q-Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.405144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.405144Z digest=sha256:c2c79820488bacdd5993180b5d70c561ea7bc5ba5dac1e959fc94ad0a440b9b8

Observation f0367503-491f-4ba0-b81e-918e5a46db32 · outbound

This paper cites RL-finetuning LLMs from on- and off-policy data with a single algorithm.

On a few pitfalls in KL divergence gradient estimation for RL RL-finetuning LLMs from on- and off-policy data with a single algorithm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.658756Z digest=sha256:dcb7c1bfb09e1564fb2e68e40d6308228860bbcbb6814b29ad93d612bfee9a03

Observation b2518048-a2e4-4371-aa52-f76b644542a7 · outbound

This paper cites d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning.

On a few pitfalls in KL divergence gradient estimation for RL d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.908197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.908197Z digest=sha256:edd5feb69c5fc5d73c078956733bede9c20beaa101137ecc866415244ec80c8e

Observation d7fb5e4e-5916-4210-8cf7-b16f66c20057 · outbound

This paper cites Details on theoretical results of variance reduced gradient estimate We provide some more details on the theoretical results of the variance reduced gradient estimate.

On a few pitfalls in KL divergence gradient estimation for RL Details on theoretical results of variance reduced gradient estimate We provide some more details on the theoretical results of the variance reduced gradient estimate

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:53:40.066058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:53:39.096962Z digest=sha256:6cc4408f14783f34c263ffc6d98f7b2d002c66fadbde904988799432ac0859ab

Observation bf10f9a3-a2d2-429f-a2f3-747c42b92411 · outbound

This paper cites The reward maximization and on-policy distillation experiments are both carried out on the 7500 MATH training prompts (Hendrycks et al., 2021).

On a few pitfalls in KL divergence gradient estimation for RL The reward maximization and on-policy distillation experiments are both carried out on the 7500 MATH training prompts (Hendrycks et al., 2021)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:53:39.906857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:53:39.197075Z digest=sha256:94ee3667c778f505691bc70e68c14fc7eb6ac321a91056a4b7a9b7be17a308c7

Observation 5240904b-dc35-4e2b-924c-87670e4cfafd · outbound

This paper cites On the design of kl- regularized policy gradient algorithms for llm reasoning.

On a few pitfalls in KL divergence gradient estimation for RL On the design of kl- regularized policy gradient algorithms for llm reasoning

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.781775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.781775Z digest=sha256:400d203dd7bfd340828ffd1f81f414dc045fff255ab9b68972fb098ae0146952

Observation 9fc43729-2309-44a1-b104-6dd6dd3a7ff4 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

On a few pitfalls in KL divergence gradient estimation for RL Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.578236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.578236Z digest=sha256:f4ebcc7e6e3aeeae50343884b2d85eea1609830c74574550a1a6c992522ee41d

Observation 47349bc3-35d0-4808-a62a-78d5df0ee905 · outbound

This paper cites The Llama 3 Herd of Models.

On a few pitfalls in KL divergence gradient estimation for RL The Llama 3 Herd of Models

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.534036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.534036Z digest=sha256:0c199f6ec76063e45c982a0159c5ca25a348367d6c61b559060e106d2c64766f

Observation 9fff336c-f0c6-49eb-ae02-0ed116cc2c0c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

On a few pitfalls in KL divergence gradient estimation for RL Fine-Tuning Language Models from Human Preferences

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:39.012582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:39.012582Z digest=sha256:e94ce59af2ad86f690eb08fcbe9e91086d5f6da33a83cacb135c0bb5fa33cfa8

Observation 1165ba8d-9c0c-44ee-84a0-a7fdd4724392 · outbound

This paper cites Soft Policy Optimization: Online Off-Policy RL for Sequence Models.

On a few pitfalls in KL divergence gradient estimation for RL Soft Policy Optimization: Online Off-Policy RL for Sequence Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.298177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.298177Z digest=sha256:d21502372477e4ba910ff3a61b1c7bf19799d4d7892a0e5063b51f069edcafb2

Observation 08f1f486-ded2-49d0-ad00-e8831d0d6165 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

On a few pitfalls in KL divergence gradient estimation for RL Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.788104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.788104Z digest=sha256:b82af44753233abbb799db538e2a43b31ae27c1ebc006044279b8d652e7704e0

Observation 19b08c8b-4762-4328-8fc2-8a8ff5760855 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

On a few pitfalls in KL divergence gradient estimation for RL PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.262205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.262205Z digest=sha256:912232256a13359b3c112782f57394d247c368a4ed77b2ed823f1dc153c131df

Observation 65b92a84-cfb6-44a7-8318-60dc0e22414d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On a few pitfalls in KL divergence gradient estimation for RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.704030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.704030Z digest=sha256:c6cb06eab12aad7710af5f09ebbf03f90a08808b31cd8f62c0b40cd9d807e3f0

Observation 261636fd-99f3-48c9-865a-8eac72226e6d · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

On a few pitfalls in KL divergence gradient estimation for RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.946422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.946422Z digest=sha256:d4673eae52fd813f502492a2103b4f338dbc7165cb9a702487e32e86f2a1b7aa

Observation c7a893ba-dd81-4779-8f18-8dd96e71abbc · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

On a few pitfalls in KL divergence gradient estimation for RL On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:37.197610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:37.197610Z digest=sha256:03922354a35bac6429b71cfacde116c5ff27a8a81ab02fca813fa749e132841e

Observation 64d1c640-c154-40cd-b440-65865129c217 · outbound

This paper cites Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion.

On a few pitfalls in KL divergence gradient estimation for RL Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:53:39.648865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T04:53:37.424674Z digest=sha256:7f7869b674e9bc12a6d535f9f6d95f3621f8b4ec6a5b771ed67e7dca75a41ebc

Pith citing papers

Observation fb93e9e2-8fda-416f-af68-8a2d11cd188d · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning On a few pitfalls in KL divergence gradient estimation for RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.560552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.560552Z digest=sha256:061b8b47d64eac2fea88940aa8471d485b0a793c9eb37c721518a1a8d7a89f7d

Observation 6d111ca8-cc59-4718-882c-f33a1c380929 · inbound

Learning to Discover at Test Time cites this paper.

Learning to Discover at Test Time On a few pitfalls in KL divergence gradient estimation for RL

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:16:04.236306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T05:16:04.001700Z digest=sha256:dfcc2b239af82582693c1f51d2ff732374ff9884fcf170354d112784dd6017dd

Observation 2ddf49ea-7940-4f64-adfa-cd484f4b793b · inbound

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents cites this paper.

Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents On a few pitfalls in KL divergence gradient estimation for RL

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:00.921568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T15:28:07.981488Z digest=sha256:71191fe010753d9c2e5f9767f36b046aa468677113ccfa23f8557583f8f92e3f

Observation f7149b28-9cc6-495b-b573-031bda18baad · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards On a few pitfalls in KL divergence gradient estimation for RL

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:26.980834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:4a8cd1b2e92ae890a0f1c1e9b7312b16fa9130f4dd899f18302aa80f3fd29d53

Observation 8be8f200-8671-4822-9aa1-4abcfff1152f · inbound

KL for a KL: On-Policy Distillation with Control Variate Baseline cites this paper.

KL for a KL: On-Policy Distillation with Control Variate Baseline On a few pitfalls in KL divergence gradient estimation for RL

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:54.210286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T03:31:07.462474Z digest=sha256:0374a7b22c81a0b514f26eb8186e8db275e5ba1dfe3a9d610d53df28e957f930

Observation 21fb574d-adc8-4751-984e-03958d073072 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability On a few pitfalls in KL divergence gradient estimation for RL

Reference 111

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:30.653037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:49efb047e31b7fdb65529bbb0508631e39deb408286691e3125a1e471b960254

Observation 6dedc3fd-0a6b-47c4-acca-3d86a6e5e0b1 · inbound

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation cites this paper.

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation On a few pitfalls in KL divergence gradient estimation for RL

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:31:23.051850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T09:29:05.360018Z digest=sha256:80c76549e40e416113c6fa61b1bde7a2ea3b59735f6c4a273986f2005e441757

Observation ec98e39a-e769-499a-81cf-4147de42a357 · inbound

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation cites this paper.

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation On a few pitfalls in KL divergence gradient estimation for RL

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.070676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T05:38:33.676797Z digest=sha256:4816956a1e4b405d5d1ed757b7fbe25d825077361d4f7c829a936e6d50f8ad69

Observation 2041392a-b305-4125-865c-1ab986934b85 · inbound

OPD+: Rethinking the Advantage Design for On-Policy Distillation cites this paper.

OPD+: Rethinking the Advantage Design for On-Policy Distillation On a few pitfalls in KL divergence gradient estimation for RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.569504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:24:21.906134Z digest=sha256:f0cc38e96dcec3e409b397c85e97f7e369c335f3db24ee78cc51bb98eea65db6

Observation 8a1953ca-efd8-4d3e-a81f-03c8585f6e50 · inbound

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions cites this paper.

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions On a few pitfalls in KL divergence gradient estimation for RL

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.596091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T10:51:30.546134Z digest=sha256:43ee267a1fc8fa23ea65d07f68dd73d5140a13b6cd0e9d978dc8f61f24a0c70f

Observation cf6f60d5-6cde-45e3-8460-3ab7a060b84b · inbound

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents cites this paper.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents On a few pitfalls in KL divergence gradient estimation for RL

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:40:07.863689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-25T19:55:51.114244Z digest=sha256:8a1ffe0fbc891ce60369c48c9b7831dc346ea1071ab60acae422105db17c3d50

Observation e19c15c6-40ff-48eb-a0e2-3d112c0596d3 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T11:23:30.063231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:23:30.063231Z digest=sha256:d2902b8ff246b1a29423e0300ce19ce355d5db07b86bdd11c858f4a78d9826cf

Observation 52d4cd12-08db-4a46-9185-101a8416eb73 · inbound

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples cites this paper.

ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples On a few pitfalls in KL divergence gradient estimation for RL

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T07:23:27.338911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:23:27.338911Z digest=sha256:19bbb1df276f02858b285516e4034cf9c686048f392add0feb8837a26c05939e

Observation f0d6b363-5e05-47cc-8ecf-0ae3bcc016e7 · inbound

Interpreting Language Model Hidden States at Scale cites this paper.

Interpreting Language Model Hidden States at Scale On a few pitfalls in KL divergence gradient estimation for RL

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:17.105396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:16:17.105396Z digest=sha256:03c40adab421c1284a16c3a92fc87a0dc782b358d013d43ac23ada503b2465ae