Pith. sign in

Paper Citation Record · LEDGER

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.04788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04788 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:07.147879Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1ad2c2e-7897-4a58-936b-c7abc4d9cffe · outbound

This paper cites On-Policy Delta Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On-Policy Delta Distillation

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.666528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:05.523794Z digest=sha256:7a52b953500f37f34738151ea48a3f3081f2b9cf8b87c1aa6c23a5edcba4bbac

Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · outbound

This paper cites Rethinking On-Policy Self-Distillation for Thinking Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.711475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.711475Z digest=sha256:ee2cb97ddd2602f4d50fbb788725e9a32dd77113053e82e6a978e55d82b08116

Observation c0342944-478a-44d2-bb1f-10c5d83b5640 · outbound

This paper cites What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.901164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.901164Z digest=sha256:ddd2739ebdfdd714b3a02e64cf33d146c92480edf1290ed41c639c01f2e2a119

Observation e7417754-ad18-4cea-89c2-5e37ef6c84f9 · outbound

This paper cites Self-Distilled Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Agentic Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.959982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.959982Z digest=sha256:ea40ab1cf16a46e3d3f67c29bf1eeb5a3288ecd2ed0a0c0db8991b55e29ce8b4

Observation 99fa61dd-c8d5-467b-9b10-929fafded276 · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.084807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.084807Z digest=sha256:0eff7230f66edc7d8548558356b9dffdc94d0ac74f5e29744e650b1e85778c13

Observation 580e6fef-625d-424c-bd32-c7274ac12b97 · outbound

This paper cites Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.431166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.198738Z digest=sha256:07fd44413e49f3babb4cf3a9de3135ec4addc89be9a9a595295e57c348a91230

Observation 59af0db5-903e-4c9c-9a80-4faba569bfca · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.275961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.275961Z digest=sha256:f63c2170383443f2e1fe7bcf050538dbda35a7a9c81bbe924a08cb0a56bdb37c

Observation ce75a16e-460f-4b37-89c4-a00d3130759c · outbound

This paper cites InInternational Conference on Learning Representations, volume 2025, 89490–89520.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation InInternational Conference on Learning Representations, volume 2025, 89490–89520

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:39:09.119131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.491333Z digest=sha256:d7e07d842e90053c76a97b504e3302eded93f24463b72d4123d46749d3ebc9ef

Observation 38de8315-b88a-46a4-82dc-19a2ec9350d9 · outbound

This paper cites PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.247916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.584525Z digest=sha256:1e5bad924606fd5e04d2ff330a515b246718abb3cdfa79e9d66490a24133d02c

Observation a54fa22c-d4a9-4e5c-b091-7fcec56b98e6 · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.054402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.711304Z digest=sha256:86701f039552f76f8b3744c0cc0b4990c23a8bcae930dbd828d5e64b9fca1313

Observation 1b4eb0f6-4658-4a91-9f08-3b6b2855efc0 · outbound

This paper cites On the Position Bias of On-Policy Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation On the Position Bias of On-Policy Distillation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.816552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.783612Z digest=sha256:43cb34cac8bdf11a537b490726614e264b06eced7279fcc431cfa46772509c84

Observation 7533a37c-84f9-42ba-ace3-070ad07a3177 · outbound

This paper cites Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.586658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.851794Z digest=sha256:be0ef1dba44bb9b039c2701be9addc26456aaa4f3e28af2cea0aad02a6dbc7fe

Observation 0bb6da02-090c-412b-b582-0e459ef46997 · outbound

This paper cites Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Beyond Absolute Imitation: Anchored Residual Guidance for Privileged On-Policy Distillation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:07.350446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:06.947603Z digest=sha256:916cbbec84232d657f0d31a61a52e653b8381b0efd54bd0101c376c243bbdf63

Observation 5f0a1dd7-5d11-46a4-a987-f8e4f6439289 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:07.025064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:07.025064Z digest=sha256:8cbc91e700a5f4f61a58045fc792f5aedb481086767b0557b5b284290f0df632

Observation ff4914d2-bbe1-4cef-8ac5-defa24f6e7f9 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:07.147879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:07.147879Z digest=sha256:b2c49abebde0fbb965707c8e84e3e96794dff0a0f61c401f261da9653cb9d76d

Observation 50411b58-ae63-4659-8875-6ab149674276 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:06.391541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:06.391541Z digest=sha256:7f5a4c0451a43e49f9a4487eb7b939ce5fb1497a11daca65a49d93b9191bb537

Observation 70e865c4-6f3f-409b-8835-4cdeb391d57d · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.816250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.816250Z digest=sha256:be97b68f9dce4812f0cddd8d98180d1ca922cf41705c4e2c0e1932f827f62294

Observation a89a410f-ac14-4e8c-a089-12281d578d34 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.647351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.647351Z digest=sha256:5577cc7551f715904bd1733f2f85a3469ed492535b9d2a2175421e01d4063ac4

Observation 51b017bc-e9e1-47fc-b9ac-05e60814ae8a · outbound

This paper cites Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy

Reference 2026

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.874780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:39:05.412903Z digest=sha256:69e5b77c6b78f467a88b0898509bc289362be848339184668ee77f9004defb1c

Pith citing papers

No inbound Pith citation observations are available.