Pith. sign in

Paper Citation Record · LEDGER

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.01837.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01837 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:59:00.345669Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b017c0d-f03d-4ccb-8932-104c8caa8a47 · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.219331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.219331Z digest=sha256:4a3390a1506384f26b0da9eb673fd5320317781d5ba66a83b95fe40a3b1f302d

Observation c9af2218-9e07-49e4-892c-592fd98de37b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.227017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.227017Z digest=sha256:96ed6c1a72c4e6e69a4ee8c30d904e6cf4debb0eb1cdab6ba879f88ed8f49647

Observation 7cb6a7a2-acd4-4b33-ba4b-13c49ecc31be · outbound

This paper cites Jin, B.; Zeng, H.; Yue, Z.; Yoon, J.; Arik, S.; Wang, D.; Zamani, H.; and Han, J.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Jin, B.; Zeng, H.; Yue, Z.; Yoon, J.; Arik, S.; Wang, D.; Zamani, H.; and Han, J

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.239132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.239132Z digest=sha256:334425129570edc508cef1d47375d89fe6fbea9cd451ccbfdb747a020fc1a775

Observation c151a5c0-e625-44b6-983b-3893a5ec2401 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.244103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.244103Z digest=sha256:abdd690549e5f040a0ba2626b5bcdfda15eed9bc7f5a8351a88dd48c9fb5d802

Observation 9ad32867-622b-4d20-b0d3-4c43daf4cd8a · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.250021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.250021Z digest=sha256:83eb353dc68468214d6a135c330c14701328bd55d9d0bd541d5dd0ace21d6110

Observation 87a09415-9b73-4196-9249-06d9fd5abf17 · outbound

This paper cites PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.257672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.257672Z digest=sha256:c4f8d1a9591c0a3df8bced98f5c399f95f71e0cfa3cb7212c499756b623f5c67

Observation 4320ecfa-9e30-413f-87b0-d8c6f7065087 · outbound

This paper cites InInternational Conference on Learning Representations, volume 2025, 79791–79821.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning InInternational Conference on Learning Representations, volume 2025, 79791–79821

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.262943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.262943Z digest=sha256:87c386924afda99b965db1ae01a3824b66411c40d8cddb1d2a19d681028c57ed

Observation 360a4500-963f-4a8b-8e26-ab48d2e3b0d2 · outbound

This paper cites InInternational Conference on Learning Representations, volume 2025, 406–441.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning InInternational Conference on Learning Representations, volume 2025, 406–441

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.269190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.269190Z digest=sha256:8629c1e0dd9d5ec1d13908eb87c7a8566ad88cd9e18a28a8de82bc1c6e638abd

Observation eb569ea2-68d2-473b-9230-e9ff688e6a94 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.281220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.281220Z digest=sha256:8b4b35c6e508f3750e4d97120849f69005cc5b4863b78e46197f1c7f7af78db3

Observation f5e41371-9c31-437b-8c28-1f4276d127cf · outbound

This paper cites OpenAI GPT-5 System Card.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning OpenAI GPT-5 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.295295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.295295Z digest=sha256:252207bef96a5c162be13cbb0a6831da26b286fd0d2729accf4e81415e002c66

Observation f0ae667c-bd3b-4388-85ab-f70d5e94d9c3 · outbound

This paper cites Wu, J.; Yang, S.; Lu, Z.; Zhang, F.; Shen, Y.; Feng, L.; Luo, H.; Lian, Z.; Zhang, S.; Wen, Z.; et al.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Wu, J.; Yang, S.; Lu, Z.; Zhang, F.; Shen, Y.; Feng, L.; Luo, H.; Lian, Z.; Zhang, S.; Wen, Z.; et al

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.302769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.302769Z digest=sha256:7b02837d202fab986a10177e2743844fda03763f2927ae6bf4a4cf2246bb7f35

Observation ae6bd321-e27e-454b-8cf6-342a6b55124e · outbound

This paper cites SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.309548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.309548Z digest=sha256:eb793d3be6fba92538932e3e0c66c7eb5bc1bffd8cd63503231737a6558de974

Observation 0dd09122-7844-423c-aa64-af12f6a62a2b · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.315255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.315255Z digest=sha256:9eb794786a86752c4e29bf810aa6c2b7be0fe420ba7c8ed721bff42d0c0de2d4

Observation e9a70feb-d516-40b4-a4bc-04fc0db54f72 · outbound

This paper cites On-Policy Context Distillation for Language Models.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning On-Policy Context Distillation for Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.321172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.321172Z digest=sha256:d0ce84654eca5806956cd838033a7dd928d3d2ee3768c88a511c40d2fc247b78

Observation cce752fe-1a0a-4975-9c80-710fe81be3db · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.326323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.326323Z digest=sha256:8627bcd7dd6e790c89a214cd94b01fa59c8b9beb112cc57be4f5d57e7c8a4f28

Observation 05c7914f-cb14-4254-a500-48cadc48c2bf · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.333551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.333551Z digest=sha256:ed3f1ceee5363d0b59ec095dc86da7f986a684a10d725874aa6f0bf51602bf09

Observation fc182e4b-37fb-4bd3-946e-c13e12c1800f · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.340360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.340360Z digest=sha256:17a66e66a5f9209573c174ea79252d778d0fa2487285045e09f7bd6ac4ea09a9

Observation b79249ad-8e10-426f-b83a-aaa012a91ea5 · outbound

This paper cites SOD: Step-wise On-policy Distillation for Small Language Model Agents.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning SOD: Step-wise On-policy Distillation for Small Language Model Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.345669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.345669Z digest=sha256:6e7589e8e2ab6d3488f2aca298d072d9cb64ee091975b542b6d096401295f1a2

Observation 92065c53-55e6-4066-8cbb-b3939c8ddb52 · outbound

This paper cites Proximal Policy Optimization Algorithms.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.275824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.275824Z digest=sha256:176617be49493c3a9d9df9d263b5cc888f17ab19da4049eedd60234063f44b8e

Observation 3041e6d0-9ab0-4bea-8acd-83e95392c18d · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.287378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.287378Z digest=sha256:6386ebcca683c6d2dfabd94306b88454c7130d44e79576ca31bb71c97aa56fc2

Observation 509ab239-d545-477e-8445-3c7f3fd7cc1a · outbound

This paper cites Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.205690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.205690Z digest=sha256:eb104016bf4a4651e1bfbc22c39340815d110590434d2964ab15260639663663

Observation 17be89a9-46cb-4218-8e82-85cde710ba52 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.212946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.212946Z digest=sha256:832b96d9b46192c22d288c689c546d6b35cc945e07940c2bff127bd75cf27f3f

Observation 702c1795-5fd7-4e5e-b1b0-0791da9e6b4c · outbound

This paper cites Reinforcement Learning via Self-Distillation.

PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Reinforcement Learning via Self-Distillation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T19:59:00.232410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:59:00.232410Z digest=sha256:43f7f57e9a032b661b63fae35ddcda36d1786b26c3cbf606438750146317865d

Pith citing papers

No inbound Pith citation observations are available.