Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:59:00.345669Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.01837.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T19:59:00.345669Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b017c0d-f03d-4ccb-8932-104c8caa8a47 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9af2218-9e07-49e4-892c-592fd98de37b · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb6a7a2-acd4-4b33-ba4b-13c49ecc31be · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Jin, B.; Zeng, H.; Yue, Z.; Yoon, J.; Arik, S.; Wang, D.; Zamani, H.; and Han, J
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c151a5c0-e625-44b6-983b-3893a5ec2401 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad32867-622b-4d20-b0d3-4c43daf4cd8a · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a09415-9b73-4196-9249-06d9fd5abf17 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning PhysAgent: Automating Physics-Based 4D Synthesis via Trajectory-Grounded Multi-Agent Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4320ecfa-9e30-413f-87b0-d8c6f7065087 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning InInternational Conference on Learning Representations, volume 2025, 79791–79821
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360a4500-963f-4a8b-8e26-ab48d2e3b0d2 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning InInternational Conference on Learning Representations, volume 2025, 406–441
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb569ea2-68d2-473b-9230-e9ff688e6a94 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5e41371-9c31-437b-8c28-1f4276d127cf · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning OpenAI GPT-5 System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ae667c-bd3b-4388-85ab-f70d5e94d9c3 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Wu, J.; Yang, S.; Lu, Z.; Zhang, F.; Shen, Y.; Feng, L.; Luo, H.; Lian, Z.; Zhang, S.; Wen, Z.; et al
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6bd321-e27e-454b-8cf6-342a6b55124e · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd09122-7844-423c-aa64-af12f6a62a2b · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a70feb-d516-40b4-a4bc-04fc0db54f72 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning On-Policy Context Distillation for Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce752fe-1a0a-4975-9c80-710fe81be3db · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c7914f-cb14-4254-a500-48cadc48c2bf · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc182e4b-37fb-4bd3-946e-c13e12c1800f · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79249ad-8e10-426f-b83a-aaa012a91ea5 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning SOD: Step-wise On-policy Distillation for Small Language Model Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92065c53-55e6-4066-8cbb-b3939c8ddb52 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3041e6d0-9ab0-4bea-8acd-83e95392c18d · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509ab239-d545-477e-8445-3c7f3fd7cc1a · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17be89a9-46cb-4218-8e82-85cde710ba52 · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 702c1795-5fd7-4e5e-b1b0-0791da9e6b4c · outbound
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Reinforcement Learning via Self-Distillation
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.