Pith. sign in

Paper Citation Record · LEDGER

On-Policy Self-Distillation without Any Supervision

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.06296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06296 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:43.839213Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9d279ec-09f0-4ddb-9790-a4860180521e · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.326448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.326448Z digest=sha256:7762e7ee78aaedf9e5106f30232a9478029a92612521410abed89d01d118b57c

Observation 1b7637fa-51c4-4807-a2f1-43ef153e9df6 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

On-Policy Self-Distillation without Any Supervision MiniLLM: On-Policy Distillation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.411447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.411447Z digest=sha256:9d4d7298add6fe3bd88604b4ae511b00d90de187137dee0aaa8bcdc9964e896f

Observation f9cb2373-0e40-4d27-8f14-1197802d626e · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

On-Policy Self-Distillation without Any Supervision OpenThoughts: Data Recipes for Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.444364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.444364Z digest=sha256:ca7000f20bf8592b883a62547f8341825e7a627e3ff839cad3d7ba2df425976f

Observation df86c3a3-12bb-4391-b060-2bf05ce0c6dc · outbound

This paper cites Large Language Models Can Self-Improve.

On-Policy Self-Distillation without Any Supervision Large Language Models Can Self-Improve

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.572940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.572940Z digest=sha256:4857028e7c87b5154a32ad33b8d17f28f18df70eeee16252d3744b8fea2c859e

Observation a710c963-c998-47cc-9513-7d4135445f1a · outbound

This paper cites UniSD: Towards a Unified Self-Distillation Framework for Large Language Models.

On-Policy Self-Distillation without Any Supervision UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.728357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.728357Z digest=sha256:900bb35a94bedde023ccdaa5272ea20b0561291561baff9023e8f8f219fd9977

Observation ae2517a6-bafe-45a7-9921-e2a36b8ff6e1 · outbound

This paper cites Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning.

On-Policy Self-Distillation without Any Supervision Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.784586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.784586Z digest=sha256:10de2956d00e1199a2fabf281c760ab52c946ab5525e289eda1b3d62ce80b2f0

Observation f2c04ddc-ed22-4a6b-8056-4cdb81fed65e · outbound

This paper cites Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models.

On-Policy Self-Distillation without Any Supervision Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.827630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.827630Z digest=sha256:f19b546efae726930e80c89934ff9f84bca50fabaeec4d8abe47598353e67442

Observation d1caa6be-f65c-45c2-a1a6-d0a3c4949e81 · outbound

This paper cites Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,.

On-Policy Self-Distillation without Any Supervision Self-evolving visual questioner.arXiv preprint arXiv:2606.13929,

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:19:44.930747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:41.853544Z digest=sha256:21fc47bdf3dd3be56f41f10947be749e1ce650fce4bb8a804f88580e2375604e

Observation cce98735-534e-4d7b-bec6-723f40a9ff5b · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

On-Policy Self-Distillation without Any Supervision Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.914010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.914010Z digest=sha256:8601c130019a606477de339bb332e10c85ddc750758a7a8b2e1700b5128bed80

Observation dbbaa0a1-f975-4e45-b0a9-8fad67c2f673 · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

On-Policy Self-Distillation without Any Supervision HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:44.623767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:41.961817Z digest=sha256:d8bfe58759a037043b4ed22607abc306fed5526158084b7e04759a497ce470a9

Observation 5409db2d-a287-478d-9367-4c9e5523aa91 · outbound

This paper cites URL https://thinkingmachines.ai/ blog/on-policy-distillation.

On-Policy Self-Distillation without Any Supervision URL https://thinkingmachines.ai/ blog/on-policy-distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.025433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.025433Z digest=sha256:45181be5d5d2d34321cc18c889e5e5fa2c5bffdcf89ff7dd856c46ae1e2da481

Observation 5ddeec39-3bf7-488d-a025-b484ce551768 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

On-Policy Self-Distillation without Any Supervision MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.100814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.100814Z digest=sha256:32e346168f81917c424ac30beb5d8d2c26280aee206a0661666ba0665b89dec9

Observation fdd5552c-0f46-4046-b580-ea0c1db55611 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

On-Policy Self-Distillation without Any Supervision Maximizing Confidence Alone Improves Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.170600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.170600Z digest=sha256:05c40ebc5525c0b6486558bc7641d91df87490fa6dca742667a15a43ec45d2f4

Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · outbound

This paper cites Self-Consistency Preference Optimization.

On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.234006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.234006Z digest=sha256:08a55ed3ad79191b2eeb17b4d020148da94675ef3a39ba696525530b3cc8f920

Observation 7b608f62-0f24-4183-bf92-ecf1227a2398 · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

On-Policy Self-Distillation without Any Supervision CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.294503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.294503Z digest=sha256:3a7d0841d38138c1dd2b55876be299b36393983f192488113b8ab5e659e97cc8

Observation 21f116ed-e599-4564-8d34-7d621f60020a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On-Policy Self-Distillation without Any Supervision DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.384243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.384243Z digest=sha256:9bd8a72aad324f2b48a7fa133bbd7afe179383acfa6ed5525b97c931535967bf

Observation a9f0db98-ed67-46b7-8b0d-98bb297156c5 · outbound

This paper cites Self-Distillation Enables Continual Learning.

On-Policy Self-Distillation without Any Supervision Self-Distillation Enables Continual Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.449663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.449663Z digest=sha256:e626148d206defd7d7b4b035e201c3bef2c2fb8f859d49a917242b189d4132b6

Observation d2ee794e-11d1-41ac-a219-3fda760934a1 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

On-Policy Self-Distillation without Any Supervision GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.513853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.513853Z digest=sha256:5ecbf4ab63a5e3f695985b3063952b53408aa7a4e8e7dbd1bf40fae19d646cf1

Observation 2154647d-37e0-44a8-9cd8-93e4ab1e3cd2 · outbound

This paper cites Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations.

On-Policy Self-Distillation without Any Supervision Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:19:44.047486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:42.587810Z digest=sha256:5547592910dc43abe920ef2a39bba7047898e5217546819956fe544f073507ad

Observation 84d0f8aa-9ec0-4549-9c80-977a7ede8044 · outbound

This paper cites SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces.

On-Policy Self-Distillation without Any Supervision SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.700966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.700966Z digest=sha256:2ef1101c0c307e3e082ece70ec8c2019340f151e651374b72b9c8b20e69dedfc

Observation 97b0ba77-58c6-42f0-9eea-bfb789c67920 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.801108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.801108Z digest=sha256:3befa69027674a66b449b0e658babfcd62f47d789bd80414c49cb41639116799

Observation e5f4537e-e154-4c97-a652-cdf41624b582 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

On-Policy Self-Distillation without Any Supervision Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.927212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.927212Z digest=sha256:5afee40770dc88f0643a9ec348b605f759a0e5b6771f366212cc2db88caf4b19

Observation 2c84e049-02db-4bd7-b671-bc8815b22c4f · outbound

This paper cites Qwen3 Technical Report.

On-Policy Self-Distillation without Any Supervision Qwen3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.007032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.007032Z digest=sha256:97aa0702489f4991cdf3519eabb0be4d74e2cf9a03012135c29833385343c347

Observation 87977f5a-2953-48d6-aa83-22770c97d75b · outbound

This paper cites Snapshot Distillation: Teacher-Student Optimization in One Generation.

On-Policy Self-Distillation without Any Supervision Snapshot Distillation: Teacher-Student Optimization in One Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:44.305382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:43.092889Z digest=sha256:712d97d3e344dcfd997098586c5b2e6a99d54185bbfd1e26a34af1d3eb27cfc2

Observation cbaaed0d-6c3f-47f9-810d-868cc4bead47 · outbound

This paper cites On-Policy Context Distillation for Language Models.

On-Policy Self-Distillation without Any Supervision On-Policy Context Distillation for Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.250443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.250443Z digest=sha256:105f30d382e70c5fcb08396f028612a7428fcd3197032baa2341e88342170470

Observation bd14b0e8-7f9d-484b-841f-0b74a17ba5de · outbound

This paper cites Self-Rewarding Language Models.

On-Policy Self-Distillation without Any Supervision Self-Rewarding Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.347186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.347186Z digest=sha256:532861c7ef534fd776b679713960c04276c21e7c7b09f5d1ea0873c2fd6fc8e1

Observation e37b18b6-0199-485d-863f-ee6df588524a · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning.

On-Policy Self-Distillation without Any Supervision STaR: Bootstrapping Reasoning With Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.516621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.516621Z digest=sha256:2b753ed5e7461eb6334ef13b9562efb1bfc3823d3469a8fea5ecc36a46d344b3

Observation 5f69390f-b506-483a-9d7d-9f611a7600fd · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

On-Policy Self-Distillation without Any Supervision Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.698887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.698887Z digest=sha256:ca53f7b9967b8eeed4743b6f566c02d6794393c826e3ef0ea22c43fd0107e195

Observation 08631336-b861-4564-ab57-28cf45aba4d4 · outbound

This paper cites Learning to Reason without External Rewards.

On-Policy Self-Distillation without Any Supervision Learning to Reason without External Rewards

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.764127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.764127Z digest=sha256:1fc3cfeb64edf12c186c068bdd1c8c51c95363c2f550f267eb461ee1a13e0369

Observation 3acd0707-5d9d-425b-82d2-f8b27eb2a3fd · outbound

This paper cites pub.” denotes the numbers published in the official OPSD repository; “ours.

On-Policy Self-Distillation without Any Supervision pub.” denotes the numbers published in the official OPSD repository; “ours

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:19:45.236912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:43.839213Z digest=sha256:63bcf60cb5f3a60092834b8be2483931d674b3b6b666bd339d8286c035b61203

Observation ceacb47f-efeb-4c0c-8e6a-932ff0750eef · outbound

This paper cites Self-Distilled RLVR.

On-Policy Self-Distillation without Any Supervision Self-Distilled RLVR

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.157443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.157443Z digest=sha256:b7f8e92e6f1e86a416cc5f1199ca2606dc3533eb400f680802136b683cc070ff

Observation 9cfbb828-d21e-44c6-8a5b-d32ece739d6f · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

On-Policy Self-Distillation without Any Supervision Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.621441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.621441Z digest=sha256:c1a5e2348b83865a0e81a05381baa06641daac9c3a17454c3cfa3a1309bbc470

Observation 125cc2a1-4970-4931-8111-eb1d3c01d637 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

On-Policy Self-Distillation without Any Supervision R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.490605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.490605Z digest=sha256:4117a3683fdc2aadb74ffe912209438d1f8ac6d5e8891408d239063f91ce4ca6

Observation 8b6d3f6b-7b5d-4a83-afb4-79b6cf5f0262 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

On-Policy Self-Distillation without Any Supervision Reinforcement Learning via Self-Distillation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.667277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.667277Z digest=sha256:41cbf51c053fc5c921da88be383d70d2bb3240a409430269ae1c7c816eaaeb6e

Observation 2a38042b-66e7-4535-a4a0-aa0bd7eafc88 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

On-Policy Self-Distillation without Any Supervision Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:43.431632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:43.431632Z digest=sha256:c12b72df7e0a60ab176b740ad7f9fe4087c202ad24a33b0021b6e4f4bfa7817c

Observation 49934ccb-49e9-4533-8497-46dfbd014348 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On-Policy Self-Distillation without Any Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.249543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.249543Z digest=sha256:59c32765581f6ea715d67f710d766df65eb3739e23d734813a85a41b6a78ddf3

Observation 2a5413ea-f2e3-4e93-a015-1e7d2c3215b6 · outbound

This paper cites Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,.

On-Policy Self-Distillation without Any Supervision Serl: Self-play reinforcement learning for large language models with limited data.arXiv preprint arXiv:2505.20347,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.277301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.277301Z digest=sha256:f77e713a98ab982e34e9bf500ab70d10e104c08564e022923db4ce4a9f5bddbe

Observation 2faf3dfb-1cba-4f8b-8eeb-6600154c7a29 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

On-Policy Self-Distillation without Any Supervision Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:41.381899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:41.381899Z digest=sha256:352df4e523784b6f4fa6c477f3794b471cfbb6769199a19adb5728f22411ebcd

Pith citing papers

No inbound Pith citation observations are available.