Pith. sign in

Paper Citation Record · LEDGER

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2608.09826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09826 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:57:51.075050Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ba62059-8747-4625-a908-af4364bfbe6c · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:50.994951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:50.994951Z digest=sha256:8881197aa5404854489bd64d8ccde5b0f005a10681155a36cc1d80efea9f9168

Observation 4f8ac343-3629-46c1-a33f-6d7448a3298e · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.002061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.002061Z digest=sha256:db8ec2bf5d25dbc2aef18cf40bbcd489a893b7f40040637b303653cdb248bb2e

Observation 04e77260-9a20-408f-8ee9-8fbe1360e585 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.017525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.017525Z digest=sha256:fc1b7f33b3a48e97ef2dae4798c24eb17a708ea9aa96b6379e64bf6cfe14d10e

Observation fe98a91a-b63a-46ce-9cba-d6518db871bd · outbound

This paper cites InInternational Con- ference on Learning Representations, volume 2024, 39578– 39601.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation InInternational Con- ference on Learning Representations, volume 2024, 39578– 39601

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.384107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:57:51.021814Z digest=sha256:fde3eb2a2fc2e406ad23e30ee268668a425a0fd20d824759feac74795c2beed2

Observation 83618215-19e0-45f4-a1b5-4f4b6aa7dfde · outbound

This paper cites Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Nam, T.; Sun, S.-H.; Pertsch, K.; Hwang, S

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.030096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.030096Z digest=sha256:128c8149d1176a436dca27725218c9081a44aaee4142776332db79b29618b6f9

Observation b7a4e63d-38d9-4e0c-9d03-483c4181436f · outbound

This paper cites Skill-based Meta-Reinforcement Learning.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-based Meta-Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.033925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.033925Z digest=sha256:ec7c4afecc0f0a9c0445ddb63e591ff8d3968fc66840f9d54099805b48856f41

Observation 74c6dd90-6804-4140-8cff-c30037ed0391 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.038075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.038075Z digest=sha256:6cc3cb4ffeba85970b99541eb548e8bb44e6f2164fb6c4c4fa4174c1c0c99a6a

Observation 0c257387-727d-4507-a38e-e0064d82d088 · outbound

This paper cites Skill-based Model-based Reinforcement Learning.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-based Model-based Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.042725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.042725Z digest=sha256:3efeb7d25e0fc26292a6b9658d8e70d1087c78f62140e1981d4f555af449e8c1

Observation d355fa74-b0e1-4ff6-a014-c3ee91c4fbb9 · outbound

This paper cites Learning by Distilling Context.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Learning by Distilling Context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.046808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.046808Z digest=sha256:58c5bac7992dceb876f1bde3b39a710ab0b3821e0c038f56de4169ee43292808

Observation c01aace5-e8f3-4764-a66e-af538adb1447 · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.055042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.055042Z digest=sha256:f0034219fb09c979f966bf83f20bda2159129e7061f6ba363fd8b86502e139e5

Observation 1ea6ecc2-2dfb-41ee-8fd8-45873a489136 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.059264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.059264Z digest=sha256:9874fffa7ef49794a182e0e6c977ec138fe99aadeb7d9f8c0e1e8f786036ae6a

Observation 3d608ef9-024e-4b91-9396-5790c6e51917 · outbound

This paper cites A Survey on Knowledge Distillation of Large Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation A Survey on Knowledge Distillation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.063267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.063267Z digest=sha256:c5055ac2f0428e6beaa5d4fe207b88473082e582e53b4c91a7be9bd76c5ae6eb

Observation 91918b64-a48b-4a63-b639-a8d3dccac5e1 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.067414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.067414Z digest=sha256:92900f0c7758a734b1c568b27b62fc54d7ef68bfcbe1ad9c318afee8cfb532f4

Observation d568653b-c3d7-4471-be70-3169ce2aee47 · outbound

This paper cites OPSDL: On-Policy Self-Distillation for Long-Context Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation OPSDL: On-Policy Self-Distillation for Long-Context Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.071097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.071097Z digest=sha256:c632c0288610aaaca16629d9f1de39478b4e961a8117ef5a7fe65ec593312167

Observation 46216a6d-52ce-417d-8b25-0848781b78bc · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.075050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.075050Z digest=sha256:e3e69844ce87364e4708baf4d7bd08a3cb7439051b574c51fcce109bc318a566

Observation 49f09919-1614-428a-9f8e-00ce3e13d919 · outbound

This paper cites Rényidivergenceand Kullback-Leibler divergence.IEEE Transactions on Infor- mation Theory, 60(7): 3797–3820.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Rényidivergenceand Kullback-Leibler divergence.IEEE Transactions on Infor- mation Theory, 60(7): 3797–3820

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.372408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:57:51.051171Z digest=sha256:bef9341a252bf96a9cf57ea7bf8d9b826bc2cece72e691f34b12b2c0cb9efa52

Observation abb5e02e-59b5-4365-9795-b310804e5112 · outbound

This paper cites In-context Reinforcement Learning with Algorithm Distillation.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation In-context Reinforcement Learning with Algorithm Distillation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.008080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.008080Z digest=sha256:67a11f862f57fd7578427bdab84ef9aa23951c9104b1ffe498deaa29eb7ee30d

Observation 39e02caa-754d-4b65-9b75-46df2de1ce23 · outbound

This paper cites Gou,J.;Yu,B.;Maybank,S.J.;andTao,D.2021.Knowledge distillation: A survey.International journal of computer vision, 129(6): 1789–1819.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Gou,J.;Yu,B.;Maybank,S.J.;andTao,D.2021.Knowledge distillation: A survey.International journal of computer vision, 129(6): 1789–1819

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:57:51.395143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:57:50.990619Z digest=sha256:185b7dad9c01118aedca9fc2f248446f77eb5664543f3d454c6e9aae6ea4d10c

Observation 49f2b6e4-87a6-4f3f-ad81-d79f9f09dca1 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.012718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.012718Z digest=sha256:af2a1b072a92c5abf5c3be7306d76909484078d57a3291e946389526b227dd33

Observation 8c4ed339-f714-4e9d-a641-632708282d8a · outbound

This paper cites an unresolved cited work.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:57:51.406340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:57:50.980488Z digest=sha256:b2fa9c147146918652f70557f598aa8e3be090be206e0727fbd50333c1b01b94

Observation 74884d22-95b9-49bc-936c-5d299c6fcb1e · outbound

This paper cites Unifying distillation and privileged information.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Unifying distillation and privileged information

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:51.025787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:51.025787Z digest=sha256:3da5dbc281776a0df23432fbc9ab16b1f2e47079e9b3737cb9ad2d9ee9e95a37

Observation f082d908-3485-4b6e-bd7f-6227454726d8 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T05:57:50.985760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:57:50.985760Z digest=sha256:dd5d8034423a9981eba0e911f72bf8927e5574cd99f2fcb0ac7793a868e8686c

Pith citing papers

No inbound Pith citation observations are available.