Pith. sign in

Paper Citation Record · LEDGER

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2608.06243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06243 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:38:13.296565Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4219dd4-1dc8-48ae-a32d-8e59cf97dc34 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.106699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.106699Z digest=sha256:0d96c4e7e324b18b74505d7aae09752486bd93fed5c46cbd6dcd7f331af2ccb7

Observation 726d5435-1ecd-4f66-8769-a2242d3f5264 · outbound

This paper cites 2025 , doi =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , doi =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.112655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.112655Z digest=sha256:9d346e18ffaa2f7f74eaa463c9742fb6c1fc66b9f71e754de201a0d14145f134

Observation eda955a7-ecd4-4746-9535-02c278dd2022 · outbound

This paper cites International Conference on Learning Representations , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models International Conference on Learning Representations , year =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.116032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.116032Z digest=sha256:692378293db023e4d82090604f8e0de0f48a3416842d5846fe472478a857cdd6

Observation 35218719-cf3c-48af-84a0-7bf1d25e9e91 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.119605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.119605Z digest=sha256:f7693d412a49fb4cdbb075ad3a8c7eb1295fcf3b62b5860e403acd56f7736dcd

Observation 77be2d5b-91d3-4f50-af29-38a046dded11 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.123104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.123104Z digest=sha256:e57562f6db0bd341b58b401e0b29e86544bfe8e4bb7eac4281fe7f074067d91e

Observation 4aa4d116-78de-46fd-8b3b-50f1884c3425 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.126568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.126568Z digest=sha256:b832f691a55ff5562ba71d52b86f1b66b667a60eb5a318e1df5ca65d17407a40

Observation a3d96e27-24ea-4989-a382-7cebde7f76b4 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.130376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.130376Z digest=sha256:e5e911ff9dd66f91db13d18266f46c740068a6e0f55d701c14e8d3b65602df82

Observation 3496a0ea-259d-4abf-8081-293c11409170 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Entropy-Aware On-Policy Distillation of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.133996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.133996Z digest=sha256:01c0cde998d4d2f3250da28136d92e56ebcd1b6c5189093dacc4d09fdcd9f0c0

Observation 658e2172-3811-4392-90c3-fc95df007c96 · outbound

This paper cites When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.138616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.138616Z digest=sha256:8b0f1e4202e65edd89d6ca24951f6c23e5cd59febb3e6bf8c0d390af675610fe

Observation b1d36558-4772-43ab-a2c1-95bc733479a7 · outbound

This paper cites On the Position Bias of On-Policy Distillation.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models On the Position Bias of On-Policy Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.143580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.143580Z digest=sha256:19b752584531634c936d43e33fcd9e630ea3a662441f7d414646e52ac6d6bf41

Observation 65004d52-a603-482a-8250-49eb9f656b22 · outbound

This paper cites Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.148001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.148001Z digest=sha256:a6f3337d87772bf41f763895fc0dda9d017d15dd02477583e81f8fdde40f3615

Observation b8f8facd-2fb7-4aa1-9442-792adc2bb604 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.152075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.152075Z digest=sha256:725cb25cdcfe1c3a62dd2539609dd73bf1d9ac06383a685949613acaa455f5c3

Observation 1486c8b0-c60f-423b-b72d-96011c922606 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.155967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.155967Z digest=sha256:9d5dc614e0cf94c340fa31f2015c285a4e0b594617317eaef703b392350b1e79

Observation 707e60fc-85eb-40d6-8cf2-3be294fd1c8c · outbound

This paper cites 2026 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2026 , eprint =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.159977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.159977Z digest=sha256:15043025fc7d31dc912d91afaf57b52d3efd45bedb33d116d616378a6f433c26

Observation 18d30972-ef0d-46dc-986c-1014289d4d7f · outbound

This paper cites Purified.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Purified

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.164056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.164056Z digest=sha256:f0154bbe96e2ac49dd60259d5fa5c5c8e9bcf20c8c489459aefb3d4d0e7e7c95

Observation 799d723b-8736-4285-a3a5-546de5330d24 · outbound

This paper cites and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.167823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.167823Z digest=sha256:2823f109975a2324babe69377891916efa26e1cd19942369af515356755ba6fd

Observation e5570b07-990d-4a48-85d0-6965be7c2fb2 · outbound

This paper cites 2026 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2026 , eprint =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.171465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.171465Z digest=sha256:cd8cd6e31a1c067d0fa2f6b1b20ca39ebc6f79fe4e81bdb6ebffaf670aee694d

Observation a463fb26-6db5-4f4b-8b61-2a110149adc1 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 18

Resolution
parse uncertain
no resolver link, observed 2026-08-15T14:38:13.175594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.175594Z digest=sha256:cf857a321b6a04f530b0f905d2de00434b1e643dec0f1b7150b3087b895cff9d

Observation d8e611ed-ac46-4b1d-a7b4-d2322e8cd760 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 19

Resolution
parse uncertain
no resolver link, observed 2026-08-15T14:38:13.179253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.179253Z digest=sha256:5cb13bb2ee714995de23b07d9a362c4376b391926b47830b226755827e8303fb

Observation 1673e8b1-bf0e-469c-9678-1629ef476dbe · outbound

This paper cites and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.182936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.182936Z digest=sha256:e7580548d5a59030f02900b3b02c8852f107b6da09899134bd77786af4f2feca

Observation 70769e61-e37a-41f5-be4c-86b17454c191 · outbound

This paper cites 2025 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2025 , eprint =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.186410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.186410Z digest=sha256:a32565f952753fe97d9d4814ab6fa72e77039922998de9275b0d89923519cce7

Observation c0d0ca39-527d-4c4f-b579-af30d59a96cf · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.189982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.189982Z digest=sha256:20b743b893cd0ced72d6c27333b421dbea6c5bb7ec5a441f6450711aeb9e3efb

Observation e021e681-8e98-4066-a4bd-602d2adb5631 · outbound

This paper cites Proximal Policy Optimization Algorithms.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.193577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.193577Z digest=sha256:afb5997832fd4765b8f657549b008d084b3402b6c0d6ec76397c0c30b385fb4a

Observation 35ef7025-daf5-410b-9df5-b12d96148eb5 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.197421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.197421Z digest=sha256:a664a5e046ec8cf842ca03b3b6d5a48edb01c4fee5151da57bb1b158da83e55a

Observation 8c4fa34c-b90c-4ec8-8dda-a5b5ec07bbf2 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.201026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.201026Z digest=sha256:07626502b97940ce01499b203ba5a97041e30216515610739c43cbbf1d398fe7

Observation 38eb1fa2-53bb-40ea-b757-be4a789de237 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.205106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.205106Z digest=sha256:f5d614c0784e4146ea70cbc6ddb39658a3341a4e068f546fb92c25fc22570c52

Observation 2e230541-cadb-4756-9b70-d2c741a229b6 · outbound

This paper cites Measuring Mathematical Problem Solving with the.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Measuring Mathematical Problem Solving with the

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.208929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.208929Z digest=sha256:bbf5c33d6f007459259aa724d15d0ebbb8c46c5a5ec898dc5618f379d78f1c78

Observation 1af62569-06fb-4acd-8790-49d4c7f075c3 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Evaluating Large Language Models Trained on Code

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.213714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.213714Z digest=sha256:392a210d564a2b2d78436d674d15991d74c8b55d0f6b75d842e44eebfcbdd74e

Observation 9e9f466b-31ca-452a-bddb-66650b598c04 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.217919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.217919Z digest=sha256:fc9ac8821eef665fa7b83df769acb87cc729e5b3d6a16ad1323d07a0da1c5ebd

Observation 449bc022-0803-4ee7-a5f8-0853ac5c1422 · outbound

This paper cites , booktitle =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models , booktitle =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.222227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.222227Z digest=sha256:aec5a93c52584097a256bae81caa180534402f75a69ad75e0004990cac55ecb3

Observation 9402f2e0-8cfe-4004-bd22-d54e179147fc · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Solving math word problems with process- and outcome-based feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.226014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.226014Z digest=sha256:80e75670469f4e2d98656b2aaf9ff0fe1e777532329be0918e99fda10635bfa7

Observation cf01f4ab-af70-47fd-999c-d4e52b0ae2fd · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.229462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.229462Z digest=sha256:379ced559603ec5a51f602f7400996bf2ae6fb01015395e88e01891853b837b0

Observation 0b3253d8-0166-4546-93b5-3bddc95a8dac · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Rewarding Progress: Scaling Automated Process Verifiers for

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.232686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.232686Z digest=sha256:35b7e47c26c295f994007e38091e0092a7c9d00cd393178c540b6a0b351701a6

Observation 053bdf03-8ec8-4c7a-9bcc-f397aa2347a0 · outbound

This paper cites 2024 , eprint =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models 2024 , eprint =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.237285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.237285Z digest=sha256:d2ab326c476705fc802a40a0a676a89ff71a33c89efd070f862813a93c862d22

Observation dfba234b-9852-4766-93de-bdd7310f3e7e · outbound

This paper cites Machine Learning , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Machine Learning , volume =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.241163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.241163Z digest=sha256:d4bc2000a2f60113739a165fe031a245ac2044bc6b43bb9873b89045e6276bb2

Observation 0c6fdf8c-64b4-40e2-a061-c0651c902dd1 · outbound

This paper cites International Conference on Learning Representations , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models International Conference on Learning Representations , year =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.244783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.244783Z digest=sha256:435c76ab8fa19afdc04c3d5d79a5b5aaa97a3ee1701527408ace19aba9297eec

Observation 15552ae1-12b0-43bf-bdb1-155e39000a17 · outbound

This paper cites Machine Learning , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Machine Learning , volume =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.248531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.248531Z digest=sha256:b999ec4f2a923dee16b87f8bfa264fa383fc55a3f4b494768cd725d2e141df73

Observation c2daa83c-01a0-4451-bf9e-cb6bf2d32177 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.252229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.252229Z digest=sha256:6950f6590ff2938d6694f5716d4060f23a35d4f968393e7cc3b628aeb90ea150

Observation f720ca1a-1281-41e2-90e7-3871528697f9 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Distilling the Knowledge in a Neural Network

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.255795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.255795Z digest=sha256:daa166ffe2de9298eede54833b928e3318da03366b9b656147f0be0d6804ef2e

Observation 6227b58b-9615-42d4-819c-f3e6439ce9cd · outbound

This paper cites Proceedings of the Conference on Empirical Methods in Natural Language Processing , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Proceedings of the Conference on Empirical Methods in Natural Language Processing , year =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.259231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.259231Z digest=sha256:a89e4b72ed8b285c86ff28549878677452c15f1073babc7d74d71d4053af0bd4

Observation 65642262-2d28-4179-b206-d993d7c2a956 · outbound

This paper cites an unresolved cited work.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.262599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.262599Z digest=sha256:ae0488f2f66ef37cebf9c95c8c8e24d21a6c724340639195196a7d8ebf8208e8

Observation 70374d51-c51e-4658-ab42-99885ab85a10 · outbound

This paper cites Proceedings of the International Conference on Artificial Intelligence and Statistics , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Proceedings of the International Conference on Artificial Intelligence and Statistics , year =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.266232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.266232Z digest=sha256:a1c9724f94bf1859f434eb093c5ec431acfe9682d9ca81677d66937603055908

Observation 9874e403-50a8-4014-aac4-aa1940815ee4 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.269547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.269547Z digest=sha256:ae2d1f1f90ba166e61f1ac6e5b5522f7676c8826480d50cfccade990350e0d31

Observation 071c1af5-d5ed-495a-869f-c18f20be8eb3 · outbound

This paper cites Neural Networks , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Neural Networks , volume =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.273289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.273289Z digest=sha256:8c4736e080df76fbbe02fea8b37f9dac87334b3a006477951b908834c15a8550

Observation 4671c1ef-63a1-4416-9fd9-f645f8167aa9 · outbound

This paper cites International Conference on Learning Representations , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models International Conference on Learning Representations , year =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.277474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.277474Z digest=sha256:88835b14cbd131c7831862687f602738f5207f601a1821d6dc8bf159ad7827b9

Observation 8587597b-2499-4693-acd1-2be192138a50 · outbound

This paper cites Robotics: Science and Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Robotics: Science and Systems , year =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.281431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.281431Z digest=sha256:4adb6e6403fb927ead48c805253526e7b50afb0beb0c168b2758df81a1abacab

Observation 42419323-f0a6-4535-b20a-e3e7d48fd639 · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Advances in Neural Information Processing Systems , year =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.285395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.285395Z digest=sha256:a3a9163ef19b091b974c22babbc3bb780667b9c798c8d877cc199aa19a62ff9a

Observation bd24587f-f8b2-4f33-9e31-5baaa52f35f3 · outbound

This paper cites Neural Computation , volume =.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Neural Computation , volume =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.289386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.289386Z digest=sha256:d55a34a5bffcf3c22151b7cf3ed42b9f94e87b2e19a73b4eb4602d0846235d35

Observation 39876256-6633-4d30-9909-20d932053f00 · outbound

This paper cites Learning Phrase Representations Using.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models Learning Phrase Representations Using

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.292577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.292577Z digest=sha256:8ec30974ed57306b67d8b753dae842261a8b6911c88f69d7ae7acf7f993c2682

Observation ad4b5534-944b-4956-a8e7-9dce4e4c1a7d · outbound

This paper cites AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals.

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:13.296565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:38:13.296565Z digest=sha256:3cab59e0a0fa030579a0d831f07e54e06f21e133e7f6ad1a65e2f260466ab9c9

Pith citing papers

No inbound Pith citation observations are available.