Pith. sign in

Paper Citation Record · LEDGER

Rethinking On-Policy Self-Distillation for Thinking Models

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2607.05184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05184 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T01:04:59.662046Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:05.711475Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:02:35.499158Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact7
  • verified fuzzy5
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbba8739-0646-4fca-bb69-bfbedf14edfb · outbound

This paper cites Forking Paths in Neural Text Generation.

Rethinking On-Policy Self-Distillation for Thinking Models Forking Paths in Neural Text Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.525524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:42fec9fbb9fba811a5ffd918770b686edeb5a071f10ccaf98ac8afbe7b4affe7

Observation 396f6a89-249f-40bd-88a8-769fa5891127 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rethinking On-Policy Self-Distillation for Thinking Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.516491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:f8626efebbd8f64417acbca8b4baf05d4cd721804dbaa3fbcb34f20e5c3c41d4

Observation eb8f50b2-3593-4b2d-b1de-da1768b4882b · outbound

This paper cites Chakravarthy, Anikait Singh, Nathan Lile, and Noah D.

Rethinking On-Policy Self-Distillation for Thinking Models Chakravarthy, Anikait Singh, Nathan Lile, and Noah D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.721941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:2313415f4e5eb81404d81a74adfe4b82ee2c34d5629a47da5723d46de95bcd59

Observation 9146e4db-699d-489e-b982-0736d154a3fb · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.515390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:bcb00ef80f465c30dd715298c815a02b4c5618fe7c132916788883d37e96f8b1

Observation f542e8fc-6641-4546-b5c8-3f025de3e865 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Rethinking On-Policy Self-Distillation for Thinking Models Reinforcement Learning via Self-Distillation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.535875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:71a5d304d98ffb14d511c9767a9f30ae7562bf573a0b3912fc5aa2b713ae250f

Observation c3e91af8-c3b6-41ab-9971-e775278d404b · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Rethinking On-Policy Self-Distillation for Thinking Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.528143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:efd49c923f42943d5f21c143545744d3936936f221648b9fa87ae308278b810e

Observation bdadb457-eb31-4c85-8b4e-19aedac412ad · outbound

This paper cites Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability.

Rethinking On-Policy Self-Distillation for Thinking Models Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.539267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:5f49aadb917c9ff59392cf0e4655e92fb29585b781aa3450df8663a228be207f

Observation e802dbe8-8581-4ab8-9b8b-9a0d04549f24 · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.

Rethinking On-Policy Self-Distillation for Thinking Models Pope: Learning to reason on hard problems via privileged on-policy exploration

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:14:27.537447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:f3ec127c94dac39225a139e2643296563b9f54dd5fde22795c8682aaed767fa5

Observation 36b68aad-dc12-4f62-80a3-5482845b9328 · outbound

This paper cites Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes.

Rethinking On-Policy Self-Distillation for Thinking Models Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:14:27.529342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:acb850b7b5021fc1611478c7614844d795353dc6cf0ca25c0b9d6f42bb2cade2

Observation be81aa2d-827e-4823-aac3-1f8519c3e5f3 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Enables Continual Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.531170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:cc28e03396d69956292a70d961d2d1928cb9f0e33d8ac67fb69a1cdb462e4f13

Observation 1a4d5b85-22a0-4110-8d15-efd0c35c8128 · outbound

This paper cites Olmo 3.

Rethinking On-Policy Self-Distillation for Thinking Models Olmo 3

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.545307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:20c8eb5ce3d962a0f44fdb5efcc288ee30899b7d641eb0b382eaf7e0ff874324

Observation afbf9a04-0dcf-4fd0-a8b4-6354c8bd52cf · outbound

This paper cites Understanding reasoning in thinking language models via steering vectors.

Rethinking On-Policy Self-Distillation for Thinking Models Understanding reasoning in thinking language models via steering vectors

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.732244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:9e71604aecbe98f741d155fb357c1dce158c21d50d633f37c1db4efeaf5dba20

Observation 9e148a67-9f7d-4410-9d2a-018a29fd6c7e · outbound

This paper cites Qwen3 Technical Report.

Rethinking On-Policy Self-Distillation for Thinking Models Qwen3 Technical Report

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.542787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:c7a6a84a37b4148e5db46f4dd3985f34d0b9abcb950e75831038f90267fc0173

Observation c8272c1a-c400-4d83-9492-cf6d3c41167c · outbound

This paper cites Embarrassingly Simple Self-Distillation Improves Code Generation.

Rethinking On-Policy Self-Distillation for Thinking Models Embarrassingly Simple Self-Distillation Improves Code Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.525726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:c26e4d023e72953726e5d5b5c7857db5d251049d8be606b2bcca4a9b8f7f1838

Observation a309d32f-8968-435e-b44c-7157562fe251 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.501845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:de0f0b33bd2367339f12a60e0e2497d15a17e233476fed8185bc1c4821a99873

Observation 37258f2f-723a-488e-ad56-29a8a099477e · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.730229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:4bee300111d86b98199f95abc72818df530895c636ed5df0220de3b139d84e5d

Observation 52c3f718-42b9-4900-b8bb-3706cb76ae18 · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.720002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:1bfcb9d2847955c50d2ac3e069a0763147cebeb5a49c02c2e2798223a7ade474

Observation dd3ad507-5f89-409f-b287-639a407cfec3 · outbound

This paper cites The Average column averages the three benchmarks.

Rethinking On-Policy Self-Distillation for Thinking Models The Average column averages the three benchmarks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.726381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:31a21d9c2c9ee8a8db7bc075414175cfbaacccb5b1697d37d47129ec1f3e21fb

Observation d7c7a81d-4e28-47fb-90ac-2186c707c26b · outbound

This paper cites Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set.

Rethinking On-Policy Self-Distillation for Thinking Models Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.723944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:2f04d5e693f2891091cfff62e26d1e58c0970b8c4c05bc09b166159cbe39785f

Observation cbf32b64-d01a-4caf-b5e1-191965f8354f · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.734544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:1cb46e7d3fb961aeeda0e38cb947597d512bc111954f019d160b02f3240973e8

Observation dcfd014f-39a6-41f3-bcd6-0904ee2b91f3 · outbound

This paper cites Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context.

Rethinking On-Policy Self-Distillation for Thinking Models Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.728454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:3da2096a7014eb19abf463662b1e5fde42d3c92524d53e089c8457ba039b79b8

Pith citing papers

Observation 713f8264-9954-41ac-9f95-00c226482406 · inbound

DAPD: Dual-Anchored Policy Distillation cites this paper.

DAPD: Dual-Anchored Policy Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T22:02:35.540632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T22:02:33.791583Z digest=sha256:8f7c3605d80a38ede5ec0d237fc597ec0f3b0e7914d0edff931d0046f6ce4f6a

Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · inbound

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation cites this paper.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.711475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.711475Z digest=sha256:7d6bed4bb01a24f3a23648f21d55d67d6a5d689469a00a17a7b1db8bbfcf1e8d