Pith. sign in

Paper Citation Record · LEDGER

Rethinking On-Policy Self-Distillation for Thinking Models

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2607.05184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05184 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T01:04:59.662046Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:59:51.920749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:02:35.499158Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact7
  • verified fuzzy5
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbba8739-0646-4fca-bb69-bfbedf14edfb · outbound

This paper cites Forking Paths in Neural Text Generation.

Rethinking On-Policy Self-Distillation for Thinking Models Forking Paths in Neural Text Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.525524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:ddf4a0d3cd9b1ba904b0b97a84b46547a34064eea78eca056e5bf38185f4283c

Observation 396f6a89-249f-40bd-88a8-769fa5891127 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Rethinking On-Policy Self-Distillation for Thinking Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.516491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:1816804fa7506d0ca351358d9e86fa8d33ccf5b068a14499d477da292481df4c

Observation eb8f50b2-3593-4b2d-b1de-da1768b4882b · outbound

This paper cites Chakravarthy, Anikait Singh, Nathan Lile, and Noah D.

Rethinking On-Policy Self-Distillation for Thinking Models Chakravarthy, Anikait Singh, Nathan Lile, and Noah D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.721941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:be4fd6ce3176e8adfe8da434e1790e6e8c91b88bd4bacee9e9e0092b94ae97e7

Observation 9146e4db-699d-489e-b982-0736d154a3fb · outbound

This paper cites Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.515390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:86ecfe78b16f6b7c5b9e8e899e2694080c5084eb2987ca4887505dec6c72ca13

Observation f542e8fc-6641-4546-b5c8-3f025de3e865 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Rethinking On-Policy Self-Distillation for Thinking Models Reinforcement Learning via Self-Distillation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.535875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:3ec316629908d3540624e5326e2d7d711a39238b363da9acca22559b0b14f6e6

Observation c3e91af8-c3b6-41ab-9971-e775278d404b · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Rethinking On-Policy Self-Distillation for Thinking Models Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.528143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:3ce894b99c46a7c21186ef78cd385ccd3bc2b465d644a050c636f4f490805f91

Observation bdadb457-eb31-4c85-8b4e-19aedac412ad · outbound

This paper cites Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability.

Rethinking On-Policy Self-Distillation for Thinking Models Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.539267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:121f868a99df3c7d842bcafe022292f71ca08f4cadde04e4c976ff37cac75dbf

Observation e802dbe8-8581-4ab8-9b8b-9a0d04549f24 · outbound

This paper cites Pope: Learning to reason on hard problems via privileged on-policy exploration.

Rethinking On-Policy Self-Distillation for Thinking Models Pope: Learning to reason on hard problems via privileged on-policy exploration

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:14:27.537447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:9e8d8e687e624cbd3be17de40917d9bb30cf01d1163153eb04e8469ad8c0791a

Observation 36b68aad-dc12-4f62-80a3-5482845b9328 · outbound

This paper cites Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes.

Rethinking On-Policy Self-Distillation for Thinking Models Reuse your flops: Scaling rl on hard problems by conditioning on very off-policy prefixes

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T01:14:27.529342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:ab254494a3d01f54643e20b98bb7adaaabbbb5830fb1914c1dbee4d9cd1c7ebd

Observation be81aa2d-827e-4823-aac3-1f8519c3e5f3 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distillation Enables Continual Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.531170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:a692efe6265f6337940d34752021d96e1151f74b0959b0b5f853e8ba4008d168

Observation 1a4d5b85-22a0-4110-8d15-efd0c35c8128 · outbound

This paper cites Olmo 3.

Rethinking On-Policy Self-Distillation for Thinking Models Olmo 3

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.545307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:8b3d704789bb17f271537e1647373492365df7db3dfae6a67069ac9f414658b1

Observation afbf9a04-0dcf-4fd0-a8b4-6354c8bd52cf · outbound

This paper cites Understanding reasoning in thinking language models via steering vectors.

Rethinking On-Policy Self-Distillation for Thinking Models Understanding reasoning in thinking language models via steering vectors

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.732244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:cbe462f0f7db71a09257d34423133176a6cbdbaf5c66a93ce9e9a27dafe8d877

Observation 9e148a67-9f7d-4410-9d2a-018a29fd6c7e · outbound

This paper cites Qwen3 Technical Report.

Rethinking On-Policy Self-Distillation for Thinking Models Qwen3 Technical Report

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T01:14:27.542787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:66dd73f7aed910f53203ca470b7b278ffe50d21a54ed1792b17ab27025e21776

Observation c8272c1a-c400-4d83-9492-cf6d3c41167c · outbound

This paper cites Embarrassingly Simple Self-Distillation Improves Code Generation.

Rethinking On-Policy Self-Distillation for Thinking Models Embarrassingly Simple Self-Distillation Improves Code Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.525726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:705f6d18ae81f905ad69665df22a7e3de80045d93063e3a7407ac43ddd85613d

Observation a309d32f-8968-435e-b44c-7157562fe251 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Rethinking On-Policy Self-Distillation for Thinking Models Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T01:14:27.501845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:825c5578da486c72e7a7bc5549fa540f68111956f9bfd17be8486d86b9eabdbc

Observation 37258f2f-723a-488e-ad56-29a8a099477e · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.730229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:f471e8f650fa7178dda9b92b154f64bc1ef321ff78aeb7b825db9ffa15db8ab5

Observation 52c3f718-42b9-4900-b8bb-3706cb76ae18 · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.720002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:d0abe78f81c68fe34724578ffdd52f44e5614d03e1d5ad718ade36216e2ee98b

Observation dd3ad507-5f89-409f-b287-639a407cfec3 · outbound

This paper cites The Average column averages the three benchmarks.

Rethinking On-Policy Self-Distillation for Thinking Models The Average column averages the three benchmarks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.726381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:fe0ef4e652f8a46ee401ffd5fa418791b76ca1de472da8e660d010b0f9bcd43c

Observation d7c7a81d-4e28-47fb-90ac-2186c707c26b · outbound

This paper cites Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set.

Rethinking On-Policy Self-Distillation for Thinking Models Epistemic-token OPD applies the same loss only to tokens in the epistemic-marker set

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.723944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:bb6cef2954fcf998693b6f41e226733dcf51c3c1a6949244715b727690e2450e

Observation cbf32b64-d01a-4caf-b5e1-191965f8354f · outbound

This paper cites an unresolved cited work.

Rethinking On-Policy Self-Distillation for Thinking Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-08T01:14:27.734544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:bf1637f1861a5a5ae56ae7fb6b31d8a7cca91ebcae585dcfe7c5a6d8dc7da592

Observation dcfd014f-39a6-41f3-bcd6-0904ee2b91f3 · outbound

This paper cites Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context.

Rethinking On-Policy Self-Distillation for Thinking Models Blue curves are base thinking models, orange curves are OPSD with full gold-demonstration context, and green curves are OPSD with final-answer-only privileged context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T01:14:27.728454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T01:04:59.662046Z digest=sha256:9ce5023f5ad68edde3a404d355d7b072ca087e10f215973c72cffbd71127a210

Pith citing papers

Observation 713f8264-9954-41ac-9f95-00c226482406 · inbound

DAPD: Dual-Anchored Policy Distillation cites this paper.

DAPD: Dual-Anchored Policy Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 76

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T22:02:35.540632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T22:02:33.791583Z digest=sha256:6e3e4bbe184c0d1f8c4c20fd10c356527c73c242ec2f8c9cde450256b36c7bab

Observation de8a8bae-340f-406b-b2b5-d2ebd0a6a75e · inbound

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation cites this paper.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:05.711475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:05.711475Z digest=sha256:492f85847737ba9bc545b22ba21959bcc483f1aa1b52b4b507dcd445ded43ab7

Observation 75ab8712-8de3-4b5a-b064-92fb51e46615 · inbound

Simple-OPD: Demystifying Warm-up for On-policy Distillation cites this paper.

Simple-OPD: Demystifying Warm-up for On-policy Distillation Rethinking On-Policy Self-Distillation for Thinking Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:49.690977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:33:49.690977Z digest=sha256:a316b7fb57ebdf0d192e81dc93fcea7756387e77781334a469d57f53c0d5a7cd

Observation 770dfc69-243a-40f7-97cf-8d7b3ebea9e2 · inbound

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ cites this paper.

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ Rethinking On-Policy Self-Distillation for Thinking Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T10:59:51.920749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:59:51.920749Z digest=sha256:317b6789cb40586dca677fd56a3bd911c337a10ec9beb4f675b30f80da43cdc3