Pith. sign in

Paper Citation Record · LEDGER

Weak-to-Strong Generalization via Direct On-Policy Distillation

As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 4 inbound Pith citation observations for arXiv:2607.05394.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05394 v2

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T07:01:56.628017Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:07:43.400511Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6384b32f-a969-43f8-a826-5270904df0ff · outbound

This paper cites DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.

Weak-to-Strong Generalization via Direct On-Policy Distillation DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:3e993b0433b1853d2f5c01b176faa3a4dcb782566dee862bf661ffd225bad799

Observation 781713bd-26e9-4bf9-a4c7-0a8c157c27ad · outbound

This paper cites JustRL: Scaling a 1.5B LLM with a simple RL recipe.

Weak-to-Strong Generalization via Direct On-Policy Distillation JustRL: Scaling a 1.5B LLM with a simple RL recipe

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:c136b64989c3416da424bf99a6cf88aaba5f4e2db5e3fd6ddb6653189994eca9

Observation 7d872834-608d-47e3-bd86-2822c9332ba4 · outbound

This paper cites Qwen3 Technical Report.

Weak-to-Strong Generalization via Direct On-Policy Distillation Qwen3 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:c53e9e9b61851261de53dd48994a1f34fc048d43b79e5b11da1007009dae0667

Observation 5aa5a17f-54ca-4445-8b18-565ae2510235 · outbound

This paper cites POLARIS: A post-training recipe for scaling reinforcement learning on reasoning models.

Weak-to-Strong Generalization via Direct On-Policy Distillation POLARIS: A post-training recipe for scaling reinforcement learning on reasoning models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:44f2507b6cd0bd7dd1deb14cfddc5256e21c2c13f45c28ae9bcb2e9a43f95378

Observation 84c1681c-4087-4f32-80df-36b75dd0da39 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Weak-to-Strong Generalization via Direct On-Policy Distillation Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:e98db766f7693261ac2b8712a7d6eb23abd9baa4e86ef5af80a73c81af476781

Observation 36d5cf22-a34d-4ab9-91ff-2a758cec96f9 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Weak-to-Strong Generalization via Direct On-Policy Distillation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:45c52ae1d8dc15b17e7f7d43e49922f690940ffbd92ecdb9b267e554f887fdb1

Observation 9653133b-8622-4b08-8a47-917e10dfefc3 · outbound

This paper cites QuestA: Expanding reasoning capacity in LLMs via question augmentation.arXiv preprint arXiv:2507.13266, 2025.

Weak-to-Strong Generalization via Direct On-Policy Distillation QuestA: Expanding reasoning capacity in LLMs via question augmentation.arXiv preprint arXiv:2507.13266, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0c78d78e50a2fcbc87d364fda1c2d2a92d7bb400c7fd1eb2da1acd8488b69c03

Observation e94af6ad-635d-4f35-9448-7b73035eb864 · outbound

This paper cites Lyng, Sanjit Singh Batra, and Robert E.

Weak-to-Strong Generalization via Direct On-Policy Distillation Lyng, Sanjit Singh Batra, and Robert E

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:5d583777a189e49027e30a1d969dc0637e3657e5d6bb11e171bdbb4c75afe12c

Observation 50deb921-7e82-4e45-a46f-85698de79d12 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:758e67c25f0e5fa9f771ad573759ca02f162ff4e05f578f9cb1cd5b0178b375f

Observation a4d3987e-c861-479b-96b7-fa783a730760 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Weak-to-Strong Generalization via Direct On-Policy Distillation Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:77f36086e539fffbf9d33e449c0aa46e7dbac94c9624a856640d69cb243281b0

Observation 94210e6c-4e52-4c8f-a8a9-bf63ba3344f3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Weak-to-Strong Generalization via Direct On-Policy Distillation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:2e66cf5d22d89d0217eeeb5d6fb471ab7ab620771a61312098c57fe6ba72a269

Observation f3eee08c-fcbf-4d05-9472-662838886b63 · outbound

This paper cites QwQ-32B: Embracing the power of reinforcement learning.https://qwenlm.github.io/blog/qw q-32b/, 2025.

Weak-to-Strong Generalization via Direct On-Policy Distillation QwQ-32B: Embracing the power of reinforcement learning.https://qwenlm.github.io/blog/qw q-32b/, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:00c619a4dd3a79119ce3d7fbe38e59f39dc4368b0df8c460467308e14ec2dbeb

Observation fd5107ec-5293-491b-8077-284ae9ad5976 · outbound

This paper cites Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/Ope n-Reasoner-Zero/Open-Reasoner-Zero, 2025.

Weak-to-Strong Generalization via Direct On-Policy Distillation Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/Ope n-Reasoner-Zero/Open-Reasoner-Zero, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:ea4fde099a5b6c4478338beb43246af30b29c140dcf200c2a49634c06926a475

Observation 79b534cb-3bb6-4c11-b93f-3ea3b04600c4 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Weak-to-Strong Generalization via Direct On-Policy Distillation DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:2bfd0056abb0fd4511f84771c0454f5cb6795adb9be27edf5449cda799240921

Observation c2ee0745-8e4e-40a9-a859-7f17b88c5fba · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Weak-to-Strong Generalization via Direct On-Policy Distillation Skywork Open Reasoner 1 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:e36d3cae6eb9d0f4231800cde52d0a56f86b604864b1fa81c62a05c78f26bf84

Observation 6cf49e12-3bfb-4301-9504-f6cef34cf187 · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Weak-to-Strong Generalization via Direct On-Policy Distillation DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:bbb24953efdf49d57bc1819e4f43fe0beea2fb9feafdf84173f8ad8dc1a63764

Observation a95b12cb-86f7-4ad0-a343-52cd6930c9cd · outbound

This paper cites How far can unsupervised rlvr scale llm training?arXiv preprintarXiv:2603.08660, 2026.

Weak-to-Strong Generalization via Direct On-Policy Distillation How far can unsupervised rlvr scale llm training?arXiv preprintarXiv:2603.08660, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:2185f791f79fabf2839b33e768dc24d2dfc19c234998972e54791515e0d36043

Observation 9a4c5fb1-4e19-45f4-aa99-5a4411c09190 · outbound

This paper cites DeepSeek V4 preview release.https://api-docs.deepseek.com/news/news260424, 2026.

Weak-to-Strong Generalization via Direct On-Policy Distillation DeepSeek V4 preview release.https://api-docs.deepseek.com/news/news260424, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d54fe944809d74b9794ad2d74524ec5724855b640cd06af1d085a205a06cc353

Observation 3b2f48df-f0f0-49ee-8791-83023a08cf92 · outbound

This paper cites GLM-5.2: Built for long-horizon tasks.https://z.ai/blog/glm-5.2, 2026.

Weak-to-Strong Generalization via Direct On-Policy Distillation GLM-5.2: Built for long-horizon tasks.https://z.ai/blog/glm-5.2, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:bd0ad21f56d0873420db2edbce69f9ac0d17c9884789180c454ea4bd7160ff69

Observation 75e7e64b-cea9-457a-b420-8432ff36543f · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Weak-to-Strong Generalization via Direct On-Policy Distillation AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:8a6f7e64fc18ba56c801aa645153b9352c0d2a7deb14e7c1e91ac5a38cfff08d

Observation 1b8404f6-722a-4ebb-a24f-a5ac0a817d8d · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation OpenThoughts: Data Recipes for Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:c15a6ab4158eec68a98c376adb41c7bf73d7d10362af7c41bff85ee5db9aad52

Observation 4581e08f-2d0a-40d5-9a7c-43a41ba0900a · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.

Weak-to-Strong Generalization via Direct On-Policy Distillation Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d24f141632fbb7077b82f1e865728300412e85d998d39a7469cea318496437a1

Observation ecc0a348-7280-492a-985b-560af29ffa29 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Weak-to-Strong Generalization via Direct On-Policy Distillation Distilling the Knowledge in a Neural Network

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:ddb194a85dcd16e520560da92ba5c72dfae8c073e5f82f2fe2f17ea9a55dcc77

Observation 2adf695e-99c8-4162-a617-959b9e5f835d · outbound

This paper cites Sequence-level knowledge distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation Sequence-level knowledge distillation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0610df3ce6ecaa624480b5cfb179ef7c10b5f977253357be87dcf2a391ba66da

Observation 0b6e95a5-65dc-4b9e-bde8-07b4ce1c8e55 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Weak-to-Strong Generalization via Direct On-Policy Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:6e12b9a006fd35cc216c0abf81e997dd084c1613fc100816ed2f16cf35b61aba

Observation a69c8719-5814-4132-ace2-c08798ad8e0f · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

Weak-to-Strong Generalization via Direct On-Policy Distillation Tinybert: Distilling bert for natural language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:e9118a9033d1eb3de157588b6672c345c20d333e08bf7e6b78c5180312396bd2

Observation 1008496c-3766-4143-be47-59767fdbbf98 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

Weak-to-Strong Generalization via Direct On-Policy Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:7f6e8384bdd74caf34b551f4fd6264037a8ef6e2aa9e46ddca933f95ab742701

Observation 05b37e31-2aab-44df-b3a7-9cbacb74238c · outbound

This paper cites Improved knowledge distillation via teacher assistant.

Weak-to-Strong Generalization via Direct On-Policy Distillation Improved knowledge distillation via teacher assistant

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:35b9c968f128f894b24749a05dd70e7beb9a7cd7e30f1145547b7ed30d2f0faf

Observation 6923504a-4d05-46ff-b1e0-b178db0f4a6f · outbound

This paper cites On the efficacy of knowledge distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation On the efficacy of knowledge distillation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:8f910f9272c19c23d6d3b385c93d761bfcebbcf105f1242a2be9ea64cabae3c7

Observation 39e87fb7-d8c1-4488-a56e-e9092c417205 · outbound

This paper cites Distillation Scaling Laws.

Weak-to-Strong Generalization via Direct On-Policy Distillation Distillation Scaling Laws

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:22ae11cd44551a42fc74e16bb1f21ba53ac4287e8152d67495086f3216961d36

Observation 50f424f5-2513-4320-b553-1367fbc27c14 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation MiniLLM: On-Policy Distillation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:b7f03f9dfebd1ae7c316ebfb7ebfe8468993d102c69bef888fd6a5c70626093f

Observation 41d7ce90-7def-4c90-9ae3-971c097dea1f · outbound

This paper cites f-Divergence Minimization for Sequence-Level Knowledge Distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation f-Divergence Minimization for Sequence-Level Knowledge Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:61c2239d39a901b0e2ebc6078611f2c753ffd2ecc9f89864a7a77d9f1176d7b8

Observation 06919646-9ff1-4cdd-b87a-59e50e152d61 · outbound

This paper cites Revisiting Knowledge Distillation for Autoregressive Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Revisiting Knowledge Distillation for Autoregressive Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:9c125e0925b3dea0b16049ea6ed56c2aa37262cce39c61cce8f03504a41847a5

Observation 5cbe331a-0496-4e58-9684-cd0b1ebcb88c · outbound

This paper cites Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Rethinking Kullback-Leibler Divergence in Knowledge Distillation for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:8b409d52887056268722676e0eb889bb6964144e265886d2f8c45df16711b14f

Observation e193e4c4-d4be-4125-9b9c-a5b283fa6f3b · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:08cb27ac9f22d81888efe39d48343b27146e32c0570db9c86eb03e8498c37532

Observation 02d101de-8bd5-4513-a448-a79bafb8a2e5 · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Weak-to-Strong Generalization via Direct On-Policy Distillation DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f05e6bb135c5613decb4254f90ce140f279015c689c6f174effb012689751223

Observation ae0c877e-fa71-480a-bd99-fb47ceb9b5e6 · outbound

This paper cites MiniPLM: Knowledge Distillation for Pre-Training Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation MiniPLM: Knowledge Distillation for Pre-Training Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:b621d9249828335677b8769acafa7da0ef1ee770d03ec81ece894706ae6c5862

Observation fe7aab0c-affa-4812-afc1-8f3420495a68 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:6891379ef7c780d204a505c8bceaecbb2cfe280e4752efa7ffaf8f7fab80f1d6

Observation f77a8127-b322-463a-9b50-9174f09f7466 · outbound

This paper cites Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.

Weak-to-Strong Generalization via Direct On-Policy Distillation Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:5d29f54e3ecf4964148ee832ee0facc4dcd6b3afdf56b757e6848568b08b1061

Observation 024037ba-1aa8-4095-a577-2a05dace3ead · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Weak-to-Strong Generalization via Direct On-Policy Distillation On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f5bcf53351175913619dc6fd28d1823ddd1f6139c096a147997272b268d937f8

Observation 728b120e-b2c7-4ebd-8f93-6ed3e04fcb5d · outbound

This paper cites On-policy distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation On-policy distillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:3a54af8afd5b8db52fadb5d1134a64d1f6854d7de0081a8ee5031513b85b9c57

Observation f79f2454-7766-4eb6-a227-81f81e99d52c · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation A Survey of On-Policy Distillation for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:8a52f186f785d3c4061b5faac7ae76c3749174e2367b30f82d2161c56379dce3

Observation 4e3d80cd-222f-4bf9-8239-ce5b1f2ac704 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Weak-to-Strong Generalization via Direct On-Policy Distillation Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d41052c8eed8668f8635968410ff0168fa368463b65ee98fc41d9fdd5ef7f709

Observation 10e7219c-baa6-428c-ada3-027eeaf43a71 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Entropy-Aware On-Policy Distillation of Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f966345acd0570c6f0c620e8c280c79e7fb82546968eb5a6131bd1d5ecd391b5

Observation 047a9d32-1acc-419c-a421-c27d9081f21f · outbound

This paper cites Stable On-Policy Distillation through Adaptive Target Reformulation.

Weak-to-Strong Generalization via Direct On-Policy Distillation Stable On-Policy Distillation through Adaptive Target Reformulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:ef18e0ac3cf08585eac8c89caf3ea5c0fb7065a8e57af256a5ab885524b257ba

Observation 634d9d0a-0f4b-4104-87d3-a211920d3e77 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.

Weak-to-Strong Generalization via Direct On-Policy Distillation Unifying group-relative and self-distillation policy optimization via sample routing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:b3feed9b7aae245c7a75446239315a3d00aefdd142d58175e843ecb6a8c29aef

Observation 590650dd-5a62-415d-917c-4b7c0c073182 · outbound

This paper cites On-Policy Context Distillation for Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation On-Policy Context Distillation for Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:590aa9ee065e81d762f7344c973664ba8b93d5d38fd726edbaa4811b8e0969d1

Observation 18c662a1-119b-458d-96fd-5c585807adb7 · outbound

This paper cites Online Experiential Learning for Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Online Experiential Learning for Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d0ba04c8be0168c40c141843d6b6fcbb38ac08a0dd769d24811f7cdcdc3d6d16

Observation edc387dc-4cc2-4c0f-b0e1-d663a455cabf · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f5bf8b5036e37325296fbcf98cb36ad4ca849db98cd9532b902789431f707ebe

Observation f80e56c1-8224-4c47-a94f-165fd0dd53b4 · outbound

This paper cites Self-distillation for multi-token prediction, 2026.arXiv preprint arXiv:2603.23911, 2026.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-distillation for multi-token prediction, 2026.arXiv preprint arXiv:2603.23911, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:fff7ced62f9a5fb159cc51a08d4aa192e0c2c2c778f7101344fa80d0202e3416

Observation 97a91287-e81b-447f-ba55-06cc450123c2 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation Reinforcement Learning via Self-Distillation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:bcca562ab41e1c289997089b86c6c0b8537a45a9994b126d1859d733fa836381

Observation 90474726-e136-48de-9e3b-6e4428f2baf5 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Distillation Enables Continual Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:36f9131958fd25ac04e48207d617d311a5ede51463afad45d5c46379f17905a9

Observation ee600dd5-d64a-4e2f-b6f9-592d71db122c · outbound

This paper cites PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence.

Weak-to-Strong Generalization via Direct On-Policy Distillation PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0905fff0836730f92fa2a85409d0f1e88c569366dbda12905e85e084886b4219

Observation 4348ce7b-e70f-47d8-80a7-4351fa53540f · outbound

This paper cites Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137, 2026.

Weak-to-Strong Generalization via Direct On-Policy Distillation Scaling reasoning efficiently via relaxed on-policy distillation.arXiv preprint arXiv:2603.11137, 2026

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:3047f5eb065adb2ea1f01d1ed8ba18eaeebc4d1a6ff8031ba78c27312de60058

Observation c5cc2cab-e0d2-4d53-a4cd-ec2741f3039b · outbound

This paper cites Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?.

Weak-to-Strong Generalization via Direct On-Policy Distillation Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:40b4e728632ba178f8d5af079886a94122e12cbe10d038c3c830b80f50c83cf9

Observation e211856f-990b-4702-8110-4a7c6b1588a3 · outbound

This paper cites Black-Box On-Policy Distillation of Large Language Models.arXiv preprint arXiv:2511.10643, 2025.

Weak-to-Strong Generalization via Direct On-Policy Distillation Black-Box On-Policy Distillation of Large Language Models.arXiv preprint arXiv:2511.10643, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0d20be4e4f5daee262d680451bfdc9975f6c28cd944c0a00f653bfc886644073

Observation 0d55ba62-ba98-4143-9c40-df5269453ede · outbound

This paper cites MiMo-V2-Flash Technical Report.

Weak-to-Strong Generalization via Direct On-Policy Distillation MiMo-V2-Flash Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d9bdc27d8434d7a56f276542739a794fac545f6f9d771f3ce1f5257c370fbc88

Observation 276f615d-744d-43a4-9f4a-e18d4b387fe2 · outbound

This paper cites an unresolved cited work.

Weak-to-Strong Generalization via Direct On-Policy Distillation Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d04b80ff5eaebc64c243b9251e3800b6d807df52ba2381bd93b2a2f1ba1b070f

Observation 64d41ed5-9b4c-4adb-a93e-dcf3a9891769 · outbound

This paper cites KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning.

Weak-to-Strong Generalization via Direct On-Policy Distillation KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:5dc9bf05acd2a7666a317f91167690f53ec2b87fbbdc769603f83ef9c3c41126

Observation 468560f9-9e50-4064-8267-4f0333143a5f · outbound

This paper cites Reinforcement-aware Knowledge Distillation for LLM Reasoning.

Weak-to-Strong Generalization via Direct On-Policy Distillation Reinforcement-aware Knowledge Distillation for LLM Reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:1f12df2479cf7b5dab25a415f76912a51412f4d7d22f7bd11d6d14db79a96fa5

Observation dae33936-2343-424a-a54f-d8b6cb275983 · outbound

This paper cites AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation.

Weak-to-Strong Generalization via Direct On-Policy Distillation AlignDistil: Token-Level Language Model Alignment as Adaptive Policy Distillation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f775266013e425be25b94275e32a04d23c80cf84441743405b25045b1460ee2e

Observation 199e62b2-4e82-4319-b42c-2c9f4923472e · outbound

This paper cites Self-Distilled RLVR.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Distilled RLVR

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:74ccf9648fa2d19e43c83c43bd79cd6f2712e16e85b4669cadc9f541a7d5f28b

Observation 90af5db4-293a-48db-b8ee-8fcac00d6648 · outbound

This paper cites Small models struggle to learn from strong reasoners.

Weak-to-Strong Generalization via Direct On-Policy Distillation Small models struggle to learn from strong reasoners

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:e845ad3b74e0a0e98a947459220b82dfb7858f7a114dccba12ac92fa7bb3229e

Observation 7f6fa9c3-f691-4d54-bab7-22445a3644f7 · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Weak-to-Strong Generalization via Direct On-Policy Distillation Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0f130a60748fc5641b412e9896c2137b59469da8b20364d137055f9fb4feef6f

Observation d4278c51-8044-46dd-8869-3e9a187f8195 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Weak-to-Strong Generalization via Direct On-Policy Distillation Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f527b0abcf8d63f0731c9641c7871b47463ddb3726d315659fe00eac5e164271

Observation 85d9953a-31c0-4406-b7bc-bab71cdbf235 · outbound

This paper cites Introducing Superalignment.OpenAI Blog, 2023.

Weak-to-Strong Generalization via Direct On-Policy Distillation Introducing Superalignment.OpenAI Blog, 2023

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:96bc8e359b519bf7486330859d52214b5e71f03b3c519bbb45616138e600ee7c

Observation 8fff07fb-add9-4446-8386-8f388a8a2e3b · outbound

This paper cites Semi-supervised learning by entropy minimization.Advances in neural information processing systems, 17, 2004.

Weak-to-Strong Generalization via Direct On-Policy Distillation Semi-supervised learning by entropy minimization.Advances in neural information processing systems, 17, 2004

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:8f1ddb372798b96a912d9a65de727bace5c18933f601ee8e0aeab20193864699

Observation 7bcd7216-249b-4b3a-a233-65f4d84875b7 · outbound

This paper cites Semi-supervised sequence learning.Advances in neural information processing systems, 28, 2015.

Weak-to-Strong Generalization via Direct On-Policy Distillation Semi-supervised sequence learning.Advances in neural information processing systems, 28, 2015

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d93d9114c68330a4a66f9f454e0b7f585a42bae2e50b8e8a71ac49ab78f46bca

Observation 34d855c2-7bc0-4846-aeed-7732e5b6d353 · outbound

This paper cites Temporal Ensembling for Semi-Supervised Learning.

Weak-to-Strong Generalization via Direct On-Policy Distillation Temporal Ensembling for Semi-Supervised Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:050e75399327ad1c480db45119e9083962c170ffb1c83f64c814e7c5260bf5b2

Observation 3028b6e0-7e46-4822-97ff-68d65d743ec6 · outbound

This paper cites There Are Many Consistent Explanations of Unlabeled Data: Why You Should Average.

Weak-to-Strong Generalization via Direct On-Policy Distillation There Are Many Consistent Explanations of Unlabeled Data: Why You Should Average

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d77b217aa1ed72df606b905d20b1223ec1e1d6262d77e33349f223507fbadacd

Observation 82abf378-ab71-4341-8a32-0c2805e6112f · outbound

This paper cites Co- teaching: Robust training of deep neural networks with extremely noisy labels.Advances in neural information processing systems, 31, 2018.

Weak-to-Strong Generalization via Direct On-Policy Distillation Co- teaching: Robust training of deep neural networks with extremely noisy labels.Advances in neural information processing systems, 31, 2018

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:aa5f5b9553633268d07ecdf3cd5ebd65fac2f2996c3532531c3879183e5a9726

Observation 00ead182-98ff-4c1c-88df-9ae548532948 · outbound

This paper cites Mixmatch: A holistic approach to semi-supervised learning.Advancesin neural information processing systems, 32, 2019.

Weak-to-Strong Generalization via Direct On-Policy Distillation Mixmatch: A holistic approach to semi-supervised learning.Advancesin neural information processing systems, 32, 2019

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0d74e3641b388348416e51e6a3beaa717fa4343fa920b8560b1671d5afaa4b72

Observation f07f7858-d323-44b4-b4d7-2cbf87c19d7a · outbound

This paper cites DivideMix: Learning with Noisy Labels as Semi-supervised Learning.

Weak-to-Strong Generalization via Direct On-Policy Distillation DivideMix: Learning with Noisy Labels as Semi-supervised Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:369bead770a477d0b45a5c0a721b8efc6783ec56402383225b55284be37a2067

Observation 950dff98-1d8e-443f-b012-6690f257a3cc · outbound

This paper cites Big self-supervised models are strong semi-supervised learners.Advancesin neural information processing systems, 33:22243–22255, 2020.

Weak-to-Strong Generalization via Direct On-Policy Distillation Big self-supervised models are strong semi-supervised learners.Advancesin neural information processing systems, 33:22243–22255, 2020

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:f755f241a83d4d02f1e66baca47e215e938d5f50f64ad6544c612f6c77a048c0

Observation 2c27b717-4ded-4395-9db9-0e282e795c8c · outbound

This paper cites Self-training avoids using spurious features under domain shift.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-training avoids using spurious features under domain shift

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:6b572eb838a3516ca762890305b601d167dbfc4a97f85db76e32ae604c90ac57

Observation 64ef02a0-7147-45f0-8af1-8039bc4c3f2c · outbound

This paper cites Debiased self-training for semi-supervised learning.Advances in Neural Information Processing Systems, 35:32424–32437, 2022.

Weak-to-Strong Generalization via Direct On-Policy Distillation Debiased self-training for semi-supervised learning.Advances in Neural Information Processing Systems, 35:32424–32437, 2022

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:6e66f3de59691a15250afa6fd7dd122bf0da63bd7d63cb673b95fe7e38c1c2c9

Observation ec5da4bb-bd31-4da6-89f0-be1b9fc965bf · outbound

This paper cites Supervising strong learners by amplifying weak experts.

Weak-to-Strong Generalization via Direct On-Policy Distillation Supervising strong learners by amplifying weak experts

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:87ecb4e2ad89e217efbcecae0f4b16c26864644295c9a46c9b17a3ddf6c4aeb7

Observation 0fbe11a8-925b-4d49-8365-804d4389dcb3 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Weak-to-Strong Generalization via Direct On-Policy Distillation Scalable agent alignment via reward modeling: a research direction

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0d0360ae82edaee6c87c1452e34e0aad981b230f3ed2c413e93b0901eb336b7c

Observation c784bae3-20a2-41a5-974b-8598dea0a1d0 · outbound

This paper cites AI safety via debate.

Weak-to-Strong Generalization via Direct On-Policy Distillation AI safety via debate

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:911a258ab69d2a44fbb40f27e1e140fba2729aefe2bc97529e8fe2f4b416776b

Observation efc4af27-2ee3-4a29-9127-9092046319d2 · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Measuring Progress on Scalable Oversight for Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:011e4319fa1f59521f13fc4c1fead1edf38184e310518c985d0e4f5a972b5f5a

Observation 0ea91f91-1b4b-4cb1-a103-43bccd9bc4dc · outbound

This paper cites Artificial Sandwiching: When can we test scalable alignment protocols without humans?AI Alignment Forum, 2022.

Weak-to-Strong Generalization via Direct On-Policy Distillation Artificial Sandwiching: When can we test scalable alignment protocols without humans?AI Alignment Forum, 2022

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:32e785c314c82447a40543a68272baa09323869f226c80a94ff9b196476273fd

Observation 99441c2c-ca2f-4586-bf71-7c1011ce5e6f · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Weak-to-Strong Generalization via Direct On-Policy Distillation Constitutional AI: Harmlessness from AI Feedback

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:6d7d3725678d730f50c48653bf45a618a71daa329216f924a8b14f510767bde4

Observation f2a85736-198b-47e0-94f0-3c6478938af2 · outbound

This paper cites Eliciting latent knowledge.

Weak-to-Strong Generalization via Direct On-Policy Distillation Eliciting latent knowledge

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:360da3cf0e20f86defbdc8deee1d36a89fa045f5ed1ee8e2a9dd7b38eed8378f

Observation 5ddb4ac6-1bdc-4ba8-918d-16f06d406da1 · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Weak-to-Strong Generalization via Direct On-Policy Distillation Discovering Latent Knowledge in Language Models Without Supervision

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:2904913bf1a72513733cd570333a4100778d3086611e9996a67731ae90c845fd

Observation d59ddbb5-b5d1-453d-8626-7b68dba12af6 · outbound

This paper cites Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks.

Weak-to-Strong Generalization via Direct On-Policy Distillation Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:b7b715e5afeff7172323ef502c36f58f15264f29409b422ecc71eef5905a813d

Observation e7e164f9-b7ce-4033-aa45-6d8ad297595e · outbound

This paper cites Datasets for Studying Generalization from Easy to Hard Examples.

Weak-to-Strong Generalization via Direct On-Policy Distillation Datasets for Studying Generalization from Easy to Hard Examples

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:50c9a71242bd65ca7de90a8ffcbccfafb61cc70e2eaddb33ef4df5fb43d5cfe6

Observation 1e73740b-de68-4537-b9e3-29f49fd9d5fe · outbound

This paper cites Towards Scalable Automated Alignment of LLMs: A Survey.

Weak-to-Strong Generalization via Direct On-Policy Distillation Towards Scalable Automated Alignment of LLMs: A Survey

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:202796e20f5ce4c2103cc5fa49e54147a0a63b2f9a016fc71d4f386252cbf57b

Observation 0e6647ff-7514-4774-80c3-60e33af2a2fc · outbound

This paper cites Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL.

Weak-to-Strong Generalization via Direct On-Policy Distillation Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:1227c2e861b499b59b52ddf818ed177eb553bc5288bbdcd1a00adc388388e0c9

Observation e7fe6e15-85a8-4a3d-97e2-033207bc1c2d · outbound

This paper cites On Weak-to-Strong Generalization and f-Divergence.

Weak-to-Strong Generalization via Direct On-Policy Distillation On Weak-to-Strong Generalization and f-Divergence

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:16848f362e5d82318e775bce5f71a368e7d1997a223ca3b4f9592cef860935b2

Observation f24e9d31-ee6b-4b32-8a90-579283d1a286 · outbound

This paper cites Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning.

Weak-to-Strong Generalization via Direct On-Policy Distillation Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:0fa7587181b300007ab7b66d6284183161b885e62cb7be52604dc3f9588fb636

Observation 05cc7a6e-dcbc-43d6-84f6-477e05ee8cee · outbound

This paper cites Incentivizing strong reasoning from weak supervision.arXiv preprint arXiv:2505.20072, 2025.

Weak-to-Strong Generalization via Direct On-Policy Distillation Incentivizing strong reasoning from weak supervision.arXiv preprint arXiv:2505.20072, 2025

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:142ef5fc0ffcdc6b390e830918bed6a765050fb0a3902c7602438cf0d638581d

Observation 36f27840-916c-40e0-a746-c4ccd5e21f1b · outbound

This paper cites Self-Rewarding Language Models.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Rewarding Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:2eb2b0ac3a8ab9971f03688d1452b082a49c96b398aef3c5f5392873cd3beace

Observation 30059df4-9f93-4e88-8475-7f9ccd21f731 · outbound

This paper cites Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model.

Weak-to-Strong Generalization via Direct On-Policy Distillation Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:b122d7bc91fb94532fd63bc8d87262f147110364c66c8ca8ca8c7ae7fc206fa5

Observation 6a17ee42-4f29-47d8-9461-8c67c3ee4b4c · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.AISTATS, 2024.

Weak-to-Strong Generalization via Direct On-Policy Distillation A general theoretical paradigm to understand learning from human preferences.AISTATS, 2024

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:5731267fca67f23afe6089ee9cb4459522f0318a45573812b9ea9df39070679d

Observation 0b6b0ef9-469d-4174-ba8c-81a441375b34 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Weak-to-Strong Generalization via Direct On-Policy Distillation KTO: Model Alignment as Prospect Theoretic Optimization

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:90e02531e06841ef041293b0900735153ccfd4387fef0ba6b2f1cc615a2c1450

Observation 7a573a63-9006-4612-9cca-b45b8d2aab0a · outbound

This paper cites SimPO: Simple preference optimization with a reference-free reward.

Weak-to-Strong Generalization via Direct On-Policy Distillation SimPO: Simple preference optimization with a reference-free reward

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:65e195bf026ab05d0fbd5b58b9fdaade552007f9b90e1fec7ce6626cdf60f120

Observation 61393b5f-49b2-49fd-bfbc-66143ca88918 · outbound

This paper cites Let's Verify Step by Step.

Weak-to-Strong Generalization via Direct On-Policy Distillation Let's Verify Step by Step

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:3550c0ec276c11f4cc3cca18ae73919d80ca4962c171e86fe9edc505776758b1

Observation 6efae79e-95ab-44cf-9b7b-66b1143a3634 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Weak-to-Strong Generalization via Direct On-Policy Distillation Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:d59ab829a4086ad69dfbef9f48be13fb803a8867f59931a518abbe0db9f7c3ba

Observation 9130062e-7c6d-431d-b775-576e47f93f5d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Weak-to-Strong Generalization via Direct On-Policy Distillation Process Reinforcement through Implicit Rewards

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:b2e5860cf1d9b8e31be41bec693fc60a500170316c9d7090b1b55a8cd63e1f6c

Observation e68e1632-e745-4550-a34c-fc003f8ded03 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Weak-to-Strong Generalization via Direct On-Policy Distillation VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:78a6ad9e9b5a4cfcf5dd622724b5868c130438049057cdbe04f89f4cbf835643

Pith citing papers

Observation d3bc2ed4-398d-4cbd-b8e9-451f2561bf49 · inbound

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals cites this paper.

Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T05:09:00.865375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:09:00.865375Z digest=sha256:a9d678e22d450ac0848fb75c1089c39b79c410c7a72afbd1f323632db673272a

Observation fc5e68fa-d86f-46ee-8b20-62d69c403240 · inbound

Visual Contrastive Self-Distillation cites this paper.

Visual Contrastive Self-Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:43.400511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:43.400511Z digest=sha256:216a0a745c2feeadb1a10354f77df90de7909bd9bd573323aa324d3b7f8246a2

Observation 9e45157b-2721-467a-bed1-ad066595600d · inbound

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation cites this paper.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:14.911227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:14.911227Z digest=sha256:661ec0e10cdf2da5cd5332340a2babc05414e862163a84c87f1ab90e74f9290c

Observation 2acfd7d2-a822-474e-a4f4-f0410f2bead1 · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Generalization via Direct On-Policy Distillation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.345094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.345094Z digest=sha256:75c31694c5f622bb9c8a2b21803d0e91ddd6243cce20577287ea0ae29d24f1f1