Pith. sign in

Paper Citation Record · LEDGER

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning

As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2505.21067.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21067 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:43.174655Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:39.414554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:31:09.136154Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e8cb2c5-5492-4022-b8e7-9b9a882ce3d1 · outbound

This paper cites OpenAI o1 System Card.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:39.894509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:39.894509Z digest=sha256:51f276c896603640b98065e7d8a1f29145c48cffa1c868a9fd7559c3afe1692a

Observation 981c39d1-d827-41f5-a62b-4912f5041cc0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.046652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.046652Z digest=sha256:52cdc4cdc8e6fb7578d74bdf4d6383afe353d450fc8a794ee08f73445c19c790

Observation ebdb321a-74b0-43d6-b6f8-4b1d477de23b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.168204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.168204Z digest=sha256:4a0cb98e1305c9b7cb4bfe11e13d1313ae849b04ae93023d69db51f5770b4431

Observation 77e1b567-0a35-4885-835a-dbfa5b57368f · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:45.826113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:40.172117Z digest=sha256:e4fb7b7812719ef3a8fb6c2db273bafe94dad25fdea1b6aeab2b3ce5c5bb80c7

Observation 51452794-1b2c-4f67-81d7-0ce7b073abdc · outbound

This paper cites Gemini 2.5 pro: Our most advanced reasoning model, March 2025.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Gemini 2.5 pro: Our most advanced reasoning model, March 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:45.637961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:40.194428Z digest=sha256:4d9039e77612d41be13a0120c99dfc95696256582f2a471dc73946d40c528670

Observation d286a298-6483-441a-a45e-2b5f44b6188c · outbound

This paper cites DeepSeek-V3 Technical Report.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.242659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.242659Z digest=sha256:1d585c4a111adefbf29ff155b057ecf45ea1763c44901493e0c741cdb746743b

Observation 7d034978-e838-4bee-a108-f5280c13d47a · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.291885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.291885Z digest=sha256:656978d93ca82575f9cd62978885f86fb2f8c77922fdd5803957efaec083be6a

Observation 12b3dd73-f9b7-4c0c-abc1-d33bcdacbf4e · outbound

This paper cites LIMR: Less is More for RL Scaling.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMR: Less is More for RL Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.358276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.358276Z digest=sha256:9051a91ee3f0a34917cc198658a58cb54789633b9c2ce8179bad0a3e9f26c1af

Observation b5f5ff59-2f3b-4638-b652-9df06b60a614 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.382950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.382950Z digest=sha256:4a9e4b0b8b1c7da6c19fe5d9ed8a500c8dccf8d8e281c744ad9533f124b3e2cf

Observation ad2aecdf-48b8-479c-bf29-d366e3f30ebf · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.408551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.408551Z digest=sha256:18211eb06d980fc457d8df54fc11e6157184179867e907c08b3e543e360ef0ba

Observation fdb7b38c-e581-4cf4-b9aa-c9cb8cccb713 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.444115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.444115Z digest=sha256:22f55945a45ba9651b44637da5a2f41e7a51933971d8020669e67fcefc2e0847

Observation 3209e50a-0aff-4c99-8a20-2a8d9c5610d6 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.447364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.447364Z digest=sha256:58c1617a51dba3f9ccb893c6e7d942b6b362cbc84e7d8300adccf480fac553ad

Observation 1212dfaf-7f7c-43df-b814-d0ab3ba330f3 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.524734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.524734Z digest=sha256:f9f2666c01185167b142fb663dcf7a259098b5e2addeb482ef529704ea8cf585

Observation de59cf55-b8b1-41c2-8bc2-ec41781e6e41 · outbound

This paper cites s1: Simple test-time scaling.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning s1: Simple test-time scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.632602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.632602Z digest=sha256:20cd34a34b613b679b6a6cf878f62325daa794e1d87b3ade6845e3d4148d7133

Observation 23a507f2-0234-4681-ab37-c7177a4073be · outbound

This paper cites LIMO: Less is More for Reasoning.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMO: Less is More for Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.704947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.704947Z digest=sha256:e498c8cc4e9570225d9391de194b0955bdf566fd89775f8b513c16a71ade1c01

Observation a26f125e-fcc6-4fb3-88fa-2c3a7cc3f445 · outbound

This paper cites Qwen2.5 Technical Report.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.781652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.781652Z digest=sha256:679272b8a89d85edbfe50b75ec5c7949bf6505c3e7a679b81d09aa3f40f8bc3f

Observation 1122bdd9-2470-4da0-a66b-0cc656f55190 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:40.944799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:40.944799Z digest=sha256:b61c1892bbcc5b1603b0742b96efd5fb4e732ab6cdfbd9ade4820e792f22b951

Observation a4230d17-5916-46f5-b094-5da04ed6c08e · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.098216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.098216Z digest=sha256:0f50e701b25f7176b2cd162764e2074ca2c210719ceebe55dcb86076ae700992

Observation c5d293f0-a5fa-4dab-bb85-22bde31f9a9e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.230911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.230911Z digest=sha256:597f0306108b5256d04cf013712e74c12a91d6be7f7a3b4cbb19603cd4ce2a75

Observation 6bda4652-793a-442a-8e04-743df8910793 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.410507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.410507Z digest=sha256:9dcbea394ca2df53ff40e574d3642adef380737db5412368684b0489d2a39394

Observation d859aa5a-3e98-4474-8094-0e91b0b685cf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.480449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.480449Z digest=sha256:6d9777fb6fca4f378fad6fa63f18a84e09740147577a3e0cec6dceb8d9366dea

Observation b9ec0550-b0f2-4b91-80a2-7ae41c75cbd8 · outbound

This paper cites Understanding Aha Moments: from External Observations to Internal Mechanisms.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Understanding Aha Moments: from External Observations to Internal Mechanisms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.603694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.603694Z digest=sha256:0a9d3fef87a147a54df59debd8fc3c5b993869a2b1be1f884cb8b739ddc02764

Observation 6ed44878-47bc-437b-9d66-0f0e4318022d · outbound

This paper cites Bespoke-stratos: The unreasonable effectiveness of reasoning distilla- tion.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Bespoke-stratos: The unreasonable effectiveness of reasoning distilla- tion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:45.479474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:41.721544Z digest=sha256:08c038ef6d4867b0a4c307eee08aaa02e990a302ef85f63139c02499ec027809

Observation 7fcb8a4d-1292-4391-ba39-3ac59747cd87 · outbound

This paper cites RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.856018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.856018Z digest=sha256:cc058d4d7be9f5cf864ad985ca55f327f9fddcf5814f52e7a0698b08804b49f3

Observation 22eae47c-d1d7-4c11-bd8d-07f167d12658 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:41.916225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:41.916225Z digest=sha256:b14f8d68ee2b0659bacda8b41cc448988fd8374c1e22639afd2bc17779174231

Observation 5dd5662e-4475-4b69-85fc-b0b28ab22321 · outbound

This paper cites American invitational mathematics examination 2024 part 1, 2024.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2024 part 1, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:45.262072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:41.980985Z digest=sha256:ef8e0328aa8af0c1f2457470935089387ad233ea44316ac484054ead0308918e

Observation 3adf5f29-287e-4e93-ad9d-aec408064789 · outbound

This paper cites American invitational mathematics examination 2024 part 2, 2024.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2024 part 2, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:45.065868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:42.054101Z digest=sha256:0381e4e876b517a92d325609e2163ebdb7733682d9a18fe1e8378c5b80a28879

Observation e73a6865-d7e6-4612-b193-3e7645b2e4b9 · outbound

This paper cites American invitational mathematics examination 2025 part 1, 2025.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2025 part 1, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:44.871083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:42.125387Z digest=sha256:7a63c1c437de93000a9dcd395282b9e5fa32c945423f33d6f5432f0ebdd51596

Observation 72bfa3ea-b9e4-4c0f-b1e4-685df00245de · outbound

This paper cites American invitational mathematics examination 2025 part 2, 2025.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2025 part 2, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:44.692831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:42.196712Z digest=sha256:d5f5693c147d3f3f394293d0c2fc02a685d5bba67213309704215ce36b89bd88

Observation 6a3c5022-58da-488e-86f6-cfc041ab58cf · outbound

This paper cites Hmmt february 2025 dataset.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Hmmt february 2025 dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:44.484512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:42.256987Z digest=sha256:6fad1147e0784e6b60f959842d00f6bb538fd32a802c8695da6820859eee0251

Observation 16ad271f-b2a9-4b39-848b-b808ffa0dd92 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Gpqa: A graduate-level google-proof q&a benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.314347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.314347Z digest=sha256:5ea97495d6dc573466ec4a73e520edbc6142890bb2683f37fe8bf5dcc160fad8

Observation f1788c9f-7d98-4d88-87f2-ffa10fc76e3d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.404452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.404452Z digest=sha256:9175730f5b97cb1ffca5ff3bf8755f079f18ed494b0c0d61df560fe1bb39cde6

Observation 4e9584e2-2f46-4080-9463-46ac2f3d115a · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.arXiv preprint arXiv:2504.07086, 2025.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.arXiv preprint arXiv:2504.07086, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.476673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.476673Z digest=sha256:3464a191b87ee2c9fc122a77f59ed80858b3f0acbed8e4ac593581d1b6d75ba2

Observation d106ebfb-aaed-4694-ba55-7bf613dea488 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Evaluating Large Language Models Trained on Code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.546131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.546131Z digest=sha256:5cfeb74c0e24e9a23573179d29943b0bf8bdb91ba18dac7745628d9505505fe5

Observation b756eddc-63a1-41ea-b5be-19d0af8f9ecc · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.641108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.641108Z digest=sha256:791298e9ca48002eefa1eb06000b27783898ba4d34c02d9e8ae8ac5932d076c1

Observation 6c4cce4b-9625-43ad-ad17-40c096faeca5 · outbound

This paper cites Assessing metacognitive awareness.Contem- porary educational psychology, 19(4):460–475, 1994.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Assessing metacognitive awareness.Contem- porary educational psychology, 19(4):460–475, 1994

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:44.322504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:42.715586Z digest=sha256:359585c957eb45be7fbd1289cc9331502bb711a3e6445db9decb0ff66706d855

Observation 0b604add-089e-44a5-9f0c-b6f06b4f223c · outbound

This paper cites GPT-4o System Card.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning GPT-4o System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.815990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.815990Z digest=sha256:21bccf6caa1f71768347de2d703ccf816c2b705bf70b09c231928af3870cbb06

Observation 88201bb7-28b0-4882-9951-9e52223f5302 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.891459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.891459Z digest=sha256:a23442405a3a6c9e69deca4d281589ce2f4443be431f70adbfc8aba9f6444bb4

Observation 2fdd44cc-4231-4b58-8921-fd5afaba7506 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Measuring Massive Multitask Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:42.990726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:42.990726Z digest=sha256:9d493c557af09b043edc2c018487035f8d815fecd53bdbfd19b9944d59434e3a

Observation 1b8199f5-b5b2-4d5e-be25-5f928a0ce9b4 · outbound

This paper cites Qwen-boxed.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwen-boxed

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:44.149742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:43.065912Z digest=sha256:8796a930927b7d1f1789d6be6bc18371dcee88567abe01af45b059b3713835a3

Observation 251e146d-c47d-4751-86ff-82cdd30e28ce · outbound

This paper cites let’s try another angle.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning let’s try another angle

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:43.925564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:43.140924Z digest=sha256:6e879bf5a6657b25d5f90f7d600ecc2357869417c3cb22d7cee373f1979d7825

Observation a3056307-a6f4-423f-80cf-ba182fd2e6bc · outbound

This paper cites wait, maybe my approach is wrong here.

Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning wait, maybe my approach is wrong here

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:46:43.710362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:46:43.174655Z digest=sha256:2f82ab46079751a9083a773926684738fc457508d35158acb2c250d779cd72be

Pith citing papers

Observation 6bdb1f8e-e9e3-4861-bf3e-f7ee7c8a74e5 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.414554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.414554Z digest=sha256:3a602b4f64efd5c47d7eb6d47b2968820751afe989d5b8d2fddf923e7122ef08

Observation 4b363ec6-1ca6-4012-bce3-8534fa828932 · inbound

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding cites this paper.

Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:09.142870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T16:29:05.186607Z digest=sha256:1c3498d6aad3d6dc2d41567c67a2c50c4173d9f9272475f2e71269c2dafe9a03