Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:43.174655Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2505.21067.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:46:43.174655Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:39.414554Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T16:31:09.136154Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5e8cb2c5-5492-4022-b8e7-9b9a882ce3d1 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981c39d1-d827-41f5-a62b-4912f5041cc0 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebdb321a-74b0-43d6-b6f8-4b1d477de23b · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e1b567-0a35-4885-835a-dbfa5b57368f · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 51452794-1b2c-4f67-81d7-0ce7b073abdc · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Gemini 2.5 pro: Our most advanced reasoning model, March 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d286a298-6483-441a-a45e-2b5f44b6188c · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DeepSeek-V3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d034978-e838-4bee-a108-f5280c13d47a · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b3dd73-f9b7-4c0c-abc1-d33bcdacbf4e · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMR: Less is More for RL Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f5ff59-2f3b-4638-b652-9df06b60a614 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2aecdf-48b8-479c-bf29-d366e3f30ebf · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb7b38c-e581-4cf4-b9aa-c9cb8cccb713 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3209e50a-0aff-4c99-8a20-2a8d9c5610d6 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1212dfaf-7f7c-43df-b814-d0ab3ba330f3 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de59cf55-b8b1-41c2-8bc2-ec41781e6e41 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning s1: Simple test-time scaling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23a507f2-0234-4681-ab37-c7177a4073be · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning LIMO: Less is More for Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26f125e-fcc6-4fb3-88fa-2c3a7cc3f445 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwen2.5 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1122bdd9-2470-4da0-a66b-0cc656f55190 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4230d17-5916-46f5-b094-5da04ed6c08e · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d293f0-a5fa-4dab-bb85-22bde31f9a9e · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bda4652-793a-442a-8e04-743df8910793 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d859aa5a-3e98-4474-8094-0e91b0b685cf · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ec0550-b0f2-4b91-80a2-7ae41c75cbd8 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Understanding Aha Moments: from External Observations to Internal Mechanisms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed44878-47bc-437b-9d66-0f0e4318022d · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Bespoke-stratos: The unreasonable effectiveness of reasoning distilla- tion
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7fcb8a4d-1292-4391-ba39-3ac59747cd87 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22eae47c-d1d7-4c11-bd8d-07f167d12658 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dd5662e-4475-4b69-85fc-b0b28ab22321 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2024 part 1, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3adf5f29-287e-4e93-ad9d-aec408064789 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2024 part 2, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e73a6865-d7e6-4612-b193-3e7645b2e4b9 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2025 part 1, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72bfa3ea-b9e4-4c0f-b1e4-685df00245de · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning American invitational mathematics examination 2025 part 2, 2025
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a3c5022-58da-488e-86f6-cfc041ab58cf · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Hmmt february 2025 dataset
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 16ad271f-b2a9-4b39-848b-b808ffa0dd92 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Gpqa: A graduate-level google-proof q&a benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1788c9f-7d98-4d88-87f2-ffa10fc76e3d · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9584e2-2f46-4080-9463-46ac2f3d115a · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.arXiv preprint arXiv:2504.07086, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d106ebfb-aaed-4694-ba55-7bf613dea488 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Evaluating Large Language Models Trained on Code
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b756eddc-63a1-41ea-b5be-19d0af8f9ecc · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4cce4b-9625-43ad-ad17-40c096faeca5 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Assessing metacognitive awareness.Contem- porary educational psychology, 19(4):460–475, 1994
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0b604add-089e-44a5-9f0c-b6f06b4f223c · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning GPT-4o System Card
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88201bb7-28b0-4882-9951-9e52223f5302 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdd44cc-4231-4b58-8921-fd5afaba7506 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Measuring Massive Multitask Language Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8199f5-b5b2-4d5e-be25-5f928a0ce9b4 · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning Qwen-boxed
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 251e146d-c47d-4751-86ff-82cdd30e28ce · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning let’s try another angle
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3056307-a6f4-423f-80cf-ba182fd2e6bc · outbound
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning wait, maybe my approach is wrong here
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6bdb1f8e-e9e3-4861-bf3e-f7ee7c8a74e5 · inbound
rStar2-Agent: Agentic Reasoning Technical Report Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b363ec6-1ca6-4012-bce3-8534fa828932 · inbound
Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.