Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-07T16:22:55.342495Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2607.05339.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-07T16:22:55.342495Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80fd619e-8d26-41d4-bcfd-4500c6ca129d · outbound
TREK: Distill to Explore, Reinforce to Refine Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5352be6f-0bc7-4a50-833a-a960d1605d4f · outbound
TREK: Distill to Explore, Reinforce to Refine DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6633503d-8a18-4a48-818e-a0717298c2f7 · outbound
TREK: Distill to Explore, Reinforce to Refine On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ff8e90aa-e930-490e-9926-7509fb356e77 · outbound
TREK: Distill to Explore, Reinforce to Refine Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 074eaec9-0a3d-4dea-b54b-31bc0f77c52d · outbound
TREK: Distill to Explore, Reinforce to Refine PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 751d1dbc-e511-49da-ab79-ee3f3978eab6 · outbound
TREK: Distill to Explore, Reinforce to Refine ReAct: Synergizing Reasoning and Acting in Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c63a4003-8702-4dab-83f6-e875f8d999fe · outbound
TREK: Distill to Explore, Reinforce to Refine Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ce4c47c0-24b1-47a2-92af-e5d57f20c716 · outbound
TREK: Distill to Explore, Reinforce to Refine Qwen3 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 248af4a3-40d6-4158-b03e-eb2f46dacadb · outbound
TREK: Distill to Explore, Reinforce to Refine Qwen2.5 Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a7fba036-a782-4d1e-bef0-4fb3cb71fdeb · outbound
TREK: Distill to Explore, Reinforce to Refine ScienceWorld: Is your Agent Smarter than a 5th Grader?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c0781527-49b7-4801-b333-8456bfad9bdf · outbound
TREK: Distill to Explore, Reinforce to Refine Learn hard problems during rl with reference guided fine-tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f0e7e388-ff58-4425-a764-6195cda5f8ce · outbound
TREK: Distill to Explore, Reinforce to Refine Privileged Information Distillation for Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e044e34-5c57-4bce-8caa-255f67e81ed3 · outbound
TREK: Distill to Explore, Reinforce to Refine Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 68754d8a-47c3-4562-9926-71451d911559 · outbound
TREK: Distill to Explore, Reinforce to Refine Unifying group-relative and self-distillation policy optimization via sample routing
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e522810-a5e4-4478-a6ae-0e83353208e8 · outbound
TREK: Distill to Explore, Reinforce to Refine Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e49b3531-1c2d-4a64-a419-f497a9e0e943 · outbound
TREK: Distill to Explore, Reinforce to Refine Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f84730e4-135e-4d18-ac39-ede5209b6f0a · outbound
TREK: Distill to Explore, Reinforce to Refine Black-box on-policy distillation of large language models.arXiv preprint
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9e19399-1ae3-4466-ab36-2a952f75ea88 · outbound
TREK: Distill to Explore, Reinforce to Refine Self-Refine: Iterative Refinement with Self-Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b249f26c-14a1-4331-a12d-4ee15eaae7d8 · outbound
TREK: Distill to Explore, Reinforce to Refine Let's Verify Step by Step
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5dba23c9-7194-410c-8aa0-94905afb19d5 · outbound
TREK: Distill to Explore, Reinforce to Refine Large Language Models Can Self-Improve
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4edc7a00-cbb0-46c6-bd20-dbf8595f699c · outbound
TREK: Distill to Explore, Reinforce to Refine Explanations from Large Language Models Make Small Reasoners Better
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c88f63a8-01cb-4ef6-a526-ebeaa3148fb1 · outbound
TREK: Distill to Explore, Reinforce to Refine Distilling step-by-step! Outperforming larger language models with less training data and smaller model sizes
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0388e6d-e7ad-4743-8324-16ab7979439e · outbound
TREK: Distill to Explore, Reinforce to Refine Knowledge Distillation with Training Wheels
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 95e4f408-5e2b-441d-b11d-8da9c6048c83 · outbound
TREK: Distill to Explore, Reinforce to Refine Codes: A context-efficient framework for enhancing small language models via domain-specific adaptation and model ensembling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 811c02bb-30bc-4395-9d72-89f511531161 · outbound
TREK: Distill to Explore, Reinforce to Refine DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49cc97da-1f47-49f0-b1ac-383d63a66841 · outbound
TREK: Distill to Explore, Reinforce to Refine Reinforcement-aware Knowledge Distillation for LLM Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e3e0f65-0415-4218-a492-0faf2c693271 · outbound
TREK: Distill to Explore, Reinforce to Refine DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82644a5a-4930-4ea6-9580-ee56a8ce3e35 · outbound
TREK: Distill to Explore, Reinforce to Refine Group Sequence Policy Optimization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd089bc9-d9f3-4a1e-8fc8-79639bc469e3 · outbound
TREK: Distill to Explore, Reinforce to Refine Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.