Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:55:58.828795Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2510.18383.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:55:58.828795Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d1c81681-92eb-4117-89a7-bbedacdd9a62 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7693d82b-32ce-43c2-bdd1-894d3401f9c5 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ba5fdc-0707-4b53-8eb9-388a0ed72549 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9ad3e8-a359-403d-b06a-4234b279a55d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ff5311-ae2a-4a5d-93d0-66f498b69373 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e62b3e33-6608-4856-aca4-e8b3cb3ed20e · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a364c0-71d5-4bcd-bd89-994a3db68594 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab00aba2-ed70-4162-8aa2-bed354f01d86 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78c4557a-f11c-4df4-84fd-6f0b44aa6948 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06dfca44-1c00-4d8c-8ece-9b7bb9ab5c00 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d0f405e-d7cb-4edf-9bda-94929f10d04a · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f29713-be10-40ec-a81b-63df1a9ee705 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88eff210-0dae-457f-a4df-1ba9d3f35292 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1e5459-173e-4b5e-98d4-bdb91eaab672 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f65d7d8-ec82-40b2-bc18-165f4449cfb3 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f28734-ff82-493c-9852-ada0b102d42f · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation LoRA: Low-Rank Adaptation of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a79947-b1bb-4788-9a65-e5b761eee3d8 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb5ef7f-c2b7-41a1-939f-4bee2fcc24c9 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86b95219-44b1-4b5e-a4a1-f06aa1cd01aa · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12a4255-856c-46ad-9ef7-aa1dfb951425 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a10d9621-0cec-4a0c-b02f-7fef2fd5fa9e · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e726e7-bd7e-4547-b60a-e73d5bc20a16 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885a75a0-e6f2-4742-a896-1cbe0fb7836a · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd25fe70-e380-489c-bccd-b94365c42041 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2e4a58-1d23-4121-9910-792e8d416aca · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35273487-6d54-4493-aa0e-a10989e94339 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Solving Quantitative Reasoning Problems with Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0887370a-54b9-4fdf-a4ac-577d9376f2f7 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec7624a-c365-482a-b2d1-4d29979e5543 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cded21b1-dc0f-4f9d-811c-f2f35c3dd36d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dafb6d8-5432-44d3-a103-d84ee6682cf7 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fabb6a06-c309-4768-ad9b-bd27b5184468 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Self-Training Large Language Models for Tool-Use Without Demonstrations
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffdab585-71ee-4ecd-a1da-ae45c1f93a2f · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201fa9f0-6e2f-4453-8b47-2cc69a32b1e2 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ART: Automatic multi-step reasoning and tool-use for large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd611c45-4fb8-49cd-94ff-9913ed299b3d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bcafa2-19ea-44fe-a9b1-9445f531367f · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f777e1f4-deb6-4a1b-ac92-a9911390a276 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ToolRL: Reward is All Tool Learning Needs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85297bdb-f9cc-4c35-9eba-cd589a144fe9 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b9cdfc19-79fd-438b-b499-37108b44385d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089cc73e-f85b-4b94-bf57-132b84e3bfa3 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Qwen3 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47449036-6866-4c3c-a19f-36b622fcfe69 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2ef7d70-0e40-4f26-9116-a9c59a73133d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Proximal Policy Optimization Algorithms
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75cd815c-e27d-4a06-aaf1-a220052d98e1 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5a4ddc-a842-4a66-a217-915d49281ec4 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation HybridFlow: A Flexible and Efficient RLHF Framework
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1322423a-7eab-4316-aea0-00406cad06da · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dddb4053-77f6-4175-a9d6-ddfc9c3e9225 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c975a41-de88-4417-bf60-f0799217a7b2 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396fc1c9-0ffc-4426-ba37-44953ac2cd76 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b8dce6-0f88-4b52-9bdf-ea5a78b62e7e · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3542fc-b4e1-4aac-9db5-679251a413c3 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7766cf9-99d0-4e62-9840-c8c132c5ee7d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b20b37-35a5-48b2-8b64-50cd07d52306 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17c3d065-b494-448f-80f4-4c536249bf21 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Emergent Abilities of Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 841bf8cc-6af8-4b60-8016-70f2a88d3ed2 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328dc584-3583-4338-abc1-416eebd8a8e6 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e141f2e5-c99d-43d4-acb4-8e21cc1a7426 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b751338d-3cd5-4acd-bc02-e6e0f32cdaea · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c90dc37-7a59-47ed-8c5d-92c2af116d4e · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938722f6-172b-4831-a5cb-e2a032a2be7e · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8482faed-1636-4af1-af6b-96c5074befa2 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f878ced-d3f7-4122-9056-02d59d5cf58d · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc9fc3a-9ae6-4abc-bdbe-ebe7479750bb · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b889f0-a615-4ec7-bbee-c4037a817e9f · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation online" 'onlinestring :=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84aa8d46-0ee5-40d7-8d82-4d16aa04f921 · outbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation write newline
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.