Pith. sign in

Paper Citation Record · LEDGER

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation

As of 15 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2510.18383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18383 v3

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:55:58.828795Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1c81681-92eb-4117-89a7-bbedacdd9a62 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.526698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.526698Z digest=sha256:614c0c5a74e15bd9aa670ad6244305c20c278fda1103aefc0a6d1514f389456a

Observation 7693d82b-32ce-43c2-bdd1-894d3401f9c5 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.535363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.535363Z digest=sha256:237751712ee88843426aec149c2af422007f9882a406f02e72f687c2e6423150

Observation f9ba5fdc-0707-4b53-8eb9-388a0ed72549 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.541079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.541079Z digest=sha256:2c80b1effbbc265548313f76ae7eaf247565fb0d558230f4e1d58dd1621a0826

Observation ae9ad3e8-a359-403d-b06a-4234b279a55d · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.545648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.545648Z digest=sha256:289da2d3e8ebc215897aafc3b7fe910e9fc4a24414a8ba5572f404338f9d8e5c

Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.550417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.550417Z digest=sha256:e9cf4b296e700fd1eb5ff69ab119b6e5698f35584341f8f67cb075dabda2dc0b

Observation 51ff5311-ae2a-4a5d-93d0-66f498b69373 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-04T08:58:31.607089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T08:55:58.555012Z digest=sha256:07dbbdea5acb95a119397451b197f986637cb09103c38d99378069a4a31fc5cb

Observation e62b3e33-6608-4856-aca4-e8b3cb3ed20e · outbound

This paper cites Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.560564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.560564Z digest=sha256:3bb741d48162f934d72b926ddff4503bdfa1048182d505a72622003e46279596

Observation b9a364c0-71d5-4bcd-bd89-994a3db68594 · outbound

This paper cites Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.565705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.565705Z digest=sha256:2a76d190815cfc3de8099937d289198f7e722f4d8c58f3151970edf0d358e417

Observation ab00aba2-ed70-4162-8aa2-bed354f01d86 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.571387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.571387Z digest=sha256:79d49f2b2fbfb60b8c792b315b4b564668326d1c56f9f8fcbe0f14d65ae7fe25

Observation 78c4557a-f11c-4df4-84fd-6f0b44aa6948 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.576083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.576083Z digest=sha256:6be66d9dde15ba362482b92c9567368c3dc0d52619f091ec392b7c8eee46a2a2

Observation 06dfca44-1c00-4d8c-8ece-9b7bb9ab5c00 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.581084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.581084Z digest=sha256:e74eae14b58219594995e47f583fdb6a841a9df4e630f76b81070e2a35e2d16c

Observation 6d0f405e-d7cb-4edf-9bda-94929f10d04a · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.586028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.586028Z digest=sha256:f25563dc3436c0563a84000de14c3a2deeb2bf20cc97c3b6a26878e0a85ea4b9

Observation c5f29713-be10-40ec-a81b-63df1a9ee705 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.590626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.590626Z digest=sha256:87e140abd631ac22713b26edcaaa2af2b3be839133820dbe4ff9c6796fa060d2

Observation 88eff210-0dae-457f-a4df-1ba9d3f35292 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.594790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.594790Z digest=sha256:e62b7a35f7bf8643fb568995838462482e4c68e02090c817edf5a6494d3ce310

Observation 1a1e5459-173e-4b5e-98d4-bdb91eaab672 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.599443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.599443Z digest=sha256:b65a0882d4e0fc3f6f783eadd1d884583b6e5f87dfcd87ee002e66d199579de3

Observation 4f65d7d8-ec82-40b2-bc18-165f4449cfb3 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.603979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.603979Z digest=sha256:a0d1b42f4df87b57cb63dd340f6080700332d0e936f7e76d8012e586c38360ac

Observation 66f28734-ff82-493c-9852-ada0b102d42f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.608828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.608828Z digest=sha256:d9305563add26780d954204002fdd778cc05879d1e66af638114b75f505faac8

Observation 91a79947-b1bb-4788-9a65-e5b761eee3d8 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.613740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.613740Z digest=sha256:07f528390d1ac553b0323621666b724790993e312360450a9a175a4e64a7da13

Observation bfb5ef7f-c2b7-41a1-939f-4bee2fcc24c9 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-04T08:58:31.475304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T08:55:58.618333Z digest=sha256:3de0ae51956b4337f24bf13b924c146a8bffeff675ea9683fc354894b875560a

Observation 86b95219-44b1-4b5e-a4a1-f06aa1cd01aa · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.622595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.622595Z digest=sha256:2c4070b4d8e18ee8be9137ef4e2fd181d71d15fe43f2c9941cadaa65fdf88635

Observation e12a4255-856c-46ad-9ef7-aa1dfb951425 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.627305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.627305Z digest=sha256:080828edf3427ba2db54ed040ee50d2e330bd05af91fcb5f3db1efb070434fbe

Observation a10d9621-0cec-4a0c-b02f-7fef2fd5fa9e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.631650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.631650Z digest=sha256:8bf67a6a165cc916c9347ffbc88e6b06e4b21da1909d89d5cc7ddd35018c03a3

Observation 13e726e7-bd7e-4547-b60a-e73d5bc20a16 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.636359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.636359Z digest=sha256:eb7652c009d8de4379231ed768c99bb60b65c855a6a649d4e0d81ca10ae9b38c

Observation 885a75a0-e6f2-4742-a896-1cbe0fb7836a · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.640766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.640766Z digest=sha256:a61b8a3c1b711ed1156f03f97286636fba2d465864da0c66a980e41d8ef15192

Observation bd25fe70-e380-489c-bccd-b94365c42041 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.645518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.645518Z digest=sha256:cce690b4c5a6e0095bb69c68370330435391b0ca00a5c84c2b30f7f2e5d3ab5e

Observation df2e4a58-1d23-4121-9910-792e8d416aca · outbound

This paper cites TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.649949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.649949Z digest=sha256:3d6c70edf6bc45d0111b5a763f53d60e149d8f5fb70e8bfeb189ca90ab47055c

Observation 35273487-6d54-4493-aa0e-a10989e94339 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Solving Quantitative Reasoning Problems with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.655096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.655096Z digest=sha256:7858ddaf5b3901843054aa2e5a67b80a67c3d92806e31d2baa12266d02c2cfe4

Observation 0887370a-54b9-4fdf-a4ac-577d9376f2f7 · outbound

This paper cites LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.659542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.659542Z digest=sha256:e98c244195d352d91fcf1f06ea4ee354200fcd0e4af221e9bdedb844ed274b97

Observation cec7624a-c365-482a-b2d1-4d29979e5543 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.663870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.663870Z digest=sha256:fa9275a4ddfd2d8c75bf685ce64118ebd7565184c7acd1001fa7ba6167b95476

Observation cded21b1-dc0f-4f9d-811c-f2f35c3dd36d · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.668140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.668140Z digest=sha256:4457e5c79d70179c4c1398f01091f938bcaf29ebc419dcfb9070311b05d1b758

Observation 5dafb6d8-5432-44d3-a103-d84ee6682cf7 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.672391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.672391Z digest=sha256:54bb919d70a43bdea967e134d47024874352fee29af7b33d98ecd65c44d51e2c

Observation fabb6a06-c309-4768-ad9b-bd27b5184468 · outbound

This paper cites Self-Training Large Language Models for Tool-Use Without Demonstrations.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Self-Training Large Language Models for Tool-Use Without Demonstrations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.676999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.676999Z digest=sha256:fd27d5fd01aee10c21f2594ff0637983b174bfff4a4bcc84a367e6f32a0e380d

Observation ffdab585-71ee-4ecd-a1da-ae45c1f93a2f · outbound

This paper cites Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.681765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.681765Z digest=sha256:1a0b2a1fe7a2f224136c77dcb05cdc9c7a775387198daed0e71f3341a8a104c2

Observation 201fa9f0-6e2f-4453-8b47-2cc69a32b1e2 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ART: Automatic multi-step reasoning and tool-use for large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.686891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.686891Z digest=sha256:a880eba8051877003205369b67c7ed81323c844f4b96a3bf07532330c1c92632

Observation fd611c45-4fb8-49cd-94ff-9913ed299b3d · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.691802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.691802Z digest=sha256:518540eafe34ce7a6e62a19a06260d27b4c65cddf31dd2896ae2c0cb52520716

Observation 98bcafa2-19ea-44fe-a9b1-9445f531367f · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.695844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.695844Z digest=sha256:5290e13208d3ffb3b99cdf25318d7ded9915ec6367bae08d910a6620d364b5ff

Observation f777e1f4-deb6-4a1b-ac92-a9911390a276 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ToolRL: Reward is All Tool Learning Needs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.700248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.700248Z digest=sha256:43ef48f00f81813861496e9edbf132197f266920981ec21ba353a05216038243

Observation 85297bdb-f9cc-4c35-9eba-cd589a144fe9 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 38

Resolution
verified exact
doi, observed 2026-08-04T08:58:31.328348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-04T08:55:58.704653Z digest=sha256:bc8bb87487420fb07bb6f2bae6485c17839ea26d85a492fac0961a48b42075ac

Observation b9cdfc19-79fd-438b-b499-37108b44385d · outbound

This paper cites AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.708980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.708980Z digest=sha256:f6ef78437344e28f266756ea2046e38aab405a2be4dde8d364670ea9f99aef11

Observation 089cc73e-f85b-4b94-bf57-132b84e3bfa3 · outbound

This paper cites Qwen3 Technical Report.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.714056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.714056Z digest=sha256:edf1c072b887bd6b233348d8e6a40e0e6fa43bb26747505eea77948703af8a77

Observation 47449036-6866-4c3c-a19f-36b622fcfe69 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.718525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.718525Z digest=sha256:df2707e435fbc68cdea472fca85be36ddcd0de653fe4b909c1f740269ef90278

Observation e2ef7d70-0e40-4f26-9116-a9c59a73133d · outbound

This paper cites Proximal Policy Optimization Algorithms.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.723046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.723046Z digest=sha256:6b05f5c1c0481b5e9f968b53eb7e25a7cba42269d06bfda323790dcfc70c249f

Observation 75cd815c-e27d-4a06-aaf1-a220052d98e1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.727453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.727453Z digest=sha256:dd53684d69b882c6c1bb2387a6d1f4409e11f61bedec6ec38bed212c678c265a

Observation 5a5a4ddc-a842-4a66-a217-915d49281ec4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.731703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.731703Z digest=sha256:b22b804c0978b2be87a03e1e63d178898142db05266161a6e46b5481407344b3

Observation 1322423a-7eab-4316-aea0-00406cad06da · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.736344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.736344Z digest=sha256:0475c8718fd5c81a3d7c0b6801cc1afb0e496bd1c3b529a96605917d014f4107

Observation dddb4053-77f6-4175-a9d6-ddfc9c3e9225 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.742622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.742622Z digest=sha256:1d0cb19757107934093405ebf650071d25086df0507a41bdea9437fb3c412df3

Observation 6c975a41-de88-4417-bf60-f0799217a7b2 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.748285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.748285Z digest=sha256:3ce26ead492ff1d1a26fdb7e95e34d02a19cb29e5eb424a794836a29d7d2330f

Observation 396fc1c9-0ffc-4426-ba37-44953ac2cd76 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.753379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.753379Z digest=sha256:974a4e1380166b2244c896533c2b580ced5ecf0935f301b2cb3e74464bf5c65a

Observation f0b8dce6-0f88-4b52-9bdf-ea5a78b62e7e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.759048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.759048Z digest=sha256:53358a101a9a730679a82c672f20ea7defd8b7aad9fa983c97010ef4ef7290b4

Observation 0b3542fc-b4e1-4aac-9db5-679251a413c3 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.764014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.764014Z digest=sha256:fece026f0a99f56449483e3f0419c7634f1cec3673f27c7d9c86bae5983ca584

Observation a7766cf9-99d0-4e62-9840-c8c132c5ee7d · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.769104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.769104Z digest=sha256:6d76eab03ec6ad727193f1947c9eaf21d27b863ecec949ff996ca30659c6af91

Observation 67b20b37-35a5-48b2-8b64-50cd07d52306 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.773353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.773353Z digest=sha256:6f7f96bc3721dcfd77a2c6ce3646ba4e1530d1cb511391c24c582efbf3c1f9c0

Observation 17c3d065-b494-448f-80f4-4c536249bf21 · outbound

This paper cites Emergent Abilities of Large Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Emergent Abilities of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.778667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.778667Z digest=sha256:9f8cdf7d92c0b74a1fdc0f6989e8f999c09758b499a731be5cb6d9b78b98e820

Observation 841bf8cc-6af8-4b60-8016-70f2a88d3ed2 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.783361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.783361Z digest=sha256:ceaaac905d0105d480bdd34f4df5ab0559299480d295ec451e90b765bca214fb

Observation 328dc584-3583-4338-abc1-416eebd8a8e6 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.787754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.787754Z digest=sha256:ff80b45f428b2addd56d1d8dfa8bda5909dbf2f91d12851f8dbf3eb372e069b9

Observation e141f2e5-c99d-43d4-acb4-8e21cc1a7426 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.792516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.792516Z digest=sha256:680023430544e5b9b27eeb2ed773618e36285e210c6ba172f7b21efe39d8d12c

Observation b751338d-3cd5-4acd-bc02-e6e0f32cdaea · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.797104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.797104Z digest=sha256:cc36ea14c5aa0fd5d443f0727b3e538195ddeb451cbfa39bc893dbfdf66d8701

Observation 9c90dc37-7a59-47ed-8c5d-92c2af116d4e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.801782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.801782Z digest=sha256:c37d82f29b20e4fd3b3bc7e90466020b1d1cc0eca76ec483ef2e83daf8be2129

Observation 938722f6-172b-4831-a5cb-e2a032a2be7e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.806320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.806320Z digest=sha256:70aab018c334c3f47cd8ee19d4bea66e54a8ac09613ba06d41f2e231501b8d13

Observation 8482faed-1636-4af1-af6b-96c5074befa2 · outbound

This paper cites Enhancing Generalization in Chain of Thought Reasoning for Smaller Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Enhancing Generalization in Chain of Thought Reasoning for Smaller Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.810427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.810427Z digest=sha256:6378b947ca733c8c1c30b3800853fabf489c0c5a65f5b7e678d9a860f95f2f00

Observation 8f878ced-d3f7-4122-9056-02d59d5cf58d · outbound

This paper cites StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.814994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.814994Z digest=sha256:66cebe2ba9b310e1803a7425654be4a2cd20f4ebfe6f51fc7badc9a10c8eb672

Observation 4fc9fc3a-9ae6-4abc-bdbe-ebe7479750bb · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.819366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.819366Z digest=sha256:f714bb3d2dd1258e3af931cdea0e4bc062fad99321d2ef879e87e7c5cef85b44

Observation f2b889f0-a615-4ec7-bbee-c4037a817e9f · outbound

This paper cites online" 'onlinestring :=.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation online" 'onlinestring :=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.823939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.823939Z digest=sha256:008e5dd2fca5078ed21b3b61af5cc475abb8ed0551a796fecea7e5fe8a8da6dc

Observation 84aa8d46-0ee5-40d7-8d82-4d16aa04f921 · outbound

This paper cites write newline.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation write newline

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.828795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.828795Z digest=sha256:fc515ebec1e37d8680f132a7aff6000a79e7feaa4b77ea9781766dec94efbd46

Pith citing papers

No inbound Pith citation observations are available.