Pith. sign in

Paper Citation Record · LEDGER

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2510.18383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.18383 v3

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:55:58.828795Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1c81681-92eb-4117-89a7-bbedacdd9a62 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.526698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.526698Z digest=sha256:f3ff0cf41238ca3c61c0d6abfb41a852ff5fe0bb7fb38e16ad4cabd544d9cb7d

Observation 7693d82b-32ce-43c2-bdd1-894d3401f9c5 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.535363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.535363Z digest=sha256:e515b11b1bdc342677d184e545c994b7056fdc25fb4a5d87e2b0d9bbc0005642

Observation f9ba5fdc-0707-4b53-8eb9-388a0ed72549 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.541079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.541079Z digest=sha256:73fb8e5a2f58179739cb4ee18dcdc5b37a7bc6c5be03081a41b73b5c7706f421

Observation ae9ad3e8-a359-403d-b06a-4234b279a55d · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.545648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.545648Z digest=sha256:9c02b66ec2d03ea264b440cb5f783dae69839eb7de55c4382ca7bdb63f1c7b8b

Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.550417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.550417Z digest=sha256:d64b8338ce3c3671622dde79bdfa81cfe3c2f7f8b45fd1ee9bf06df0447d02f6

Observation 51ff5311-ae2a-4a5d-93d0-66f498b69373 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-04T08:58:31.607089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-04T08:55:58.555012Z digest=sha256:0ef45bc8fcf03698d71d21fbfcf016b0a4f0a53408fa36a66f199283d7e92af0

Observation e62b3e33-6608-4856-aca4-e8b3cb3ed20e · outbound

This paper cites Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Beyond Imitation: Learning Key Reasoning Steps from Dual Chain-of-Thoughts in Reasoning Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.560564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.560564Z digest=sha256:70a44c9752a733c9aa899a1c33f560d7ca95dddc82b3b63af2fa88c4d771a96c

Observation b9a364c0-71d5-4bcd-bd89-994a3db68594 · outbound

This paper cites Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Improve Student's Reasoning Generalizability through Cascading Decomposed CoTs Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.565705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.565705Z digest=sha256:78d932f151deb65c294b98b1d9128864655387da645c1c9d7e5200f64b1206ec

Observation ab00aba2-ed70-4162-8aa2-bed354f01d86 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.571387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.571387Z digest=sha256:084044ccd049703f7fde9ab1dfb23c5baf3860f0fc43052549e396b7058a98be

Observation 78c4557a-f11c-4df4-84fd-6f0b44aa6948 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.576083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.576083Z digest=sha256:3c54081e6c6679c81dad2309de4710fda981a8d25e74fe0f8741c78f348c2d02

Observation 06dfca44-1c00-4d8c-8ece-9b7bb9ab5c00 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.581084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.581084Z digest=sha256:6d4868212177fd8808ef4f0777ae8c5d112b90627b59c190183de76e7d50ac9b

Observation 6d0f405e-d7cb-4edf-9bda-94929f10d04a · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.586028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.586028Z digest=sha256:0f6ad3b1ee93b98a54ba0b945fd977fa20239f93cdda8421c6bfbbf76eed99e9

Observation c5f29713-be10-40ec-a81b-63df1a9ee705 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.590626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.590626Z digest=sha256:72c141aa8e9102a82c6f1d15d6b49fb03723ddb07f3f9a075a59b5a6c1866df8

Observation 88eff210-0dae-457f-a4df-1ba9d3f35292 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.594790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.594790Z digest=sha256:a8b0d281ce3cb7fe6feaf13eec18ad40b739ae2558532ddc9517887778ae15d1

Observation 1a1e5459-173e-4b5e-98d4-bdb91eaab672 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.599443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.599443Z digest=sha256:e9a33cf68c14f4e5f148468e98b13fdb21807f2385a73ebccc277ccaf7e7ec0d

Observation 4f65d7d8-ec82-40b2-bc18-165f4449cfb3 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.603979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.603979Z digest=sha256:d2a7b8e8d3cf99ab95c249c9c6490dfd941d37039a7d033fe82ca79978d7d21e

Observation 66f28734-ff82-493c-9852-ada0b102d42f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.608828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.608828Z digest=sha256:5704bbcada4154dbf8ea1bd09fe02dba2a37262302f231c86a5e29baae670afe

Observation 91a79947-b1bb-4788-9a65-e5b761eee3d8 · outbound

This paper cites R-Zero: Self-Evolving Reasoning LLM from Zero Data.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation R-Zero: Self-Evolving Reasoning LLM from Zero Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.613740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.613740Z digest=sha256:e7d1359b45dbee8dc17cc313507528bc6a2b5cbc34dd1a41da45c3a63fa1cdf4

Observation bfb5ef7f-c2b7-41a1-939f-4bee2fcc24c9 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-04T08:58:31.475304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-04T08:55:58.618333Z digest=sha256:89a6749cdf3ea1e7bef3f28f95bb84b07f9adf485c65b584ee2024b769a2cad1

Observation 86b95219-44b1-4b5e-a4a1-f06aa1cd01aa · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.622595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.622595Z digest=sha256:568026ad2344d058d9a0c3fa0dcb4a7a564146fcdb2b1d7ed56a1b0d913d69fe

Observation e12a4255-856c-46ad-9ef7-aa1dfb951425 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.627305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.627305Z digest=sha256:2adf9d074792a06e3826d1120bb0229457e58f6cf786a66329a30f3b1cc7e302

Observation a10d9621-0cec-4a0c-b02f-7fef2fd5fa9e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.631650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.631650Z digest=sha256:e1c8431a96f370fe1cc43ec3a7c1926c4fc6026f38e23a4bbfcca4f67954ebb7

Observation 13e726e7-bd7e-4547-b60a-e73d5bc20a16 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.636359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.636359Z digest=sha256:ec98f329bd5fca8fb506f137f9849e65acac0bae8a9347c961332128849a4a3a

Observation 885a75a0-e6f2-4742-a896-1cbe0fb7836a · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.640766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.640766Z digest=sha256:ac47f4082bff3e2a255b1df9d1aa0338c148bb1a2667be7b590c559a4d933b96

Observation bd25fe70-e380-489c-bccd-b94365c42041 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.645518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.645518Z digest=sha256:500e8c8e8f421033ce5779bd40ce84934dd69066c88f849721942ce526564346

Observation df2e4a58-1d23-4121-9910-792e8d416aca · outbound

This paper cites TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.649949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.649949Z digest=sha256:c4e9547d2979f7073f216ada48be06e830927b91edb31a601fb07a2085f8f518

Observation 35273487-6d54-4493-aa0e-a10989e94339 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Solving Quantitative Reasoning Problems with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.655096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.655096Z digest=sha256:4f50224943ebafbca4e81149131b24676f15b010be88655cc8a2e8412a25fb97

Observation 0887370a-54b9-4fdf-a4ac-577d9376f2f7 · outbound

This paper cites LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.659542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.659542Z digest=sha256:bd8da5c258d25fd9cfeebd0aada73726abaa0932fc1527228c37fff9bd00f1a7

Observation cec7624a-c365-482a-b2d1-4d29979e5543 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.663870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.663870Z digest=sha256:5d165e9369b0cb0012bab09870ad423779a5e2e67dfda2ff9fd826822ce58b18

Observation cded21b1-dc0f-4f9d-811c-f2f35c3dd36d · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.668140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.668140Z digest=sha256:e240eab4dee7a67c051b5b9349b2b5735988505e95b46c9e37d23984c7740747

Observation 5dafb6d8-5432-44d3-a103-d84ee6682cf7 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.672391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.672391Z digest=sha256:076da4f42d5def30b121bd5708bf0abce6536e2739007a3e3529bb62ab1fbdd6

Observation fabb6a06-c309-4768-ad9b-bd27b5184468 · outbound

This paper cites Self-Training Large Language Models for Tool-Use Without Demonstrations.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Self-Training Large Language Models for Tool-Use Without Demonstrations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.676999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.676999Z digest=sha256:9b5bfc919712b5a8681b5295ea152c2f8f651ccf6f063a7ea51c3dd2504d5cdc

Observation ffdab585-71ee-4ecd-a1da-ae45c1f93a2f · outbound

This paper cites Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.681765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.681765Z digest=sha256:22f1142b6f6d3775f9fbf482f6a89b1e51fb0d8be9990b140011e9296b5fd177

Observation 201fa9f0-6e2f-4453-8b47-2cc69a32b1e2 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ART: Automatic multi-step reasoning and tool-use for large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.686891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.686891Z digest=sha256:db52a0e53b4badf1c86916d2c2f0d81f9241f6b5462d70915f51ee4bd037b33f

Observation fd611c45-4fb8-49cd-94ff-9913ed299b3d · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.691802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.691802Z digest=sha256:d1389333b0e2a4ad007dbb88078d08bb974e2b144e2127fe343add243cb87931

Observation 98bcafa2-19ea-44fe-a9b1-9445f531367f · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.695844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.695844Z digest=sha256:d20e6054d3906968800490c7f1a6fb58bd9bff82c54d554a56a05566005299f0

Observation f777e1f4-deb6-4a1b-ac92-a9911390a276 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation ToolRL: Reward is All Tool Learning Needs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.700248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.700248Z digest=sha256:276969ac93b82124664b886893a35168d81b45d2b7499af1d2c9daae15d7db13

Observation 85297bdb-f9cc-4c35-9eba-cd589a144fe9 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 38

Resolution
verified exact
doi, observed 2026-08-04T08:58:31.328348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-04T08:55:58.704653Z digest=sha256:7b4718af8e0c4b73bffab5ad1dd5f6e7246e4b0e2cc2cfed81cdde75d31deb7e

Observation b9cdfc19-79fd-438b-b499-37108b44385d · outbound

This paper cites AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.708980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.708980Z digest=sha256:be35410044601b674efe24e356ff8596f2e029b8ad23f042df84e8af5644c218

Observation 089cc73e-f85b-4b94-bf57-132b84e3bfa3 · outbound

This paper cites Qwen3 Technical Report.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.714056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.714056Z digest=sha256:bb15f3f1079898ccac72766569fdc4b04eedc432987938cc7b6dbef5d1e6b0c8

Observation 47449036-6866-4c3c-a19f-36b622fcfe69 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.718525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.718525Z digest=sha256:bcbda9678569246bede6c387957fae52dd632fd6942d5d8e4feb1926173ba39a

Observation e2ef7d70-0e40-4f26-9116-a9c59a73133d · outbound

This paper cites Proximal Policy Optimization Algorithms.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.723046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.723046Z digest=sha256:52b7a6d5964606e6e703be25e2c71f47bf4bc7d4f128482bcf30106a4db5408c

Observation 75cd815c-e27d-4a06-aaf1-a220052d98e1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.727453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.727453Z digest=sha256:04293443cd2ac6a723c81a7d04cba4410f8f407205840da87745091c3582e6ab

Observation 5a5a4ddc-a842-4a66-a217-915d49281ec4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.731703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.731703Z digest=sha256:b82f6578838ff9b4b63463c6f20a6076d927d730d8159371ab0445fe3e87ac1e

Observation 1322423a-7eab-4316-aea0-00406cad06da · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.736344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.736344Z digest=sha256:5a802c125aeaa5c07717a90c428b25a806520607befb1e881485bee2a458f2dd

Observation dddb4053-77f6-4175-a9d6-ddfc9c3e9225 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.742622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.742622Z digest=sha256:8deef854d6b3ef00b9378a08a3fbdee26d00c272afb405971f39fe48286f09e5

Observation 6c975a41-de88-4417-bf60-f0799217a7b2 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.748285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.748285Z digest=sha256:f1df0d274505e81818c74f201f9d1b6e3d9780905c602e88bd5ced69ff9a4913

Observation 396fc1c9-0ffc-4426-ba37-44953ac2cd76 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.753379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.753379Z digest=sha256:9872f19b28e28d6aecafbb84f627bf6b2f9c5d8c375798a5caed95227ec37345

Observation f0b8dce6-0f88-4b52-9bdf-ea5a78b62e7e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.759048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.759048Z digest=sha256:8da8b193068a02f715b55df14579a02bd7f1ba164eb592577643b3312c5f9e75

Observation 0b3542fc-b4e1-4aac-9db5-679251a413c3 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.764014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.764014Z digest=sha256:95204a50815344e147eb172a14f745ee8a8c094c31efdeedd296397c66c683f1

Observation a7766cf9-99d0-4e62-9840-c8c132c5ee7d · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.769104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.769104Z digest=sha256:1ada79dd9d05bf9ab4639dc90931975abab059ab008fbe70f42954b2c6d2ea2a

Observation 67b20b37-35a5-48b2-8b64-50cd07d52306 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.773353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.773353Z digest=sha256:1b9001e6d2d323ff699fb355f6e79e1947900e599cddb2917c75ad3bf82cecbd

Observation 17c3d065-b494-448f-80f4-4c536249bf21 · outbound

This paper cites Emergent Abilities of Large Language Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Emergent Abilities of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.778667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.778667Z digest=sha256:eeeb0526f7c3ba21ca3b27a1f38e128f5cd3583ac45764538cb14368764726c9

Observation 841bf8cc-6af8-4b60-8016-70f2a88d3ed2 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.783361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.783361Z digest=sha256:d9cbcb5906fc3f6542ebf105408b269ef173787fbde1243b19212602217e5b01

Observation 328dc584-3583-4338-abc1-416eebd8a8e6 · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.787754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.787754Z digest=sha256:b3ebba1dbfce18fd0fb27a2f7106210b1acaf84be81dccfc0afa9a1336dcf17d

Observation e141f2e5-c99d-43d4-acb4-8e21cc1a7426 · outbound

This paper cites Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.792516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.792516Z digest=sha256:cbcd000bdbbd347028040cf850c97d49516adb0803f56b962110f7fca2ce58bb

Observation b751338d-3cd5-4acd-bc02-e6e0f32cdaea · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.797104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.797104Z digest=sha256:2ae73924ed721853e6cccd80ee051ebf72cfc144823d90dc0e51e993fff1adca

Observation 9c90dc37-7a59-47ed-8c5d-92c2af116d4e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.801782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.801782Z digest=sha256:fc84ba15a0826c07934ca3411cb0634fa74bdd87f84ef162d07343a730f3f000

Observation 938722f6-172b-4831-a5cb-e2a032a2be7e · outbound

This paper cites an unresolved cited work.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.806320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.806320Z digest=sha256:7e54da984716e4ef6c69f80350a43306aaca2655e36c2181b94b41a4fe3d0edb

Observation 8482faed-1636-4af1-af6b-96c5074befa2 · outbound

This paper cites Enhancing Generalization in Chain of Thought Reasoning for Smaller Models.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Enhancing Generalization in Chain of Thought Reasoning for Smaller Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.810427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.810427Z digest=sha256:8d8ae8dfa0c4cc57da6fb6d9f71944de6bb65dd73f875591f1de57c46066b8d0

Observation 8f878ced-d3f7-4122-9056-02d59d5cf58d · outbound

This paper cites StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.814994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.814994Z digest=sha256:9fa7bf63fee79548c8ea14a2f0831a2da2edf5cb59ee7ea540100b0ce7c6ce53

Observation 4fc9fc3a-9ae6-4abc-bdbe-ebe7479750bb · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.819366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.819366Z digest=sha256:6745494b099fdd91f4740b6c41fb51b2ac4fb1f243f7aa2dd20c11dcd29f5f42

Observation f2b889f0-a615-4ec7-bbee-c4037a817e9f · outbound

This paper cites online" 'onlinestring :=.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation online" 'onlinestring :=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.823939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.823939Z digest=sha256:a557a5f8f6ef025bb8765ddb92b1d4aa5c4b0b0360b6031e65cf27038f8393f6

Observation 84aa8d46-0ee5-40d7-8d82-4d16aa04f921 · outbound

This paper cites write newline.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation write newline

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.828795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.828795Z digest=sha256:6bdd579d06379111291adc2fd60b30a034b8f0efad4c5956a8de160d86c0457a

Pith citing papers

No inbound Pith citation observations are available.