Pith. sign in

Paper Citation Record · LEDGER

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 38 inbound Pith citation observations for arXiv:2505.16410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16410 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:08.486342Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:26:56.575865Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a768c1a7-f587-45d9-afb8-9244aff5719c · outbound

This paper cites Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:02.921197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:02.921197Z digest=sha256:d685c3a5533fa1925f30a863ab56e20e926b1824b9eb880d578de01af27ffa23

Observation 9a52689b-1077-4ebc-ab31-0003e6de246c · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:02.965213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:02.965213Z digest=sha256:c18cfce1360dc5b37ab2f30bcedebb5d2568a55c7c8627047c571d4c7acdcfa9

Observation c5f52244-f3d1-4478-92df-621630b3f6d6 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.057637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.057637Z digest=sha256:887bd9cc334609f86cb7344664a147a9fad5eb8adcecbccae92d5df395ab423a

Observation 08d50b1f-5230-48c6-9bf5-b7e9303ae65f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.258543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.258543Z digest=sha256:635f66d77976adc33324225c81bb09ce5aa89050360e3f7a2867152bcd41602d

Observation 6e0db75b-ba54-40fd-a9c3-04d5116bb108 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.338887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.338887Z digest=sha256:34e1c3c251366d6abf6a19bb200087a92890998e58a21e7e69ed8e4634dd0bbc

Observation 3f2a9b67-4572-412c-9ed6-5b8da5fb3627 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.410193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.410193Z digest=sha256:fd4a1999c8ba13a2b1dd5386dbeaefb0d7d80c12fa7d6cd1d2530f118b6ef44e

Observation 9db8e6d8-9076-4853-be05-13d988fca364 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.377945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:03.487236Z digest=sha256:961082cc73acf6c1f9a54f37371b6217b04696f1074fdd44a2960297333627d5

Observation 8a01c378-6eea-464c-945e-8b952fd37e6d · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.575092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.575092Z digest=sha256:99ab95bf076847c20c2966eb9523f05cb607cf53ac9f19f28d70b15d76dfae7d

Observation 463ec17e-ac30-4bdb-8ea2-81a9d406ba19 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.644180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.644180Z digest=sha256:b1ea941d8f9ddffa608a5ceccf2cd58ec34a07ac2086b8c5376234937a9f3d9e

Observation 59afab32-de7d-400d-9f4e-bda9cdd99418 · outbound

This paper cites How abilities in large language models are affected by supervised fine-tuning data composition.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning How abilities in large language models are affected by supervised fine-tuning data composition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.722976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.722976Z digest=sha256:9526d28338ab3eb1a041163b5a83a13085e754df6a714262931d6dcb8c346e6c

Observation 534947a6-a771-46e7-89c1-5c0506bb8b8a · outbound

This paper cites Progressive Multimodal Reasoning via Active Retrieval.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Progressive Multimodal Reasoning via Active Retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.782011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.782011Z digest=sha256:26b7db99e8368938f0b570ed1910fc358f2ee5787370cbbe923c470f79a375b1

Observation 05b45d04-ff58-4c3c-9d94-db92a37592e5 · outbound

This paper cites Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.853422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.853422Z digest=sha256:ae5c04b5c4ca683e6b91828cb5c02775d20c7ce0f9c2b8f3b9be338669969496

Observation 3a8ccf65-cd6a-4971-93df-8ea0870efad9 · outbound

This paper cites The Llama 3 Herd of Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.926256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.926256Z digest=sha256:c6b06f07c394ecacc655394aaf2deb984f54a6530d8a2c3b3b3cc07447bfed1d

Observation 913923bf-d874-4927-8775-ad0923d608cb · outbound

This paper cites Concise reasoning via reinforcement learning, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Concise reasoning via reinforcement learning, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.356157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:03.998962Z digest=sha256:779f22cb35e28f1f9ced6fca68b3315f170a10d25612cc930feffc3b8dbe0afa

Observation f7aea587-87f7-4ccd-9cd7-cf89a904e6da · outbound

This paper cites Retool: Reinforcement learning for strategic tool use in llms, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retool: Reinforcement learning for strategic tool use in llms, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.079207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.079207Z digest=sha256:b5f37d5c2b51a65395cf356d0f79be213787c7a2a630978f320e67117b2fc9bf

Observation 23cafa84-7db2-4b24-bf89-2225134fdcf0 · outbound

This paper cites Tora: A tool-integrated reasoning agent for mathematical problem solving.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Tora: A tool-integrated reasoning agent for mathematical problem solving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.339342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:04.133757Z digest=sha256:1ecac1594fa626b59cb923531960c8ea247103dba2c23b1563bc2f6d73f81c7f

Observation f0ec14f4-adee-4938-ae80-09f21d6f692b · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Measuring mathematical problem solving with the MATH dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.244504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.244504Z digest=sha256:57d9e0b3eaa7356348b44d73fc178e893b0402f1683944f6ae67efd4071f3bb3

Observation 8458cacf-0f84-45b9-8de3-68df07e66012 · outbound

This paper cites Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.304466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.304466Z digest=sha256:36208069c607c7b494d6c03d0fbca79bd0e1c69119e92b711c3516cc8de1ac1d

Observation f8588b63-1591-4eb6-bc10-acab3f7332ea · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.343006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.343006Z digest=sha256:d762b91e5003121b6f1f8aeeddd2fb293589c4c6a9e16da6d619f36676e108be

Observation e9160051-d954-4e11-b638-b95197f87cbb · outbound

This paper cites Towards reasoning in large language models: A survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Towards reasoning in large language models: A survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.459705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.459705Z digest=sha256:545620587edb4f6aa4086308f2570d1d4a2dfc9402942b90ae8695295b6bc78b

Observation 3cddfc57-598a-4628-8a98-68a78fd6e4d9 · outbound

This paper cites RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.555452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.555452Z digest=sha256:087ec1866e0ec5cf18f488f2c77179d86a74a3e1bf5afd2e5dc56c459cea5802

Observation bffe5425-3732-4e3f-9c2d-fde963dae4b9 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.638038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.638038Z digest=sha256:2b5ed87afd95e2c8a7e6594c36f60d84563821643c5c715025ac1df80ed04478

Observation 83ed383c-4019-413e-a777-e91fd6877bd9 · outbound

This paper cites FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.709451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.709451Z digest=sha256:1813bffc077d67d08048781f7dd34bf7010ad64348de5cddd4c097b02883d3bc

Observation 200f13a6-785c-47f8-b91a-5eff891f837a · outbound

This paper cites InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.761722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.761722Z digest=sha256:01f25de1c8718ffd52812e6dcd7a24f17bee37e405a2eaec33a0f3afcdd5ff5f

Observation 8416c664-e224-46e3-a43d-0bdebc16618d · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.311394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:04.856927Z digest=sha256:2abef9d8218a2ac9320cceec5da1c25b819dca32b08f2e379bbfda01874c5ba5

Observation daa241b4-b288-4055-81b0-54ead7f9de7c · outbound

This paper cites DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.931191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.931191Z digest=sha256:acf93bf7edab318ce689d1f1d682e09410878dd4e250ed38f4d2611e9b038168

Observation 25f7b39c-0cb8-4f13-a9ed-05c22fd440b6 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning START: Self-taught Reasoner with Tools

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.048539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.048539Z digest=sha256:c3449d1ad6f707ee84761984a2b97a481e98ba217383645477bb112148aca5ee

Observation 283ff077-586d-43e1-8ac4-58ba3cee46c4 · outbound

This paper cites Chain of code: Reasoning with a language model-augmented code emulator.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Chain of code: Reasoning with a language model-augmented code emulator

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.302008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:05.120040Z digest=sha256:bf2a5faeb9d70f072bab69c98788615e6d2c9d4358742174a525ab64eeb2b91c

Observation 57c7b695-3dee-4172-bfcc-48b6d2e5a6fe · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.190265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.190265Z digest=sha256:29e1a2744e1c01a395a452c13f3fe873468dc4300b721394d20aa542f862d249

Observation 04539f72-ae22-48f5-b753-e1a7d0010b5e · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.232033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.232033Z digest=sha256:b126191901b88232c69b8cf41a810da8b8c806108bfd697b9a4c4db0d6b5d7d9

Observation 3575f901-c584-4012-9934-210449d77720 · outbound

This paper cites RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.273622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.273622Z digest=sha256:1978bcd7c6ee8351c66c88e30884673623127ba46ff3aa3aa43dcce72fedda81

Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.361350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.361350Z digest=sha256:471db1c7fea2392c7f8285e8d8b3a3f9c882f237b581664431741a6c9ec667d1

Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.394679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.394679Z digest=sha256:b37c06378c58cee9f21d2086ab84cb6bedbf34e1f739de8af8d36f97f13ee4e0

Observation 745a06c8-a473-4d00-abe6-0858cc0119fd · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.495850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.495850Z digest=sha256:a8e777c63968a45df46e02e5dc7ca525a3738b5129a2c45acc027bccaa17cc90

Observation aea5fb85-94ca-4a23-9392-ebdc726e520c · outbound

This paper cites Let’s verify step by step.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Let’s verify step by step

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.292687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:05.585309Z digest=sha256:1756403ff0dabda2c70ebb88cdd5e532912d7785004730eed3c47cd747dc17c8

Observation 6b7fd29a-5863-4e42-96ef-e1de6797f489 · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.698406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.698406Z digest=sha256:89d83b53ccd07748e6be1018c467cfa1bdd88b0bc49210407730998b4f47bb06

Observation 445726c9-8530-4733-8eca-b261283df9e9 · outbound

This paper cites GAIA: a benchmark for general AI assistants.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning GAIA: a benchmark for general AI assistants

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.280556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:05.804990Z digest=sha256:d31982000f4804c5584dc1697592816cc65e1fa48bfda7dc92de93c810f4893d

Observation 93a9665b-d25d-40fa-bd45-651c6f9aef19 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.887519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.887519Z digest=sha256:3032b87aeb4541818d9bacf2094f7a893e2d5795f5afb7af32c74fdb70307cf1

Observation cf84693b-90e0-48cd-a54a-a6034290062d · outbound

This paper cites Learning to reason with llms, September 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Learning to reason with llms, September 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.988114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.988114Z digest=sha256:af50e3c08bd478eb919bf75bf725e37a042e2b46877a66fcef658af96affd950

Observation 3144e645-50d1-48ce-885a-71094a43d092 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ART: Automatic multi-step reasoning and tool-use for large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.095753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.095753Z digest=sha256:eebb625ca28c5eb84115706a5f5440c1165fe692ac940e5d301df3582c8580d1

Observation 5f3d20e1-635b-4c5d-a32f-33d56633289f · outbound

This paper cites Humanity's Last Exam.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Humanity's Last Exam

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.199430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.199430Z digest=sha256:c7ccdb895f42a4181b3fddfde766896959d8adbb522b4e6483f20e53259f2458

Observation b8539a7f-3075-48f5-bbce-6105aff6cb86 · outbound

This paper cites Smith, and Mike Lewis.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Smith, and Mike Lewis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.237798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.237798Z digest=sha256:2d543372af12b6a1355d0247b8060c90ee144941f48d726ddccaa55a0edc2b05

Observation b5233d9b-b8cd-4656-88ee-e72349f1e852 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.338068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.338068Z digest=sha256:9236a4e7151fe0e9280add3f6b7cce7dd3aa3e9c3be965f498ba07c82a6fd040

Observation a7298ef0-0a11-46c6-a798-9ec61ac14a85 · outbound

This paper cites Toolrl: Reward is all tool learning needs, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Toolrl: Reward is all tool learning needs, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.259059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:06.394366Z digest=sha256:21700dc9547ab5353e2903b36d68235f8737370878a70209d43f4f85ce391124

Observation 0a372f50-8758-4b4f-b9b3-60ce4a8d6509 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.397967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.397967Z digest=sha256:483b418f18ab308da37c60e1380b12d950d201ef5b712bbfba8f1158d5d7c6c2

Observation 0adb45c0-3add-4d97-9c99-26536730064d · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.440420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.440420Z digest=sha256:db1d1618417b103b6553c03155213fb33000a5d1d453f9b4caf18ff1e675e407

Observation 2997ecbc-83d1-4c34-bf61-4ee3967d1677 · outbound

This paper cites Qwen2.5 technical report, 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5 technical report, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.491523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.491523Z digest=sha256:9704e61c6132f2ef046fc054630fcbb50ed9dca893cfe5d231485db2370dd16b

Observation f7e10529-5f86-489a-8250-4815c38c2293 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.244120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:06.528602Z digest=sha256:0923a3b1573dd7d7f864b3833d6553b6e57b64440f84cf6fc0a13b03d1326365

Observation c0e5f2f0-0796-41dc-aa5e-5b70c7d0bdbd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.594525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.594525Z digest=sha256:58e2895baf1bf2c7c448a6e60ec41352b05fd9ef8d8199b8ad162ff2c4385859

Observation 00a5d9b7-a885-40fe-b01c-681f3acdd004 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.664381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.664381Z digest=sha256:b2c7bab28447390446475888183b81a6f551264b885350418321113e28220e47

Observation 93374031-0617-423a-82a1-ccf6d4d99499 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.718929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.718929Z digest=sha256:9fee945692bd48e832890ae09f861c312c946e40219913aa123bef1540191aeb

Observation 3fc324cf-9944-49e5-aea7-6b77dd04f0ab · outbound

This paper cites Curriculum learning: A survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Curriculum learning: A survey

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.233430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:06.770243Z digest=sha256:8722174c36e29fc543c35805c24f3be6fb26e8d520ee531ea379bf8fd633610a

Observation a1453fd0-e7d7-43f6-806d-6f5e890cfc88 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.794401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.794401Z digest=sha256:ca840055a13056dffe8531ca338d721ac247a531695d78ca77b3f881d9fbecee

Observation c7c6dfe8-d0b4-4b8e-8f90-86b78c232e7b · outbound

This paper cites Zerosearch: Incentivize the search capability of llms without searching, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Zerosearch: Incentivize the search capability of llms without searching, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.835005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.835005Z digest=sha256:1e1a999f81150338bbb05d946f4b47e9c9dfd5c07462435492a41c01134edcd6

Observation c6e6190f-f8ce-4ad7-b7b1-f53ae29ec45d · outbound

This paper cites A Survey of Reasoning with Foundation Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning A Survey of Reasoning with Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.912225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.912225Z digest=sha256:07559327c9c01660f846cb709517970b6755cf72d1737509f03e9bd1f05256ac

Observation 98a456f2-71ba-4ab3-acf7-9470f5701e26 · outbound

This paper cites Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:06.939970Z digest=sha256:efe47a47ace384aae962e4852247948eae50ab2c128570c98d0a911be874b1f2

Observation 2c61a8ef-fdc0-4fec-bd37-f51574ecfb16 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.075956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.075956Z digest=sha256:bdb77affc4abea72961d622b66b7c563fac20546a8b5111e5539d5dad1cedb2e

Observation c4530ae1-730f-487c-b9de-cd35b72a215c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.190061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.190061Z digest=sha256:8325ecf1ad974a7a6bdbd7925930c1221585ba576f32fd928663f4e6948fac11

Observation 8e086096-b12f-4f26-82b3-691746961c7b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.292322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.292322Z digest=sha256:b254a492e18c2d0bf4a549aa5e9927fb71fe6820d5bc36ab46b802aa2dfcb6d1

Observation 41c61fc1-3617-4a5e-a0f4-9210c9366594 · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.353123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.353123Z digest=sha256:15ed40975dd07d234fca0944b916c66063c3834f06f388b658887dbe8b96407e

Observation 291d80f0-80dc-4ebd-a067-e19d68876a68 · outbound

This paper cites musique: Multihop questions via single-hop question composition.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning musique: Multihop questions via single-hop question composition

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.492226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.492226Z digest=sha256:f088ef323fb4770d1ac6fab07d712a67c05cfd6a8b201be2e31f73a072b8b72a

Observation e3927b1f-8cf6-41e1-a29e-5b85e26a419c · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.633692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.633692Z digest=sha256:b11d4307b40d11931dd588adbc7457c507e515fb884e7621f6213794e699a4eb

Observation abd5b097-0a92-431b-afce-308df4f64937 · outbound

This paper cites Text embeddings by weakly-supervised contrastive pre-training, 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Text embeddings by weakly-supervised contrastive pre-training, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.793407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.793407Z digest=sha256:9394d44cb92c64da5671e133a62e587393d545d0f74b6dac15a68b3f6cffbfd4

Observation 6865c8fb-0fe2-458f-82c7-17d3d5577937 · outbound

This paper cites Reinforcement learning for reasoning in large language models with one training example, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in large language models with one training example, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.888830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.888830Z digest=sha256:39ee6149b58cad4fee99fee3e800f93cec806db05b4f576af71bb5351502b539

Observation 1c6d98e7-73d1-4a9f-823c-a9a0497ebfc8 · outbound

This paper cites Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.183390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:08.048505Z digest=sha256:bf13abf3b7b346cab5e6746f9b25b16fcbc90ac8d8883e4041ec317c5bb7c437

Observation ce961658-b063-4860-8900-c03507313a72 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.136388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.136388Z digest=sha256:16e0459e7687f5313ec97a30acd8a13c4dfe8be5c8f2f1e2ac880fd150ac03f1

Observation fc269afc-78f1-4380-933b-e35f5fe7a0d2 · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.224735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.224735Z digest=sha256:26e8c61c5de8e8acb55e2166a68222d3d1ae322b876733022ab152863e342532

Observation 2fc8f876-fd04-4022-a5c7-510118aae02a · outbound

This paper cites Qwen2 Technical Report.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.283937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.283937Z digest=sha256:10c9474e91008f0534b104011393fd901278e80e940b3f8fa6f676a537f8514c

Observation b498f759-afe6-4708-9ac1-e06713c18d97 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.324488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.324488Z digest=sha256:350356ba79b138a40d8ba9eb1e0eb83aae97fb3c61fb9a698575b25dcf723a1a

Observation 307a74a1-7a9a-47b0-b6d7-5d55cb513d6a · outbound

This paper cites Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.363789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.363789Z digest=sha256:e105abfb6e71787fa70ea8f5b1f7c6970b05156ff014238197793acd3bb17b42

Observation e548cbf4-4dc3-4a94-8cd7-2d6e1c820185 · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.408176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.408176Z digest=sha256:a959bd7d4d5ff7a6bf8841bb7f65d0c0a897e768959669136dde71446cc97ab6

Observation 81248e37-f9ac-4e80-a775-969a6c60d58d · outbound

This paper cites LIMO: Less is More for Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMO: Less is More for Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.440907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.440907Z digest=sha256:8e3646094391d238412157eb62c042ba93e53371451a25a80cab41f3dca1a508

Observation d4d0aa70-6d7d-4dfd-a556-9709db7a630e · outbound

This paper cites Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.455553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.455553Z digest=sha256:1ca545816c6e817114ca3fac9a17fe3c57a04e9a0df8b3d4a70e109cbe88cd1a

Observation c60b9b13-4423-4230-8314-19ace93877e5 · outbound

This paper cites SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:06:08.651537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:08.459251Z digest=sha256:8f5a752c6e712822516673b65fa5dfe227907abedfe888f592062b7f5d343737

Observation 4ee74034-adfa-4c42-8c6b-7a8371bd8dbd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.462641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.462641Z digest=sha256:8109e581c16863e106e1ac55584a2d8af17b07f71da5a7dc8a18498e5f79d15f

Observation 981b328b-5cad-4846-a3ff-9e4652bff27c · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.466277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.466277Z digest=sha256:f144ea2872825241cbb03b6bfcaef58ad31dcbfb2c7c5f19fa07c2c28fb685c0

Observation d58cba6d-1eda-493c-91ec-c37fb3e07075 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.469560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.469560Z digest=sha256:17f2156a5184c298aae4a55dd86f63938b4ff60d7c620af5bbcc52d86d41e0de

Observation ba4b92eb-7b3e-4d32-9517-c76a3d75a398 · outbound

This paper cites Agent models: Internalizing Chain-of-Action Generation into Reasoning models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agent models: Internalizing Chain-of-Action Generation into Reasoning models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.472839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.472839Z digest=sha256:42a0bf6a41cf8a930435451f1118289a623ac4803181b678f11b320a27dab5c0

Observation 23835f8d-1b01-4f55-8421-2fea2308e4d4 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 81

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:06:08.476018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.476018Z digest=sha256:d384968c0f13bff2292e11d5fdaccb3ee8022fe22a1ef89f87ab7eddc44ede8b

Observation 98931d4b-53cc-4d13-815c-8d322e3a3bf0 · outbound

This paper cites - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours).

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours)

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.167010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:08.479995Z digest=sha256:981cd6976cb568a3cfce514c4d3c3850420f2e47f460c6c2a6c18ccfc99ccedf

Observation be2b3f86-a863-4906-a4a4-aec64f33799c · outbound

This paper cites Converting t minutes to hours, we get t 60.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Converting t minutes to hours, we get t 60

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.155451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:08.483160Z digest=sha256:0e591a63de8f0d481f2bb707a4c5a961770f9cc23175bca5c70e22f47aebbd22

Observation e10bd72b-7cc0-464e-9a4c-f2c71d769f71 · outbound

This paper cites {numerator}/{denominator}.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning {numerator}/{denominator}

Reference 84

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:06:08.583497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:06:08.486342Z digest=sha256:39bc3f2d232ed6c15c44c2dcd090c911e077252b5b60936adcde24ffca079e85

Pith citing papers

Observation 26238ddb-55ef-4b2a-9f10-2735e6547f87 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.575865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.575865Z digest=sha256:7555fec650d07aa54194eef783001a914a1bb5b3305fddb6abe64c61c926a645

Observation 57099b37-beae-4b05-a5c4-9e9ca5be68b5 · inbound

Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation cites this paper.

Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.509403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.509403Z digest=sha256:c31ff76c68836d22beaa33fa014fa43e18c8bf8819d608c2b1b80bd01872e78f

Observation 8249ae1b-0f50-430e-a455-2913c4bae6bd · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.317119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:8a811ce22498f436810b4e8f5cd6011df4c6a953a61d45555cc922494eee3b7b

Observation 61e64d85-0ef9-4e7b-8d37-94c49eb4807c · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.066336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.066336Z digest=sha256:708ec126fbc76d4a00de4fb9fe46330a3464a7ef78cb29d50b6fdfbb4d25727d

Observation 1ec53e26-e08c-4eaa-b6f8-4349080788a5 · inbound

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning cites this paper.

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:13.551033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:17:13.551033Z digest=sha256:192b0157186f71fd3f0c96c9688751f562cb2b5a3fc00e1d10ab90e83d804296

Observation 5c8e51c0-4b72-44f0-baa2-9ff8928beed1 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:21:42.329827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:69aa8dde159fc4f171995d8bc612bcbd9484fc90fc02e5f043915038f693e29a

Observation 6d12bb68-7528-4280-9e78-e7c7d7d238ba · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.860107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.860107Z digest=sha256:8aa2238bc77c9a125278189d8558227f686022f6e9b2b2b6e90710c343ab0a9d

Observation 8c997f8d-da85-4462-888f-99a8a655b5c8 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.278117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:2d1440a0bfe5db9e493a76e9b5270990f290067bf48b55e39f5ba576432786db

Observation 4cf73df6-8d84-4a8c-9504-70b53a7eb141 · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.334039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.334039Z digest=sha256:a030840f72e8e7e42ac80b03c86d38a93c3f9fab79bfaa1f57a51cc0ec9d8113

Observation c39a617e-8a38-456c-a1b9-3b25bf595665 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.753936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.753936Z digest=sha256:c0f52e05f82099f622e5fff9aa22a6836faf685d4cc9ac6674b73584e7a6f8c3

Observation d53bfd2c-ef7d-496d-9924-bbd31e9e9dc6 · inbound

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning cites this paper.

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:49:51.555062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:49:51.555062Z digest=sha256:a456ec4563984146cda606eb3f5b52bee3487a4bcd369bae5237c6990497bf16

Observation 72f33c36-2fab-44a6-93a7-49a1d4ea5b14 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:36.415695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:36.415695Z digest=sha256:04f7023fb34befd9136498a063f591426541447ece18e91c51a5b71721672986

Observation cca8698f-4717-4727-b186-7865c2f30559 · inbound

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning cites this paper.

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:56:57.071345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:56:57.071345Z digest=sha256:6062000a6565c8d69d08bf4ba7e2f23a8677f362556cc3ea743a0848e294d7c1

Observation 2eee2387-aef1-4049-93d4-754602ccb4fc · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.747261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:ca7c4064c2031ab80038731ad91fc05921135c5cdc6840836066f9fb0f35e8d6

Observation 782c5788-7aec-4c3c-ac47-15f9e354c63a · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.288717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:21fc0b81d55ba0458b9b72b19a86c3fa2acf42fb3ead77de5a3ffa84f542dee2

Observation 084cdba8-6810-472e-b25e-6bcfd79a999d · inbound

Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA cites this paper.

Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:49.508970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:25:26.811390Z digest=sha256:146c821a34f0e8b193203e773c4c50e364825edae50e369143d4b2ef6f0482c4

Observation 31bd0abe-80a2-4ab8-9381-690740a6df9a · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.937831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:e509014ddcbcb9fb35488d19cc521532ecf8296e9b093ad605f2f458d279dbd4

Observation 3bee09b4-f1f4-4348-9163-07973725b8ee · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.424047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:86c536fbf15a7fe596bb80c4c2c6247698d35a557c1e08f6557eea8446f21ce3

Observation e9ac6733-1ad4-4aa6-b9d9-15bdb1d04c7b · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.223482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:128caf765ac0ad930cc29de48e0d09d268d45d18c1726d60598f747889ff0eb3

Observation 407fad62-21f3-4e68-b6b4-9bf344089a81 · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:56.498224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:33:02.995360Z digest=sha256:50fb3e59620c5ef15820570b52277d71bb594297b66c6bbb4f398c88bb5a08c1

Observation c63639f7-9229-49a5-8728-2f23fb64545b · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:28.390643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:26:11.265781Z digest=sha256:64b47283ca806a8742a9b60eeeeff030646ae476bfe5400731e68caac717e3b1

Observation bab25d3d-c1ec-460f-8e49-fc9827819e64 · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:23.986138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:e45d96ab79e8643480133ff6b2d47eacad4a989c9ece473782761a6a3d5ea30a

Observation 2154c720-8d36-4496-9cab-3c05b7e3268e · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.278748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:47f96ee804174885bec3cf6b6b665ff3de335764f5635994bbd5edd06fbb1583

Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.727198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:216cca5255c83c13a448fc10d6e6d41192f0c89e2deb050f01e356f6fd252073

Observation 0a0d90ec-834b-4ada-b8af-32f410603583 · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.303055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:16332032f3223d726819abe68c728023a758f73a103e68424643f847bdb210ca

Observation 707e0454-297d-4895-a8f1-af98616163da · inbound

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction cites this paper.

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.517513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T06:05:45.046156Z digest=sha256:e53c747616cdc2366fc52de4aeccb0778248582a7b77cb72e5604966345f5b4f

Observation 898960cb-a66d-4f3e-8cb9-5dd058afa2b4 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.553725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:d6dd2d0d609cd59acb379dfab0e8b14e96969981ebb24383a3b1bf0ef7a3aae5

Observation 639eeddd-856a-470d-8cbe-1fd351a1493e · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:02:18.700268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:aeda1364fa987bcdd3259e7aa85b7e98e60c8fec237c8fc4cdd94b6e66213bd8

Observation 80e8188d-8c77-4e2a-b6bc-3d98bda5218c · inbound

ToolFG: Towards Well-Grounded Fine-Grained Image Classification cites this paper.

ToolFG: Towards Well-Grounded Fine-Grained Image Classification Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.338611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T14:52:06.186594Z digest=sha256:7aa9e4e34301d1779e651ed1ba1eda47639371a35dc51f69ca5458473e0ca127

Observation 27969eb6-8c88-47cd-8b7e-643059f52c85 · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.289106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:f8fa1a389f7a83ec0410deaad6e705f541bd55a0dcdf9017ed8e8be46dd91c61

Observation 5d04b888-c649-42c2-b7ca-ae6f7837d636 · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.884050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:e91bcfff46ed5fc529b98a401378d28cb71010279e11e4d73e305b292b35383e

Observation 0e47b6b6-929a-44e0-b716-146985896a3f · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.295388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:08cbe6270d53567ec4f6949a819b84a50e1643a46bc3c94af3ba5b375b0884fd

Observation 44a3994b-2af1-4c3e-9f49-65b04b8794c1 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:26.150859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:26.150859Z digest=sha256:d7e2676f2b2eb2549a274b6b51993c2aad2744fe334f18325ed490ef4df19d62

Observation a6de5c47-ed61-469e-a188-460913426f03 · inbound

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents cites this paper.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.509228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.509228Z digest=sha256:871bfa365e45e0be7934e4adcdafed73a013d2e14d7fd582fe9a8403f453db77

Observation e01b077e-76b7-4348-b052-0bcf74aafbde · inbound

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs cites this paper.

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:38.022901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:38.022901Z digest=sha256:db607732df84b9af8c9da95250871455cecc190435071c18c2b9ecffaab01a2f

Observation 032d3af7-9043-4df6-94e3-156e805cafde · inbound

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation cites this paper.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.315097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.315097Z digest=sha256:e794f95c1b8814aa07155d9f9861e8f1031440e1b579e6eef90cc0cf6a059a89

Observation b6ccbbc4-bf13-4618-b828-bd37700535d2 · inbound

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation cites this paper.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.144305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.144305Z digest=sha256:a02bb7118ad0507f71e47787126c7949d4bf67a4b3ce4b7da1c0ed49083ea92d

Observation f9f149f2-ea60-4886-ad4d-d4362bbd38fb · inbound

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning cites this paper.

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T18:36:23.555649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:36:23.555649Z digest=sha256:5751895c5b9d3356d83d6667687df8c46395978475cd190fa8b24cea3f237692