Pith. sign in

Paper Citation Record · LEDGER

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 38 inbound Pith citation observations for arXiv:2505.16410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16410 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:08.486342Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:26:56.575865Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a768c1a7-f587-45d9-afb8-9244aff5719c · outbound

This paper cites Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:02.921197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:02.921197Z digest=sha256:427786f196c46c8dbba990282ba34bf106b33a5f36debc20cfe858ad1c331e48

Observation 9a52689b-1077-4ebc-ab31-0003e6de246c · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:02.965213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:02.965213Z digest=sha256:7a94106902e7e936625f17326512ab00649fe53ad0038fce6fc58f7a6d097243

Observation c5f52244-f3d1-4478-92df-621630b3f6d6 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.057637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.057637Z digest=sha256:a7520ee907182378e323a6ced5ab7ac4d85c30fae149f78eaaa3094c74a948c3

Observation 08d50b1f-5230-48c6-9bf5-b7e9303ae65f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.258543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.258543Z digest=sha256:12de33bfd8d64b0bc2ad21294eb8421874b81f0a85f6fa7ae6664ac74bbec18d

Observation 6e0db75b-ba54-40fd-a9c3-04d5116bb108 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.338887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.338887Z digest=sha256:08d9dec24b6e5131572fbffadffeaadbb125bc186ac583bacbe8e0bd760fd429

Observation 3f2a9b67-4572-412c-9ed6-5b8da5fb3627 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.410193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.410193Z digest=sha256:a6960f9855643dea8404eeaee3c9b73b1eda08be693e1cbb29f75222704455bf

Observation 9db8e6d8-9076-4853-be05-13d988fca364 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.377945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:03.487236Z digest=sha256:e77f6c64915b2e52e8b35a871d923e4b05fda86ada281547a5bb60914819ea4c

Observation 8a01c378-6eea-464c-945e-8b952fd37e6d · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.575092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.575092Z digest=sha256:78eadc5a0e9954b66eba49d54ebf57810335e73f38bbfb69df02e847e1198490

Observation 463ec17e-ac30-4bdb-8ea2-81a9d406ba19 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.644180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.644180Z digest=sha256:f21e2ca9e400780261eef9c7321e893f15fac8cdbb22675a61500a95770e3a47

Observation 59afab32-de7d-400d-9f4e-bda9cdd99418 · outbound

This paper cites How abilities in large language models are affected by supervised fine-tuning data composition.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning How abilities in large language models are affected by supervised fine-tuning data composition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.722976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.722976Z digest=sha256:0380cfb916ccc92a20d29c9f771dc74ed241f8861ac73f2ca5d86446d10a7e3c

Observation 534947a6-a771-46e7-89c1-5c0506bb8b8a · outbound

This paper cites Progressive Multimodal Reasoning via Active Retrieval.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Progressive Multimodal Reasoning via Active Retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.782011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.782011Z digest=sha256:0a4e3258ba64945ec285266575d4572d5b64139b526694eb53c95e10407bc2e2

Observation 05b45d04-ff58-4c3c-9d94-db92a37592e5 · outbound

This paper cites Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.853422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.853422Z digest=sha256:ae392945e426095a7fdf93c41488ff626fb427e213e684716759dd16a36e07b5

Observation 3a8ccf65-cd6a-4971-93df-8ea0870efad9 · outbound

This paper cites The Llama 3 Herd of Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.926256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.926256Z digest=sha256:499228bbd0181a1e78e1d63f1bb904826e4ebb8e5338e92dfeceb757fb72e57f

Observation 913923bf-d874-4927-8775-ad0923d608cb · outbound

This paper cites Concise reasoning via reinforcement learning, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Concise reasoning via reinforcement learning, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.356157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:03.998962Z digest=sha256:34c4e2de6ccf5c875e28f74eea4c819889780cedc944ef0a93084218233ea672

Observation f7aea587-87f7-4ccd-9cd7-cf89a904e6da · outbound

This paper cites Retool: Reinforcement learning for strategic tool use in llms, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retool: Reinforcement learning for strategic tool use in llms, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.079207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.079207Z digest=sha256:d08e7c2bcd84289beef2133da6f8a03dd76c9742a7ebd2c07517d5ced7535f17

Observation 23cafa84-7db2-4b24-bf89-2225134fdcf0 · outbound

This paper cites Tora: A tool-integrated reasoning agent for mathematical problem solving.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Tora: A tool-integrated reasoning agent for mathematical problem solving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.339342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:04.133757Z digest=sha256:6ffae671430a01febbe071ad40e5a0e1e1905b0e02ff9cddcd15f0943e169a10

Observation f0ec14f4-adee-4938-ae80-09f21d6f692b · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Measuring mathematical problem solving with the MATH dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.244504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.244504Z digest=sha256:764f5ca0faa2eb9cee644e5cce8538131fc9690915f8e48c4e673b0dbfc05078

Observation 8458cacf-0f84-45b9-8de3-68df07e66012 · outbound

This paper cites Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.304466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.304466Z digest=sha256:42636e1e9dec2942556be014c7fef4d2d931a700e0ea4628ec856d7f50446f2b

Observation f8588b63-1591-4eb6-bc10-acab3f7332ea · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.343006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.343006Z digest=sha256:799cac34de96bb02004b0802fe35f508c627f156c9457c1cbac19f5fc3d60a21

Observation e9160051-d954-4e11-b638-b95197f87cbb · outbound

This paper cites Towards reasoning in large language models: A survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Towards reasoning in large language models: A survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.459705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.459705Z digest=sha256:bfa8260da85f7479a6d8505fa4ce0f9977f04a9044905ca4a731528de9743837

Observation 3cddfc57-598a-4628-8a98-68a78fd6e4d9 · outbound

This paper cites RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.555452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.555452Z digest=sha256:f1607dd5bae475f7e855f98088431c63ff3787e61e90928448bf630ff3564024

Observation bffe5425-3732-4e3f-9c2d-fde963dae4b9 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.638038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.638038Z digest=sha256:a829e11c02ce1765a09d77ddf8f941e667e6c6c494e296bf37e6e6fb6042af56

Observation 83ed383c-4019-413e-a777-e91fd6877bd9 · outbound

This paper cites FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.709451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.709451Z digest=sha256:f11103778f2293f584bca1bb57b29e0ff7b9c6a81d1f3e374a1959124234df95

Observation 200f13a6-785c-47f8-b91a-5eff891f837a · outbound

This paper cites InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.761722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.761722Z digest=sha256:27466a6ebb8458f32669a272f3055a9b011316b9611ef9345640c063a5cfb121

Observation 8416c664-e224-46e3-a43d-0bdebc16618d · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.311394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:04.856927Z digest=sha256:88a9ad7f9e3a326f3e7527fe8fae67da85a783690e23de3a717281099411595f

Observation daa241b4-b288-4055-81b0-54ead7f9de7c · outbound

This paper cites DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.931191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.931191Z digest=sha256:f3c65829b6e44f6cf4be56d19daaf8b7bf720b20db978af57f30efa8c3c8d6a3

Observation 25f7b39c-0cb8-4f13-a9ed-05c22fd440b6 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning START: Self-taught Reasoner with Tools

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.048539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.048539Z digest=sha256:e07b12d5ab75c3e6cf9bd7d576fed1aecfc310b13bdeedb3bfa606055d209f27

Observation 283ff077-586d-43e1-8ac4-58ba3cee46c4 · outbound

This paper cites Chain of code: Reasoning with a language model-augmented code emulator.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Chain of code: Reasoning with a language model-augmented code emulator

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.302008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:05.120040Z digest=sha256:ea63afd38f66f240bffb703ef900644b1dc5b07436d0365abcae5bdf6c21a1c3

Observation 57c7b695-3dee-4172-bfcc-48b6d2e5a6fe · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.190265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.190265Z digest=sha256:5dc363e3e50379f9bb91f0c81214b7a4e266054a14aaa9a34af6b5f783d6dfa6

Observation 04539f72-ae22-48f5-b753-e1a7d0010b5e · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.232033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.232033Z digest=sha256:8d54f3b79f59456d250f5d594f8c3c78024df66025b253018b9bd3e89016fa92

Observation 3575f901-c584-4012-9934-210449d77720 · outbound

This paper cites RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.273622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.273622Z digest=sha256:38e1425cdcd97f8252393854c32b6f0b34c5fc2020a95c4cf8decbb049a84717

Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.361350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.361350Z digest=sha256:5e6f46373c100a4a16fc3e854dc1cbe22f5fd027287895d91b2621424234cbd3

Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.394679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.394679Z digest=sha256:e69bfa0a48d3cc69ef966eed9a83b0c13df2f2c410eb91b7f879b41eaac99213

Observation 745a06c8-a473-4d00-abe6-0858cc0119fd · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.495850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.495850Z digest=sha256:7ba485eb0475e5ea480dcf40c9202d75e3da8670b7762bcf5cfad1a9c84a34c2

Observation aea5fb85-94ca-4a23-9392-ebdc726e520c · outbound

This paper cites Let’s verify step by step.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Let’s verify step by step

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.292687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:05.585309Z digest=sha256:bf5b9a8de9fb46475b5aaebca7260bd06b41a968bd9625b7df92fb24e99f20e1

Observation 6b7fd29a-5863-4e42-96ef-e1de6797f489 · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.698406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.698406Z digest=sha256:96db4fb16f24bbf50b9ce7c9d15fa2b65a28a8ed3284d6d0563c506754dfa0e9

Observation 445726c9-8530-4733-8eca-b261283df9e9 · outbound

This paper cites GAIA: a benchmark for general AI assistants.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning GAIA: a benchmark for general AI assistants

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.280556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:05.804990Z digest=sha256:12fc1ee4475ec8eb260aa35ce3aafa32bcef6c9e902024198817e7d6779f9506

Observation 93a9665b-d25d-40fa-bd45-651c6f9aef19 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.887519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.887519Z digest=sha256:cb50916634c337706c76201702d34983cff9faffdedcfca84f45212cee3794b2

Observation cf84693b-90e0-48cd-a54a-a6034290062d · outbound

This paper cites Learning to reason with llms, September 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Learning to reason with llms, September 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.988114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.988114Z digest=sha256:19a82a53e4e024dec3f2d1cf70da9b49b14ed2d61bfbffb133acd6c07581f446

Observation 3144e645-50d1-48ce-885a-71094a43d092 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ART: Automatic multi-step reasoning and tool-use for large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.095753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.095753Z digest=sha256:94292f00d778ec271d1aa651639ab564c402918504e1a2afc2e0744135c3a0a6

Observation 5f3d20e1-635b-4c5d-a32f-33d56633289f · outbound

This paper cites Humanity's Last Exam.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Humanity's Last Exam

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.199430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.199430Z digest=sha256:3fc1594b2fb24d2480e2f11d3dbbdd520aeadd8fc3c048cdcca954a813623dac

Observation b8539a7f-3075-48f5-bbce-6105aff6cb86 · outbound

This paper cites Smith, and Mike Lewis.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Smith, and Mike Lewis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.237798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.237798Z digest=sha256:3a1a6983de11fb3f6ef1c7b3713a187a5d994c7e59d60530ab240e7901a932e0

Observation b5233d9b-b8cd-4656-88ee-e72349f1e852 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.338068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.338068Z digest=sha256:26a0bc79a08171b0f6f7e4d5cf4f23bae5e5e6090f93b136e06b06ca5f30c75c

Observation a7298ef0-0a11-46c6-a798-9ec61ac14a85 · outbound

This paper cites Toolrl: Reward is all tool learning needs, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Toolrl: Reward is all tool learning needs, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.259059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:06.394366Z digest=sha256:f1ea6b2fc9fac914dc02f1194683f37e38ee2ab82d7292fd744069dd701301c6

Observation 0a372f50-8758-4b4f-b9b3-60ce4a8d6509 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.397967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.397967Z digest=sha256:391cf2fb74d89f921cce81e6473494c49e5cb7bbe8bf99b5b49c6b7047d03237

Observation 0adb45c0-3add-4d97-9c99-26536730064d · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.440420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.440420Z digest=sha256:9961e59d928f12183f2e0f08c8bec294c0393095a905852a4672d820e0c23c90

Observation 2997ecbc-83d1-4c34-bf61-4ee3967d1677 · outbound

This paper cites Qwen2.5 technical report, 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5 technical report, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.491523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.491523Z digest=sha256:628701d3c4ddf296cf49f61ccb732de76959f8577eee0ea09e367f4ad8677385

Observation f7e10529-5f86-489a-8250-4815c38c2293 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.244120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:06.528602Z digest=sha256:8243d8a6b915e7d9f6f0b5047aa802ce6cc2be102cdeb0c135d7079e4eace996

Observation c0e5f2f0-0796-41dc-aa5e-5b70c7d0bdbd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.594525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.594525Z digest=sha256:242823951c3e4eec7f54569f795de896437ebafe88dd5e20dc778b29b431be16

Observation 00a5d9b7-a885-40fe-b01c-681f3acdd004 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.664381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.664381Z digest=sha256:60c04bb92def6adeba03a57c0413024e54ae6d8d76fdf3203158b4b680c015b5

Observation 93374031-0617-423a-82a1-ccf6d4d99499 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.718929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.718929Z digest=sha256:9254f0a7aa23c433b06bc66e61985e64939f9915c1a2a3105395288e614d0bc9

Observation 3fc324cf-9944-49e5-aea7-6b77dd04f0ab · outbound

This paper cites Curriculum learning: A survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Curriculum learning: A survey

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.233430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:06.770243Z digest=sha256:a71efad3c4366dae5b67ec5f2b1e59135a14ba084b94fe8f17d77be21a895d19

Observation a1453fd0-e7d7-43f6-806d-6f5e890cfc88 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.794401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.794401Z digest=sha256:8dac5f8274fceb9e8250457426aa5dadfc51e84d99715a40922c9d7979804a57

Observation c7c6dfe8-d0b4-4b8e-8f90-86b78c232e7b · outbound

This paper cites Zerosearch: Incentivize the search capability of llms without searching, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Zerosearch: Incentivize the search capability of llms without searching, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.835005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.835005Z digest=sha256:66c03fe1a7afb25b0c3fc0173c07cce0bc0d8eeb109a36db0e86dc927696a215

Observation c6e6190f-f8ce-4ad7-b7b1-f53ae29ec45d · outbound

This paper cites A Survey of Reasoning with Foundation Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning A Survey of Reasoning with Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.912225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.912225Z digest=sha256:f008997b8f66be3f9d40914fd6429f5ab13a1997e5a6e28ce203ef076a5a0677

Observation 98a456f2-71ba-4ab3-acf7-9470f5701e26 · outbound

This paper cites Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:06.939970Z digest=sha256:a929ee559948ba4aa7a5061003bf40b201b8a7d1dadaf42bebe6b06227ef7be4

Observation 2c61a8ef-fdc0-4fec-bd37-f51574ecfb16 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.075956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.075956Z digest=sha256:0e5e278df768ce826acb3938ddf32a67edc0536360452535c631385625dd60b5

Observation c4530ae1-730f-487c-b9de-cd35b72a215c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.190061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.190061Z digest=sha256:e83590fc45f2fcc698c2350c17e65960557e49efc7bf7c771c7ede5b58a3b79a

Observation 8e086096-b12f-4f26-82b3-691746961c7b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.292322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.292322Z digest=sha256:f74fde5a1f135cb9626fd413d9c5d888a95f82734d4582c3fc8b8318d66d6d3e

Observation 41c61fc1-3617-4a5e-a0f4-9210c9366594 · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.353123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.353123Z digest=sha256:fc74fe567571ac45414b4dff0247f29f7a30e18351e3d508572e3ecd50eefb30

Observation 291d80f0-80dc-4ebd-a067-e19d68876a68 · outbound

This paper cites musique: Multihop questions via single-hop question composition.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning musique: Multihop questions via single-hop question composition

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.492226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.492226Z digest=sha256:6346fe557ea13dfa9f0067560ebdf1cf65b4a8b99f2356add870a450db3c1fdc

Observation e3927b1f-8cf6-41e1-a29e-5b85e26a419c · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.633692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.633692Z digest=sha256:dfe21703594ebfd3c3c1535d98546fe417755a389a097ce0556e018d6cc9d740

Observation abd5b097-0a92-431b-afce-308df4f64937 · outbound

This paper cites Text embeddings by weakly-supervised contrastive pre-training, 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Text embeddings by weakly-supervised contrastive pre-training, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.793407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.793407Z digest=sha256:99b0d0fb40c7b4ef108589af86f314f34b9a960fab23c6f751c3e2393483bbc7

Observation 6865c8fb-0fe2-458f-82c7-17d3d5577937 · outbound

This paper cites Reinforcement learning for reasoning in large language models with one training example, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in large language models with one training example, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.888830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.888830Z digest=sha256:e03d84ada33f62bedcc5573996ad6f026158033358eee1c637f088435cf85cf2

Observation 1c6d98e7-73d1-4a9f-823c-a9a0497ebfc8 · outbound

This paper cites Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.183390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:08.048505Z digest=sha256:e300fe4366dfffca597a53b2d6dbda9986b7cfdd2536762f2afdcd453401bf1e

Observation ce961658-b063-4860-8900-c03507313a72 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.136388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.136388Z digest=sha256:3e9b1bbb69cea25c5064c13f99c78f87bafac9d48eff98cae3a03e6fda2af6cf

Observation fc269afc-78f1-4380-933b-e35f5fe7a0d2 · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.224735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.224735Z digest=sha256:6347c83acd9af6d2b2ef5b72d12b6a9fad2012255885566a4143823f806f16b8

Observation 2fc8f876-fd04-4022-a5c7-510118aae02a · outbound

This paper cites Qwen2 Technical Report.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.283937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.283937Z digest=sha256:2f9003c7dd2f0dda1fa33817a251e3b8c19b0711b8fa0700eea79f3951f52a6e

Observation b498f759-afe6-4708-9ac1-e06713c18d97 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.324488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.324488Z digest=sha256:fc48f140347627b9a004f072c372d9d91663af9253dbbf5f0952116fc7828f5e

Observation 307a74a1-7a9a-47b0-b6d7-5d55cb513d6a · outbound

This paper cites Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.363789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.363789Z digest=sha256:d8552f9016291cbbbcf98bd1cce827ed0d6c8aa0133485cbb0c0401e71f24315

Observation e548cbf4-4dc3-4a94-8cd7-2d6e1c820185 · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.408176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.408176Z digest=sha256:8eddffa32c29b039cedd8c37257e3a03567ce4f228926c5126f0a3044107d698

Observation 81248e37-f9ac-4e80-a775-969a6c60d58d · outbound

This paper cites LIMO: Less is More for Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMO: Less is More for Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.440907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.440907Z digest=sha256:70ee0d2af90420001f6b9532ab08417439ef7b81600b10b7e43d1bddd83a8a86

Observation d4d0aa70-6d7d-4dfd-a556-9709db7a630e · outbound

This paper cites Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.455553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.455553Z digest=sha256:ca5b7c8d061cb2868b2835fe228bb26d256cb027bf893781bd3b6162b4abe4e5

Observation c60b9b13-4423-4230-8314-19ace93877e5 · outbound

This paper cites SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:06:08.651537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:08.459251Z digest=sha256:c4c04c70addf70b08eb92cb0bf13f02cf3044cc69775d561cadb7c2403c2f5ae

Observation 4ee74034-adfa-4c42-8c6b-7a8371bd8dbd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.462641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.462641Z digest=sha256:8458d421478b91ae1cad47608bedd815521977dd4f737367fa42d9dcbf519aae

Observation 981b328b-5cad-4846-a3ff-9e4652bff27c · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.466277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.466277Z digest=sha256:ccf3d242805e4e8147bf59b3344c18bb2fadab30cf6f4dd3f16d3221159f2250

Observation d58cba6d-1eda-493c-91ec-c37fb3e07075 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.469560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.469560Z digest=sha256:4cf92491d0167973efadb103b3aacefcae4daebfbba6a3514f8f58fdf11b6d30

Observation ba4b92eb-7b3e-4d32-9517-c76a3d75a398 · outbound

This paper cites Agent models: Internalizing Chain-of-Action Generation into Reasoning models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agent models: Internalizing Chain-of-Action Generation into Reasoning models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.472839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.472839Z digest=sha256:aabe0c57b6f973ecd6ea91ce943394339ab7f8e31b375e477d43d0603d91cd74

Observation 23835f8d-1b01-4f55-8421-2fea2308e4d4 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 81

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:06:08.476018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.476018Z digest=sha256:755405a3cbd7b7544e353787391fa3e5823ba4559c8a4200d5222d7c4817bf61

Observation 98931d4b-53cc-4d13-815c-8d322e3a3bf0 · outbound

This paper cites - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours).

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours)

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.167010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:08.479995Z digest=sha256:27c6ee2bd62f282252aab96ad0300c805de066191acba5a569d33b169bd029b7

Observation be2b3f86-a863-4906-a4a4-aec64f33799c · outbound

This paper cites Converting t minutes to hours, we get t 60.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Converting t minutes to hours, we get t 60

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.155451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:08.483160Z digest=sha256:abf392d4ea7cb50658162d15d0ef1053d6750c17ec4efc4ef1e4f5da6cf17e22

Observation e10bd72b-7cc0-464e-9a4c-f2c71d769f71 · outbound

This paper cites {numerator}/{denominator}.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning {numerator}/{denominator}

Reference 84

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:06:08.583497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:06:08.486342Z digest=sha256:ef8fc407be2e975d949908e666c571880d6f36ed9647230189b83a4bc17ce94d

Pith citing papers

Observation 26238ddb-55ef-4b2a-9f10-2735e6547f87 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.575865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.575865Z digest=sha256:407b321c32e2dc05c6dda9a9c7b2e34acb8b70461abe68c3be2e8df375b77e15

Observation 57099b37-beae-4b05-a5c4-9e9ca5be68b5 · inbound

Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation cites this paper.

Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.509403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.509403Z digest=sha256:71cf64724d09f81fd7d6ddf200da1a5b4f8194317a51b8755a79b94e5fa5d4a8

Observation 8249ae1b-0f50-430e-a455-2913c4bae6bd · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.317119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:c6667d93aa47f842ed7e43ee0df8598049a05cbd581f4f730cf6db1a559a1547

Observation 61e64d85-0ef9-4e7b-8d37-94c49eb4807c · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.066336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.066336Z digest=sha256:c0cd0a80a5a1efae97322ef69408e90de86cea1004026c8f047a1c745c4af904

Observation 1ec53e26-e08c-4eaa-b6f8-4349080788a5 · inbound

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning cites this paper.

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:13.551033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:17:13.551033Z digest=sha256:8dd4e33fbb5a8b6ca87f346c0190d3a00f2f894df6af02a2471eb39bb565fd5c

Observation 5c8e51c0-4b72-44f0-baa2-9ff8928beed1 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:21:42.329827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:99be3a41b683b011b3f4a9d109ff2c1e51e1838454a3dfe34f87bd3dfa8877ae

Observation 6d12bb68-7528-4280-9e78-e7c7d7d238ba · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.860107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.860107Z digest=sha256:06a3a9f41ce9ed016afa9391fb662549e5236e0957f65b43de725d9877ff5a09

Observation 8c997f8d-da85-4462-888f-99a8a655b5c8 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.278117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:e87d5d5d2204e8dcad913a55657a6e0b24a37bb27f72a1ef8704cbd56c057572

Observation 4cf73df6-8d84-4a8c-9504-70b53a7eb141 · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.334039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.334039Z digest=sha256:85d99174a1ece12bca09c21d3ada6be015b3ed182d21a0694f8b0b1d75802831

Observation c39a617e-8a38-456c-a1b9-3b25bf595665 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.753936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.753936Z digest=sha256:f4eef3a45a2eab47c3c69fa7785f812157c615c6a61a8f47b0197eb5e07dc92b

Observation d53bfd2c-ef7d-496d-9924-bbd31e9e9dc6 · inbound

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning cites this paper.

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:49:51.555062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:49:51.555062Z digest=sha256:ba47cbb5d3b8a7d421a4e1d26ebb2e5c739fbcfe6717b4658e41ab2fa1f35ab8

Observation 72f33c36-2fab-44a6-93a7-49a1d4ea5b14 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:36.415695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:36.415695Z digest=sha256:203534fef91fff8153693fa4ac53de62034b0ff8f97737e69889ef5049f8654f

Observation cca8698f-4717-4727-b186-7865c2f30559 · inbound

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning cites this paper.

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:56:57.071345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:56:57.071345Z digest=sha256:f50ac201479ba93d8fd5d2a26082d0607c2aaa725ce4ee6b8ec2b8d3a03732e4

Observation 2eee2387-aef1-4049-93d4-754602ccb4fc · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.747261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:8b7dbc67da8d6f53465c54ef7f504bb3f45bd937b9bc4207fe0a9d58e8f4761a

Observation 782c5788-7aec-4c3c-ac47-15f9e354c63a · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.288717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:6aeee45aa4307aee523f8b2b9a6cea1e0d4549df16a437975904cc56df2dfee9

Observation 084cdba8-6810-472e-b25e-6bcfd79a999d · inbound

Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA cites this paper.

Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:49.508970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:25:26.811390Z digest=sha256:468e8908e309607e2555d14d1002eaabaa7436492cbbc40e87488ebe02ab380e

Observation 31bd0abe-80a2-4ab8-9381-690740a6df9a · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.937831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:5ce4765ffd5e8b92d1df5d08773236c7a9267535f964988e319a1cc1b664ced9

Observation 3bee09b4-f1f4-4348-9163-07973725b8ee · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.424047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:7ba3d0ef9057a3d7d7d839daeedb8748a5238a5065e0a4e8b46b2b981c501136

Observation e9ac6733-1ad4-4aa6-b9d9-15bdb1d04c7b · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.223482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:1e28534e28fa1058f27e875f194744abb602f3a39dd14a0eccfb8f4c02511f68

Observation 407fad62-21f3-4e68-b6b4-9bf344089a81 · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:56.498224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:33:02.995360Z digest=sha256:48113bb3907a7a4ddbc36d1559b016a3e21314dcb5a1e992ac0ea44d18a78a27

Observation c63639f7-9229-49a5-8728-2f23fb64545b · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:28.390643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:26:11.265781Z digest=sha256:baf5cb983364d14d2e87ced53e277f337f0672217a01a15717635fc386199b44

Observation bab25d3d-c1ec-460f-8e49-fc9827819e64 · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:23.986138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:220eab9fc3fbeeb8b7aa8c3811027f9a2c82f8524bc5132b6e0f1b227d216065

Observation 2154c720-8d36-4496-9cab-3c05b7e3268e · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.278748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:55b501881292d1da50248a6ad7ca108b22d1483cdbc020423a91e65ac123a092

Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.727198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:864a0276c77e15e4a61b3e4723993979c4304c0da04902c7b7588cf632f5c574

Observation 0a0d90ec-834b-4ada-b8af-32f410603583 · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.303055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:02f19b375e25ffcd3e41e735898638a557ade6247952072b8c670223d15465f7

Observation 707e0454-297d-4895-a8f1-af98616163da · inbound

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction cites this paper.

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.517513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T06:05:45.046156Z digest=sha256:9f8633225b7cd86b7d670da2337ba460dd1cbfc9df892e67d0199043a3054ce1

Observation 898960cb-a66d-4f3e-8cb9-5dd058afa2b4 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.553725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:50dddd6e0174a3d47d91108249b1972caacf7a85766276dbffc6ddfaa14cb541

Observation 639eeddd-856a-470d-8cbe-1fd351a1493e · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:02:18.700268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:4b593b81bde678bda4d87dc9ddf312663482699fe9ff1b479a91b4f79eb61df1

Observation 80e8188d-8c77-4e2a-b6bc-3d98bda5218c · inbound

ToolFG: Towards Well-Grounded Fine-Grained Image Classification cites this paper.

ToolFG: Towards Well-Grounded Fine-Grained Image Classification Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.338611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:52:06.186594Z digest=sha256:339085d0c95fb6cd3ea5b853018b5478c6c097426dff39eefa8d89dfaa72f187

Observation 27969eb6-8c88-47cd-8b7e-643059f52c85 · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.289106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:39167bc4c3ecc51627c97e62b0c31dd836e3b4f969c692fe7008c1ab59131727

Observation 5d04b888-c649-42c2-b7ca-ae6f7837d636 · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.884050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:8565cbe2c3db2a40360ff7e945fee82944de34525a1962f2ed06a45f0b7fa629

Observation 0e47b6b6-929a-44e0-b716-146985896a3f · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.295388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:9861932bf555e84667cb84273b60c900f4538acbf58f362fb364707e01fe9117

Observation 44a3994b-2af1-4c3e-9f49-65b04b8794c1 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:26.150859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:26.150859Z digest=sha256:dc1893f954ed60eb896c9f4eb0b653d7dd4132e5c00d37ff5ef873e0ae682ef3

Observation a6de5c47-ed61-469e-a188-460913426f03 · inbound

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents cites this paper.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.509228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.509228Z digest=sha256:50a807f3f1f859b3617381b9027edfbe6da6b409e14c2bdbd9d09b6a5c14baaa

Observation e01b077e-76b7-4348-b052-0bcf74aafbde · inbound

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs cites this paper.

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:38.022901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:38.022901Z digest=sha256:e7fe08874288bc4c9ae82ed974e3f2942827a8ee69584a225016b1dd81beec43

Observation 032d3af7-9043-4df6-94e3-156e805cafde · inbound

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation cites this paper.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.315097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.315097Z digest=sha256:c4a06505d08c034452caf1342ba4de2acf442433e81ca1e4787894b0786410c0

Observation b6ccbbc4-bf13-4618-b828-bd37700535d2 · inbound

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation cites this paper.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.144305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.144305Z digest=sha256:b7ea50fd5be1b6daa9c0a91753509407372b5c2ce25b96e0db41bc05d98e4354

Observation f9f149f2-ea60-4886-ad4d-d4362bbd38fb · inbound

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning cites this paper.

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T18:36:23.555649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:36:23.555649Z digest=sha256:6072fdd74ce11ddf0dff4d486e06e02dc5ad0bfc5bd87d384809d2c13653dc05