Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:08.486342Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 38 inbound Pith citation observations for arXiv:2505.16410.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:08.486342Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:26:56.575865Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
82 of 82 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation a768c1a7-f587-45d9-afb8-9244aff5719c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a52689b-1077-4ebc-ab31-0003e6de246c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f52244-f3d1-4478-92df-621630b3f6d6 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d50b1f-5230-48c6-9bf5-b7e9303ae65f · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0db75b-ba54-40fd-a9c3-04d5116bb108 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2a9b67-4572-412c-9ed6-5b8da5fb3627 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in small llms: What works and what doesn’t
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db8e6d8-9076-4853-be05-13d988fca364 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Flashattention-2: Faster attention with better parallelism and work partitioning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a01c378-6eea-464c-945e-8b952fd37e6d · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463ec17e-ac30-4bdb-8ea2-81a9d406ba19 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59afab32-de7d-400d-9f4e-bda9cdd99418 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning How abilities in large language models are affected by supervised fine-tuning data composition
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534947a6-a771-46e7-89c1-5c0506bb8b8a · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Progressive Multimodal Reasoning via Active Retrieval
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b45d04-ff58-4c3c-9d94-db92a37592e5 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a8ccf65-cd6a-4971-93df-8ea0870efad9 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 913923bf-d874-4927-8775-ad0923d608cb · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Concise reasoning via reinforcement learning, 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7aea587-87f7-4ccd-9cd7-cf89a904e6da · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retool: Reinforcement learning for strategic tool use in llms, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23cafa84-7db2-4b24-bf89-2225134fdcf0 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Tora: A tool-integrated reasoning agent for mathematical problem solving
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0ec14f4-adee-4938-ae80-09f21d6f692b · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Measuring mathematical problem solving with the MATH dataset
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8458cacf-0f84-45b9-8de3-68df07e66012 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8588b63-1591-4eb6-bc10-acab3f7332ea · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9160051-d954-4e11-b638-b95197f87cbb · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Towards reasoning in large language models: A survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cddfc57-598a-4628-8a98-68a78fd6e4d9 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bffe5425-3732-4e3f-9c2d-fde963dae4b9 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ed383c-4019-413e-a777-e91fd6877bd9 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 200f13a6-785c-47f8-b91a-5eff891f837a · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8416c664-e224-46e3-a43d-0bdebc16618d · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retrieval-augmented generation for knowledge-intensive NLP tasks
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation daa241b4-b288-4055-81b0-54ead7f9de7c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f7b39c-0cb8-4f13-a9ed-05c22fd440b6 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning START: Self-taught Reasoner with Tools
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283ff077-586d-43e1-8ac4-58ba3cee46c4 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Chain of code: Reasoning with a language model-augmented code emulator
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57c7b695-3dee-4172-bfcc-48b6d2e5a6fe · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04539f72-ae22-48f5-b753-e1a7d0010b5e · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3575f901-c584-4012-9934-210449d77720 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745a06c8-a473-4d00-abe6-0858cc0119fd · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aea5fb85-94ca-4a23-9392-ebdc726e520c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Let’s verify step by step
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b7fd29a-5863-4e42-96ef-e1de6797f489 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 445726c9-8530-4733-8eca-b261283df9e9 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning GAIA: a benchmark for general AI assistants
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93a9665b-d25d-40fa-bd45-651c6f9aef19 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf84693b-90e0-48cd-a54a-a6034290062d · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Learning to reason with llms, September 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3144e645-50d1-48ce-885a-71094a43d092 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ART: Automatic multi-step reasoning and tool-use for large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3d20e1-635b-4c5d-a32f-33d56633289f · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Humanity's Last Exam
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8539a7f-3075-48f5-bbce-6105aff6cb86 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Smith, and Mike Lewis
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5233d9b-b8cd-4656-88ee-e72349f1e852 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7298ef0-0a11-46c6-a798-9ec61ac14a85 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Toolrl: Reward is all tool learning needs, 2025
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a372f50-8758-4b4f-b9b3-60ce4a8d6509 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0adb45c0-3add-4d97-9c99-26536730064d · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning O1 Replication Journey: A Strategic Progress Report -- Part 1
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2997ecbc-83d1-4c34-bf61-4ee3967d1677 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5 technical report, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e10529-5f86-489a-8250-4815c38c2293 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0e5f2f0-0796-41dc-aa5e-5b70c7d0bdbd · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a5d9b7-a885-40fe-b01c-681f3acdd004 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93374031-0617-423a-82a1-ccf6d4d99499 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fc324cf-9944-49e5-aea7-6b77dd04f0ab · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Curriculum learning: A survey
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1453fd0-e7d7-43f6-806d-6f5e890cfc88 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c6dfe8-d0b4-4b8e-8f90-86b78c232e7b · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Zerosearch: Incentivize the search capability of llms without searching, 2025
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e6190f-f8ce-4ad7-b7b1-f53ae29ec45d · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning A Survey of Reasoning with Foundation Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a456f2-71ba-4ab3-acf7-9470f5701e26 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c61a8ef-fdc0-4fec-bd37-f51574ecfb16 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4530ae1-730f-487c-b9de-cd35b72a215c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e086096-b12f-4f26-82b3-691746961c7b · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41c61fc1-3617-4a5e-a0f4-9210c9366594 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 291d80f0-80dc-4ebd-a067-e19d68876a68 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning musique: Multihop questions via single-hop question composition
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3927b1f-8cf6-41e1-a29e-5b85e26a419c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd5b097-0a92-431b-afce-308df4f64937 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Text embeddings by weakly-supervised contrastive pre-training, 2024
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6865c8fb-0fe2-458f-82c7-17d3d5577937 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in large language models with one training example, 2025
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6d98e7-73d1-4a9f-823c-a9a0497ebfc8 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce961658-b063-4860-8900-c03507313a72 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc269afc-78f1-4380-933b-e35f5fe7a0d2 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc8f876-fd04-4022-a5c7-510118aae02a · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2 Technical Report
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b498f759-afe6-4708-9ac1-e06713c18d97 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307a74a1-7a9a-47b0-b6d7-5d55cb513d6a · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e548cbf4-4dc3-4a94-8cd7-2d6e1c820185 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81248e37-f9ac-4e80-a775-969a6c60d58d · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMO: Less is More for Reasoning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d0aa70-6d7d-4dfd-a556-9709db7a630e · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60b9b13-4423-4230-8314-19ace93877e5 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ee74034-adfa-4c42-8c6b-7a8371bd8dbd · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981b328b-5cad-4846-a3ff-9e4652bff27c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d58cba6d-1eda-493c-91ec-c37fb3e07075 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba4b92eb-7b3e-4d32-9517-c76a3d75a398 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agent models: Internalizing Chain-of-Action Generation into Reasoning models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23835f8d-1b01-4f55-8421-2fea2308e4d4 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98931d4b-53cc-4d13-815c-8d322e3a3bf0 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours)
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be2b3f86-a863-4906-a4a4-aec64f33799c · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Converting t minutes to hours, we get t 60
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e10bd72b-7cc0-464e-9a4c-f2c71d769f71 · outbound
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning {numerator}/{denominator}
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26238ddb-55ef-4b2a-9f10-2735e6547f87 · inbound
Deep Research Agents: A Systematic Examination And Roadmap Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57099b37-beae-4b05-a5c4-9e9ca5be68b5 · inbound
Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8249ae1b-0f50-430e-a455-2913c4bae6bd · inbound
A Survey of Context Engineering for Large Language Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 231
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61e64d85-0ef9-4e7b-8d37-94c49eb4807c · inbound
AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec53e26-e08c-4eaa-b6f8-4349080788a5 · inbound
MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8e51c0-4b72-44f0-baa2-9ff8928beed1 · inbound
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d12bb68-7528-4280-9e78-e7c7d7d238ba · inbound
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c997f8d-da85-4462-888f-99a8a655b5c8 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4cf73df6-8d84-4a8c-9504-70b53a7eb141 · inbound
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c39a617e-8a38-456c-a1b9-3b25bf595665 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53bfd2c-ef7d-496d-9924-bbd31e9e9dc6 · inbound
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72f33c36-2fab-44a6-93a7-49a1d4ea5b14 · inbound
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca8698f-4717-4727-b186-7865c2f30559 · inbound
Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eee2387-aef1-4049-93d4-754602ccb4fc · inbound
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 782c5788-7aec-4c3c-ac47-15f9e354c63a · inbound
From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 201
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 084cdba8-6810-472e-b25e-6bcfd79a999d · inbound
Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31bd0abe-80a2-4ab8-9381-690740a6df9a · inbound
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bee09b4-f1f4-4348-9163-07973725b8ee · inbound
Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9ac6733-1ad4-4aa6-b9d9-15bdb1d04c7b · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 407fad62-21f3-4e68-b6b4-9bf344089a81 · inbound
Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c63639f7-9229-49a5-8728-2f23fb64545b · inbound
Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bab25d3d-c1ec-460f-8e49-fc9827819e64 · inbound
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2154c720-8d36-4496-9cab-3c05b7e3268e · inbound
PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · inbound
Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a0d90ec-834b-4ada-b8af-32f410603583 · inbound
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 707e0454-297d-4895-a8f1-af98616163da · inbound
Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 898960cb-a66d-4f3e-8cb9-5dd058afa2b4 · inbound
Trust Region On-Policy Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 639eeddd-856a-470d-8cbe-1fd351a1493e · inbound
Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80e8188d-8c77-4e2a-b6bc-3d98bda5218c · inbound
ToolFG: Towards Well-Grounded Fine-Grained Image Classification Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27969eb6-8c88-47cd-8b7e-643059f52c85 · inbound
Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d04b888-c649-42c2-b7ca-ae6f7837d636 · inbound
When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e47b6b6-929a-44e0-b716-146985896a3f · inbound
APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44a3994b-2af1-4c3e-9f49-65b04b8794c1 · inbound
APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6de5c47-ed61-469e-a188-460913426f03 · inbound
TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e01b077e-76b7-4348-b052-0bcf74aafbde · inbound
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 032d3af7-9043-4df6-94e3-156e805cafde · inbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ccbbc4-bf13-4618-b828-bd37700535d2 · inbound
Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9f149f2-ea60-4886-ad4d-d4362bbd38fb · inbound
ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.