Pith. sign in

Paper Citation Record · LEDGER

ToRL: Scaling Tool-Integrated RL

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 77 inbound Pith citation observations for arXiv:2503.23383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23383 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 77 of 77 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:49:39.554840Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0ec26f3-934f-475c-a272-d028e2921326 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs ToRL: Scaling Tool-Integrated RL

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:42:39.069463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:a1fa6c5b429d6618b1eae94f635a14215898bb33e6b278716c7c32b900a6433b

Observation 7df66722-4f73-4a07-886d-78b479392da2 · inbound

WebThinker: Empowering Large Reasoning Models with Deep Research Capability cites this paper.

WebThinker: Empowering Large Reasoning Models with Deep Research Capability ToRL: Scaling Tool-Integrated RL

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:14:25.328698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T19:14:25.283645Z digest=sha256:2cd418ecda92d9507c40920c05cce8fb0490fc6819dca4798387346bc3b0efff

Observation 242dc923-cbeb-4a5f-93d5-edf32159a6d5 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.826306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.826306Z digest=sha256:e9b54a5af929689e8c2fd2ce76a2f24335ad650506a469ad75caadc0127ec551

Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.394679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.394679Z digest=sha256:e69bfa0a48d3cc69ef966eed9a83b0c13df2f2c410eb91b7f879b41eaac99213

Observation f88a0921-bfd6-4689-b5af-88624a1ad566 · inbound

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation cites this paper.

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:16.110667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:16.110667Z digest=sha256:22a5a1dd079c531854363386eda2d693565f893e58c108aebbac2badaddb9b6d

Observation 79512292-bbd6-4bc9-af93-958343527375 · inbound

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking cites this paper.

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking ToRL: Scaling Tool-Integrated RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:19.031069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:19.031069Z digest=sha256:15086349adf8b8292616ddf2418691f7a0b744a167a2e49222f8291c945a27cd

Observation 5ce54d20-9228-4fce-bc41-0c6183dc2f59 · inbound

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers cites this paper.

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers ToRL: Scaling Tool-Integrated RL

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:20.079725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:20.079725Z digest=sha256:b1cab690b9b71caa8a2073c55f639fff19cbf75140b8b34f108c95f1aee40881

Observation 7dd91fe7-fa36-4dd1-b7cf-d4b0f6f84d79 · inbound

Reasoning LLMs are Wandering Solution Explorers cites this paper.

Reasoning LLMs are Wandering Solution Explorers ToRL: Scaling Tool-Integrated RL

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:12.497316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:12.497316Z digest=sha256:6dd318feb0ab1de4d1b7b799fa771f56e6f07f1883a4588bad8727583632802d

Observation 91035334-31c4-4bf9-8900-7c926ce8639f · inbound

Towards Effective Code-Integrated Reasoning cites this paper.

Towards Effective Code-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:26.808955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:26.808955Z digest=sha256:645b61435fd778e57e2adf48ddfff078e0ff7c065418544380451f2d4d4437bb

Observation 0b82f9e9-4aab-4d41-aa95-17b48f455999 · inbound

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning cites this paper.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.503027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.503027Z digest=sha256:8bf8d3228d2b8b784db268386c9ccc67b6fca27c53f027becb9696bee072cd76

Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · inbound

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents cites this paper.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.665196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.665196Z digest=sha256:d739c7f109fadbaa96c0f6ca58e226091c6ce54a569dc653fd9de811cb04dbdb

Observation aca1611a-9a7e-42cc-aac2-3e1b73269c53 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning ToRL: Scaling Tool-Integrated RL

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.327873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.327873Z digest=sha256:522a48e938d75d9c908adb7f4dc0efc7664f3959e7545af07f9e62bcad644850

Observation 215f307f-beb2-4e89-8e06-91b84257bee5 · inbound

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following cites this paper.

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following ToRL: Scaling Tool-Integrated RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:18.797764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:18.797764Z digest=sha256:f112158ef63fee4d8952d8877c13e4ea4b17e873077c74f4cbee1b290b77a3ee

Observation 601b0de8-2bed-4708-a7fe-8102af2e7d01 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications ToRL: Scaling Tool-Integrated RL

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:20.031722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:20.031722Z digest=sha256:0947eea22f5af56978f06eb80e607acf2e5cfef9d79926ec428db61334cf9f04

Observation be98b4be-dba1-44c8-bb83-0f574e20da63 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:31.307164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:31.307164Z digest=sha256:3e640a22d68cdf354b046b3724cbba5206ad683101281c3417a8d83d25d182ac

Observation 7c9e36bc-fb98-4b7f-90c5-0b5834acc2fa · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges ToRL: Scaling Tool-Integrated RL

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:04.518537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:04.518537Z digest=sha256:571e6e347929a1e3d94af2a3d40d9d56cbce3691aa75c90bb0cf9d22c6b301cc

Observation 5d1b38b3-89b2-4ed8-a746-8d265c087767 · inbound

Distilling Tool Knowledge into Language Models via Back-Translated Traces cites this paper.

Distilling Tool Knowledge into Language Models via Back-Translated Traces ToRL: Scaling Tool-Integrated RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:58.308371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:58.308371Z digest=sha256:c6c6c4064983e71abdde237ff3f9c74809d2249a28e1755e4beb39e40ff4ad13

Observation 380fa6b6-651c-49c2-94f3-e4a6ef083491 · inbound

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization cites this paper.

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization ToRL: Scaling Tool-Integrated RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:12.930533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:12.930533Z digest=sha256:a515b0d4f8fd823b80023a116ead2e3b50887d7b9501673d993a14b931afd4dc

Observation 05729999-5f89-4d04-a2cb-2cda13c47ab4 · inbound

StepFun-Prover Preview: Let's Think and Verify Step by Step cites this paper.

StepFun-Prover Preview: Let's Think and Verify Step by Step ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:38.217893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:47:38.217893Z digest=sha256:2c0dcfe888a204ec9811d99aff4180abf001394dac2a8d0f940b126012ba75a2

Observation 4c924188-6d0e-4605-9afb-89eb632f15d7 · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.175839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.175839Z digest=sha256:349dfc2a4b4e83b97e6d9bb62506d157b778c82d994143ced5df314d080ae94a

Observation 77514045-5d44-4f40-be4c-22d915a50560 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.376258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.376258Z digest=sha256:aba3ea8aaf96dc100977ab9c45adb464769b346b81935136c753d6e92621c316

Observation 7a0b5e96-f4a5-4364-a336-25872a9514fc · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ToRL: Scaling Tool-Integrated RL

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:29.999533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:29.999533Z digest=sha256:0aaa28096f512f453f757602ecbdca2df88ee388d160ced5ac75cc9b6eb711b7

Observation 0824de45-a673-4ba2-a2a3-3a113de407a4 · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.485908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.485908Z digest=sha256:9a8c0aefdf3c1d75a651b9dac84703a2e0a2ea4ae74dba4860ef0b32c561ff0d

Observation a7d18c52-045b-4390-a996-eb97c0cad2ef · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.584383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.584383Z digest=sha256:995b6457f708ee151afbdfd9c26e8d100f1350a0d83ee29ce11970c7af2e8977

Observation 6a3e08fd-da83-46f4-8a76-dfd42cea623b · inbound

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction cites this paper.

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:16.817954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:16.817954Z digest=sha256:28a5c93d0a9fb7571768dba108ca3b70397bd1514c9a4bd5a33689e6218feb11

Observation f5cc4b08-122f-4dbe-8bac-4bcd6feb5e3f · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.848125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.848125Z digest=sha256:3a0486081bcd11ab7ebbfedbb1658840a262a250e962da8bdc7103bc32cd2e32

Observation 0b4a37c3-6a47-4f09-a775-3a90ffedb885 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.427791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.427791Z digest=sha256:91324ff5eba0ba29d9405f2c42b099d1b673ce6c1dae3bdd457a7bec5fbfad8e

Observation aa12c784-d0c1-47d4-9e2a-41d1e08b1339 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.953751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.953751Z digest=sha256:414775970abbadb4a01b92da0d86fbd3eb6c771fa45cd0ec0447dfa2b719e539

Observation b9f49a87-c1ee-4848-adad-06a39be94915 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.036699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:0a9d155c8d3fa71b388a587dacf5b66515664212530bff41fe88f88733e03c1f

Observation 7d4cf452-a9e5-41dc-835e-67e5a5324a0f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ToRL: Scaling Tool-Integrated RL

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.647828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f8a1d7dbabaae6a11e56156ec4ab86d43964f778b75538bffc6448f74d7b0cee

Observation e44d5100-5f37-478c-8a4d-474f38a0c806 · inbound

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents cites this paper.

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents ToRL: Scaling Tool-Integrated RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:38.101554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:38.101554Z digest=sha256:2a12c51e15728e118cd2d97fa47c05894f7827ff8457695ea9e5a3ca5e7937c9

Observation 70fe5237-d565-4d4d-8b25-e3e90c7d63e2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models ToRL: Scaling Tool-Integrated RL

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.780909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:41d34e679564c60877d0e666841b969d3f16957d6ef7d1f705cd7df9e3224e75

Observation ac6d0e7b-4034-494d-b8cf-bf0226db0283 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ToRL: Scaling Tool-Integrated RL

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.433285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.433285Z digest=sha256:e4b397cfb6cbf6aa3179e74eff2c5655266f42e758be3b545128fe85d5fe9ad2

Observation 4017163c-30ac-4cba-9d59-a1bcd7e95c4e · inbound

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions cites this paper.

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions ToRL: Scaling Tool-Integrated RL

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:01:31.394429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:00:51.162221Z digest=sha256:6270b343114f145861dd98e1f3fa1935e7b1192640faba5caaefe85712864f3f

Observation f3054fae-3073-4eae-82ed-0095c656cfff · inbound

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination cites this paper.

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination ToRL: Scaling Tool-Integrated RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:10:51.356802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T04:09:50.183494Z digest=sha256:46a70ed186234efcc580b15ef4644b5dfe4146a3d4fd8cdc9ba2b70ebd2e2f8d

Observation d543d618-6245-4541-a094-87a016794b19 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ToRL: Scaling Tool-Integrated RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:24.762093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:24.762093Z digest=sha256:0827994082ed44f40f0f17fdbe45a892568d37caf339c201ff189b8d537b8809

Observation e8fec37e-6cc5-425d-950a-b976230eb3af · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation ToRL: Scaling Tool-Integrated RL

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.934885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.934885Z digest=sha256:9953c5bb235892319ca2cceabeeacd8eb407f3c74f844fa3425ad1d92d61ade6

Observation 4fb04403-fcde-49e9-9366-19131b72d040 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation ToRL: Scaling Tool-Integrated RL

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.765848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:ba5786a7dad204f7d274761a232fe3dc97e38606114e07c0c5a901c535149632

Observation 8ef8c2e9-57e1-49e3-8c4c-3e37788322b3 · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding ToRL: Scaling Tool-Integrated RL

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:08:12.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:ff034f06a32a683bdcefd40e529eee16356ecf7dec0afe53a03cd9470a022823

Observation 706f8afb-266d-4440-9298-f4ad4fa4182c · inbound

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning cites this paper.

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning ToRL: Scaling Tool-Integrated RL

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:51.855593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:40:04.944348Z digest=sha256:1006d65d51acd6d3cf714ac1daad10f29f233149f6f52cf83f53e6265a5562be

Observation 218d5a03-420e-463c-9b2f-0f0cf5ad4642 · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.826281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:0a7daf91a7804132622da4ab49c20556477bc4c8a327497116b2f1167bc40922

Observation 187e6633-ab0d-4b30-9cb5-7edc87a06037 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence ToRL: Scaling Tool-Integrated RL

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.216722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:ff0091fe50996fe7a7305fa7dc712738ac020e20c2d886e64a135967eab325d4

Observation bcc06785-14f6-4835-beb0-663c880a8336 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.366674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:c585f139b5bd69464f7b63cc06f6de0b221ab2c4dd6ca8e50959ed13608fd25e

Observation 2001a2f7-cab5-4e06-993e-153caaf1c8bb · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.761107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:3f2d473479ca4b5c05e941844ac4b711171c92a7a667f3fa9869c76c7ae94dc9

Observation 7a1abce0-7186-407a-84da-7f16f046f19a · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.032153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.032153Z digest=sha256:3ab26841753eb5d735ebb13f874186d74bd0cf0931e4daa3b384acefcaf4da22

Observation 8d5585b9-95a3-409d-81b0-3bb2dd92a7b2 · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:23.221704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:f82ff9b7a6b306f8d62e94bbf5aff86fb4bea63590f8db1f8a3b314d9530b598

Observation fa50e7bb-92cb-4f0a-9e52-486162a28127 · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:23.444039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:42:47.496496Z digest=sha256:86362e126a045384cbcab6b5b4c46e3bde5cec1c7f4f69d80208af9f5da2f256

Observation e38e9970-45c9-415f-8c2f-b14b8d94550d · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.260787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T08:08:04.544564Z digest=sha256:595adc6d847437687e290c7f754d5db2bb7555e04771e9e494bd569a198a83db

Observation 160a545c-9bd6-4ab7-ba7d-e17491a1ca7a · inbound

Harnessing LLM Agents with Skill Programs cites this paper.

Harnessing LLM Agents with Skill Programs ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.372406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:26:19.382463Z digest=sha256:6d5565f236b369796c05438af2477f91332294f29ae180273efa2de525b4e8f1

Observation d07a6f08-b72c-426f-944b-2693cdfc2f61 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use ToRL: Scaling Tool-Integrated RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.597051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:0237be8b6744a0569bc86b4d809c8bc3c3a44effd0a39711d09f3187767d9ed1

Observation 0cd270cd-f03b-4a05-b28a-9734d0724f15 · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ToRL: Scaling Tool-Integrated RL

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.097141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:9679eae18790ec9b9e391826a7f2ba0ca7010e26606e1ee9998dbd1e051fb7f2

Observation c528edd2-8f50-4d0b-aee0-6209ce4db83c · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ToRL: Scaling Tool-Integrated RL

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:23.709428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:cdfc0bd6006618ad861710a1bc0476a2e930e7dc76585ffdba03ed8cd5a603de

Observation 89728c42-fa89-46bf-94cc-d12ce2f5c9bb · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.298738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:c7ce1cffa18710ef4765397a2f8a8598eb45990e52763a209ff848b6f301e2ff

Observation 3385bd01-37da-4c78-8fe7-45aac4978f0c · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating ToRL: Scaling Tool-Integrated RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.740016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:fda2d1d75985931f08ae03b453933b803a63b12ffc72f049c8e76166b3547c56

Observation 527b852c-249b-406e-bffd-26a8d7b36d08 · inbound

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs cites this paper.

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.699154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:43:00.259139Z digest=sha256:3ae9fffe8b2587dc3f13b6ded2ee4ba75b742a353cf00223f57213fcd7557a6d

Observation 8813d5a7-129f-4fa4-8f6f-300c3a5cb865 · inbound

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation cites this paper.

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation ToRL: Scaling Tool-Integrated RL

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.261630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:23:42.745788Z digest=sha256:b92918d7f82f62cb098b3427ff4c56e7d83f5257fd0d66044a82a850879dddc3

Observation 89052348-3a9c-4647-871f-62b814bbd5d7 · inbound

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents cites this paper.

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents ToRL: Scaling Tool-Integrated RL

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.064660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:57:55.875707Z digest=sha256:1647910ca7e2e7a28fe92b9e2c1164471c26c46a7dedb7444e00ecdf43975b13

Observation 8a2e1419-c0c5-4db1-9992-b40160bc9eb4 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.403443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:19ff6b775874fecf54c5e3019fb7f3feb390c2fa984c7d619216bb15d07fb671

Observation 73e0d108-c07f-4b76-95e9-cbeb0f34d7fa · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:28.554279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:28.554279Z digest=sha256:02cacc1c03a04379660505cda464b49e522e19d39d580539858fd4ee821d6297

Observation f3056329-2a84-46e7-a6ac-8d65e3d85169 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:33.209084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:a91f110ee87c61dace71142b8b1f9a9db09311e2b182ff6ac0304c5a83aeec29

Observation 206a8438-71f8-4e3a-8575-8313c113c327 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ToRL: Scaling Tool-Integrated RL

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.636889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:68d3933e375c39ec3160f273f3c91a31f7e3b59331a011c88b9fe27c75a86eab

Observation 23264379-35df-4bd4-bf00-2e6d31f696a2 · inbound

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It cites this paper.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It ToRL: Scaling Tool-Integrated RL

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.369980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:28:36.499352Z digest=sha256:2b70d450e3e361b70edb4c9ad5ec89e93efcd967898bbb884d6051760d381e98

Observation baa37537-78f8-422b-b35c-ec386e6386c2 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.244054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:13:54.960194Z digest=sha256:f9e23084ac62fbb29f3083cecbd44ca3faf617ba3513a6a72cb75e7b0acddb7b

Observation 6ebedfcc-fc11-4d90-97d7-91509a9fdb01 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T07:12:34.663753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:12:34.663753Z digest=sha256:b0e786f97cea864a2bbc09eb3d88a638e5362de10ec904e7387a54002278f28f

Observation e3b45eb5-c77a-4597-bb91-d81577474950 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:12.937563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:12.937563Z digest=sha256:d60225291f69a94e8b5fcdcad020d6879d3aa3ac202c57317959124109653fc3

Observation 85bdd02c-b90a-4538-b798-a0eae5c34095 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:05:44.909805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:05:44.909805Z digest=sha256:90109fa91d92f01f254b63fcd89a7d7c53b594ce4ccfde5c6fa1b64a14fc02f4

Observation bcee6077-162f-49d4-b9ec-4b8b61e6d5cd · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:36:25.292766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:36:25.292766Z digest=sha256:9a10687fdbcf0343641db34a569b58a70e8c88f5a95f665987f4d102a5dda462

Observation cbe8b569-abbf-4213-be6e-ee19d8d30bbf · inbound

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use cites this paper.

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use ToRL: Scaling Tool-Integrated RL

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.326575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T12:16:47.299349Z digest=sha256:ebf18f236abac181663f43310a8234ad5afcf3fce285b3a97f363fe9aeb26381

Observation b57a85d3-e8bc-44e0-b596-f4f14bc7aa84 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception ToRL: Scaling Tool-Integrated RL

Reference 134

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:069733d3ed8108675968cb9a6d0c9612765cda7a194f5ed8f115643a534e5fd4

Observation bc658aa5-335e-456c-bbc8-569f751eb622 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:460375a8c56b9f6544dc413566c47d395a6bf82a55192dbb25036b95f67e303b

Observation 012ef096-fe13-4c84-b804-db4c759c726e · inbound

Knowledge-Centric Agents for Workflow Generation in ComfyUI cites this paper.

Knowledge-Centric Agents for Workflow Generation in ComfyUI ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:12:53.833936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:12:53.833936Z digest=sha256:ed2e86c52835e5f35358209445d8b505ee3f9d5679f486f8854d7706e32c0aac

Observation e8d2fc97-d1e1-45c2-985b-edf2ec06025b · inbound

H$^2$SD: Hybrid Hindsight Self-Distillation cites this paper.

H$^2$SD: Hybrid Hindsight Self-Distillation ToRL: Scaling Tool-Integrated RL

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:49.091047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:49.091047Z digest=sha256:bfc4164aa6ac9d581f5d0b50167941f57eda9ff25dec1c6e0dd8f52d3d015f34

Observation f134a709-de0c-4f9d-9fbf-e0f22b86e1b4 · inbound

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents cites this paper.

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:56:53.271564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:56:53.271564Z digest=sha256:38f15c4e41c086449b5b32d1fb2eaa40af06ed6be95a9f896be98713ccc54bb1

Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.645199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.645199Z digest=sha256:60fb1e7064837d55574518dc3bbebd087655e69a9c589ed487ceb6177f9ece5e

Observation 05877cd5-5ffe-4f1a-9b12-286a45b46ca7 · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.937398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.937398Z digest=sha256:ac8e10c7fe3c2bd569e12a65f227e66653f7c12051f59f3afe60322049f1933e

Observation 927b7980-bf7f-4d01-b64b-5900562bc05e · inbound

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents cites this paper.

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents ToRL: Scaling Tool-Integrated RL

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:49:39.554840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:49:39.554840Z digest=sha256:7dba6b764263d16fb29a298925d91308388b070bbc0ba0f571ca8413d654374c

Observation ffc6c40f-2636-40c7-a5d1-2b2f38aca1c8 · inbound

Contextual Information Policy Optimization for Search Agents cites this paper.

Contextual Information Policy Optimization for Search Agents ToRL: Scaling Tool-Integrated RL

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:44.357730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:27:44.357730Z digest=sha256:e50b9f696807ed5ec7a166f20858c24c75739b3f923c53712d18bc4554363c05