Pith. sign in

Paper Citation Record · LEDGER

ToRL: Scaling Tool-Integrated RL

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 85 inbound Pith citation observations for arXiv:2503.23383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23383 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 85 of 85 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.322456Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0ec26f3-934f-475c-a272-d028e2921326 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs ToRL: Scaling Tool-Integrated RL

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:42:39.069463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:54e55f946115d9651d32005c381dbd5132db4167439d7a5dd07cfd7dfdf45c07

Observation 25e6e5a0-2022-4752-86d8-c3a998fcf750 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering ToRL: Scaling Tool-Integrated RL

Reference 181

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.322456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.322456Z digest=sha256:c1350c7dc0f611b82b4a2c3ca41396663de4e9ff9f57417fbf21523fd6a6dbc6

Observation 8f1ff404-6ca0-451a-b0be-ea32b7af48b7 · inbound

AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset cites this paper.

AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:58:59.031839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:58:59.031839Z digest=sha256:508b216701a4e45c3b8eed6788b5f8f3dd122823750f7597bf1a5eb514d5e2f7

Observation 7df66722-4f73-4a07-886d-78b479392da2 · inbound

WebThinker: Empowering Large Reasoning Models with Deep Research Capability cites this paper.

WebThinker: Empowering Large Reasoning Models with Deep Research Capability ToRL: Scaling Tool-Integrated RL

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:14:25.328698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T19:14:25.283645Z digest=sha256:7b7ead639688e03fe872b09d145a9488039f7387da7c1059b67f177d1eb43a54

Observation 6c980726-5d3a-4f9d-9940-6fce536c2387 · inbound

Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving cites this paper.

Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving ToRL: Scaling Tool-Integrated RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:54.092671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:54.092671Z digest=sha256:5afbb9c7e1260004b5b81ced654da583362723209c5bb94cacc0daa77087f13f

Observation b2062e6d-11ff-4ab5-bf6d-f0fe946c1bf1 · inbound

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs cites this paper.

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs ToRL: Scaling Tool-Integrated RL

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:13.059774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:13.059774Z digest=sha256:c6cff27547c723d2b07f5aefca0854246056209a3135c53bb9b298f91fd3fea3

Observation 242dc923-cbeb-4a5f-93d5-edf32159a6d5 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.826306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.826306Z digest=sha256:72096da0cac7eb7ec3000be96bbf9121e985740aca2c71af0f8023599f272579

Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.394679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.394679Z digest=sha256:613ca9317e72cdc4e4da21920640dc2420c4f546cfdbff039cbadcf57bf37915

Observation f88a0921-bfd6-4689-b5af-88624a1ad566 · inbound

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation cites this paper.

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:16.110667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:16.110667Z digest=sha256:ecaec81836124a310eef74db9bb7487185fc3f47e503ea3b63347f97a42cedde

Observation 79512292-bbd6-4bc9-af93-958343527375 · inbound

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking cites this paper.

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking ToRL: Scaling Tool-Integrated RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:19.031069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:19.031069Z digest=sha256:823b8e7d5c88b9085fdee592b4e4049ca364d5808d6dfdb315dc94e8aacc3f31

Observation 5ce54d20-9228-4fce-bc41-0c6183dc2f59 · inbound

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers cites this paper.

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers ToRL: Scaling Tool-Integrated RL

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:20.079725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:20.079725Z digest=sha256:e163ab9c5a9ed3395feadd5679d25b713b50c1bfbd7973c4591edfec27bf187b

Observation 7dd91fe7-fa36-4dd1-b7cf-d4b0f6f84d79 · inbound

Reasoning LLMs are Wandering Solution Explorers cites this paper.

Reasoning LLMs are Wandering Solution Explorers ToRL: Scaling Tool-Integrated RL

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:12.497316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:12.497316Z digest=sha256:b6ea8a844f5f73b6fd2ff5739b8ddbf23cf7f897cc120f2efe6a3dce407145f1

Observation 91035334-31c4-4bf9-8900-7c926ce8639f · inbound

Towards Effective Code-Integrated Reasoning cites this paper.

Towards Effective Code-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:26.808955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:26.808955Z digest=sha256:c977e1eec6dcf4dff05e86234d8bb18f7b8a9993ab21c2c9e3ad08a83560c20d

Observation 0b82f9e9-4aab-4d41-aa95-17b48f455999 · inbound

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning cites this paper.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.503027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.503027Z digest=sha256:7a06a9b2d6b6f8d5ae866d42c76d334b6799d60aa120562da817377f74cb94fb

Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · inbound

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents cites this paper.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.665196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.665196Z digest=sha256:4d9430a1beca6561f32a1b2d86f4a6cf42d491fcc0d3993703c9f0b39405ad49

Observation aca1611a-9a7e-42cc-aac2-3e1b73269c53 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning ToRL: Scaling Tool-Integrated RL

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.327873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.327873Z digest=sha256:01d494a13027ae5a9f411ce9eab9c1abb6400691c5df669fa35eb10ed151470f

Observation 215f307f-beb2-4e89-8e06-91b84257bee5 · inbound

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following cites this paper.

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following ToRL: Scaling Tool-Integrated RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:18.797764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:18.797764Z digest=sha256:a1492a1b51558addd6888fe4d73dd4c7000f1b8d45ae4d669df81b21067b416d

Observation 601b0de8-2bed-4708-a7fe-8102af2e7d01 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications ToRL: Scaling Tool-Integrated RL

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:20.031722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:20.031722Z digest=sha256:b69ee48761bc84e743a3599321cc4e5ed2d3d5b4e9ea4d958bd8e0491ceca3f2

Observation be98b4be-dba1-44c8-bb83-0f574e20da63 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:31.307164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:31.307164Z digest=sha256:b43773dab32844ec51860bf02663731778921359b39f6a306e29102d3d4b3b2a

Observation 7c9e36bc-fb98-4b7f-90c5-0b5834acc2fa · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges ToRL: Scaling Tool-Integrated RL

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:04.518537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:04.518537Z digest=sha256:80a3af143b457324c6b0bee36bc623aac30029697f300b9f14e319b7e6e0d39f

Observation 5d1b38b3-89b2-4ed8-a746-8d265c087767 · inbound

Distilling Tool Knowledge into Language Models via Back-Translated Traces cites this paper.

Distilling Tool Knowledge into Language Models via Back-Translated Traces ToRL: Scaling Tool-Integrated RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:58.308371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:58.308371Z digest=sha256:e7ca381c2ed9538e5059de038716af6235381c3edc40c5555c12f5a6e41cf046

Observation 380fa6b6-651c-49c2-94f3-e4a6ef083491 · inbound

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization cites this paper.

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization ToRL: Scaling Tool-Integrated RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:12.930533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:12.930533Z digest=sha256:a8fb818631a780d7dbb1c873da0e4bd92aab93b9c16a1efc445cf5399dcc8ea1

Observation 05729999-5f89-4d04-a2cb-2cda13c47ab4 · inbound

StepFun-Prover Preview: Let's Think and Verify Step by Step cites this paper.

StepFun-Prover Preview: Let's Think and Verify Step by Step ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:38.217893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:47:38.217893Z digest=sha256:fefb0862fcaf3a08157d6524898d4ae8c6613904983e3d1c4e7a02e579cf99b9

Observation 4c924188-6d0e-4605-9afb-89eb632f15d7 · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.175839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.175839Z digest=sha256:10b371c7d635234c4b819d82c2d29837ad5cd74ea8564d5dfbaa902ad2eefcd2

Observation 77514045-5d44-4f40-be4c-22d915a50560 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.376258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.376258Z digest=sha256:de4979e44dea02aa662f2348f5a8f049921cebd2155c1f9bd1a71055ce04035c

Observation 7a0b5e96-f4a5-4364-a336-25872a9514fc · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ToRL: Scaling Tool-Integrated RL

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:29.999533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:29.999533Z digest=sha256:b8a54006d822bb5654a416e88edc86d985198c49273a628726117dd236c63934

Observation 243c78a3-5ade-4f3d-b10a-0bff989019c9 · inbound

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance cites this paper.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance ToRL: Scaling Tool-Integrated RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.597081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.597081Z digest=sha256:3b53857c516e829d6803d3f2fd7dd03aab8721091071e6fa5e67c6c38e40a4a5

Observation 0824de45-a673-4ba2-a2a3-3a113de407a4 · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.485908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.485908Z digest=sha256:6e6ffe6c0611078b305f500d46f7aaee67a811d763a76ea93770437914901341

Observation a7d18c52-045b-4390-a996-eb97c0cad2ef · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.584383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.584383Z digest=sha256:cceabda8d23522f2daf833534367ab7cd71258bb330e3615d324244c0f355152

Observation 6a3e08fd-da83-46f4-8a76-dfd42cea623b · inbound

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction cites this paper.

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:16.817954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:16.817954Z digest=sha256:43d26c73ab15d7c863efd8a01e5d79136acae871a726402b4ca4be9265edd3a7

Observation f5cc4b08-122f-4dbe-8bac-4bcd6feb5e3f · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.848125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.848125Z digest=sha256:95bae912ee7d497637863205426b94f343582bcd66b77b77bb3e800ae9b1198c

Observation 0b4a37c3-6a47-4f09-a775-3a90ffedb885 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.427791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.427791Z digest=sha256:fdba44308f36edbcac3cb6273122438749ef024a12204b55a41701a29d82f539

Observation aa12c784-d0c1-47d4-9e2a-41d1e08b1339 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.953751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.953751Z digest=sha256:511c59ac987121f8e91548d36a3440dad27439cc14b54bcc8c088825a6455ac3

Observation b9f49a87-c1ee-4848-adad-06a39be94915 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.036699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:7ad98ef669ae1d6c318a8e18af92f3792f06e8e591370df1d29ff37855759c04

Observation 7d4cf452-a9e5-41dc-835e-67e5a5324a0f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ToRL: Scaling Tool-Integrated RL

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.647828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:dc81d42f943cde006457a1e48ab5590235e6611d89f7a926fcf8b2a1aeca8709

Observation e44d5100-5f37-478c-8a4d-474f38a0c806 · inbound

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents cites this paper.

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents ToRL: Scaling Tool-Integrated RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:38.101554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:38.101554Z digest=sha256:e31f3e78bf36039ac4d8485422e5a22925eaabc6896fb10b2f41d8c0bb60b657

Observation 70fe5237-d565-4d4d-8b25-e3e90c7d63e2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models ToRL: Scaling Tool-Integrated RL

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.780909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:5ec3b1c92f4420817cfbd149c06cc1d3dcf76e55450fbc2838e1a40495f1f5aa

Observation ac6d0e7b-4034-494d-b8cf-bf0226db0283 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ToRL: Scaling Tool-Integrated RL

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.433285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.433285Z digest=sha256:7b80546abf25428848c087cbb7010663bae218666456df9067075ec6e013555c

Observation 4017163c-30ac-4cba-9d59-a1bcd7e95c4e · inbound

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions cites this paper.

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions ToRL: Scaling Tool-Integrated RL

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:01:31.394429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T15:00:51.162221Z digest=sha256:c93c9aa60ad56fac0edc8929f849ff594d6a6a79a897aaa1e804f39f57d56cd0

Observation f3054fae-3073-4eae-82ed-0095c656cfff · inbound

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination cites this paper.

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination ToRL: Scaling Tool-Integrated RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:10:51.356802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T04:09:50.183494Z digest=sha256:7c2c5c1594bc5a52f7f3cf568c31c60ce0273f0aba42d3cc5f54817912a7b3d1

Observation d543d618-6245-4541-a094-87a016794b19 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ToRL: Scaling Tool-Integrated RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:24.762093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:24.762093Z digest=sha256:e8acc520ee419cf91f6c79a293052d1751fa82a0da522bd1de4edbdb4e0bb7f4

Observation e8fec37e-6cc5-425d-950a-b976230eb3af · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation ToRL: Scaling Tool-Integrated RL

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.934885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.934885Z digest=sha256:0ad329ef10fd8b0d8d295e81ffab83b40301315dabfd77bee25fa1aa463191c8

Observation 4fb04403-fcde-49e9-9366-19131b72d040 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation ToRL: Scaling Tool-Integrated RL

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.765848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:514e3e1f214eb7df1bd099c95bfda6d1dd819d4f181e7cb0c5e3bf7e6813b4be

Observation 8ef8c2e9-57e1-49e3-8c4c-3e37788322b3 · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding ToRL: Scaling Tool-Integrated RL

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:08:12.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:13ae4f814b6811b8e0e22bf8c9cf9e07b1721b51a397405ae386328b2ab63c4e

Observation 706f8afb-266d-4440-9298-f4ad4fa4182c · inbound

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning cites this paper.

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning ToRL: Scaling Tool-Integrated RL

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:51.855593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:40:04.944348Z digest=sha256:667a16f31dc28109ba5d104ba67170e2bbb2e525183045a3da33ed8477225293

Observation 218d5a03-420e-463c-9b2f-0f0cf5ad4642 · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.826281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:d4cea009162cbe6a4df90eb36f88102fea5c868ce29dfa1ab0712037e390fefb

Observation 187e6633-ab0d-4b30-9cb5-7edc87a06037 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence ToRL: Scaling Tool-Integrated RL

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.216722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:025d181254c057f4e13acb72996ebe41a68f0d7a1adecc91da1a91ef17e0a6b1

Observation bcc06785-14f6-4835-beb0-663c880a8336 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.366674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:c166b70e650d808ed8734fc38cedcb76c07d28f8dbefaa75b7d8a2ed05dcb14f

Observation 2001a2f7-cab5-4e06-993e-153caaf1c8bb · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.761107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:fea9910a331c6d1d33a2756540f9fd1dedd2f02113b4e7635aba22c72130c7a6

Observation 7a1abce0-7186-407a-84da-7f16f046f19a · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.032153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.032153Z digest=sha256:f2dd8c69dd5072517087087a033be0c2849d9f7b9b79a5328018f282b2bbe6f6

Observation 8d5585b9-95a3-409d-81b0-3bb2dd92a7b2 · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:23.221704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:fbd7201ffaa53da26f8ff3493d78ab39cf6f099eb85ea72598d483b0aef6f67c

Observation fa50e7bb-92cb-4f0a-9e52-486162a28127 · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:23.444039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T04:42:47.496496Z digest=sha256:8cdcf61a135db98193a221664decb69fa77813cbd5d57e4a1c00627a87278cfd

Observation e38e9970-45c9-415f-8c2f-b14b8d94550d · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.260787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-21T08:08:04.544564Z digest=sha256:ae11a5097e18e4720e9e817000bb75ded27a5aa3e365a15ce7b0a43075e01a3d

Observation 160a545c-9bd6-4ab7-ba7d-e17491a1ca7a · inbound

Harnessing LLM Agents with Skill Programs cites this paper.

Harnessing LLM Agents with Skill Programs ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.372406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:26:19.382463Z digest=sha256:5a67a6d557e79d840a12d1c71c1a9fbb1a1b21178ff54277ce370ddbbdc65a46

Observation d07a6f08-b72c-426f-944b-2693cdfc2f61 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use ToRL: Scaling Tool-Integrated RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.597051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:4fac79ee39439928e641bd5e3137f677e5284f8545b113a6d7e447e4bf37a470

Observation 0cd270cd-f03b-4a05-b28a-9734d0724f15 · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ToRL: Scaling Tool-Integrated RL

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.097141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:268f0b2ba8b3b5bfdaca85e25788592622b4e796e4721bf941a24d85619657cc

Observation c528edd2-8f50-4d0b-aee0-6209ce4db83c · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ToRL: Scaling Tool-Integrated RL

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:23.709428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:c86bfa2140e14ec15392662eb3517912cbb2017c4030e05ae58cf937a19c035e

Observation 89728c42-fa89-46bf-94cc-d12ce2f5c9bb · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.298738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:0bba11f792f4bc8c1a721b6b5d48a5bda660bfe41e5a3dafd3f1429093c4f6fb

Observation 3385bd01-37da-4c78-8fe7-45aac4978f0c · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating ToRL: Scaling Tool-Integrated RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.740016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:943134b0299a87bc412cbda0ad3be8c007e50c821f30a453a8e902b4d4cb3165

Observation 527b852c-249b-406e-bffd-26a8d7b36d08 · inbound

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs cites this paper.

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.699154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T16:43:00.259139Z digest=sha256:9247e6b891daa007680ffddd2bd516f389cf93c1068cbf00a58da7500771bf93

Observation 8813d5a7-129f-4fa4-8f6f-300c3a5cb865 · inbound

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation cites this paper.

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation ToRL: Scaling Tool-Integrated RL

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.261630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T13:23:42.745788Z digest=sha256:af62503f3531731d561f3b681fd36f739e96b5148eb29fd92befc8148f775e06

Observation 89052348-3a9c-4647-871f-62b814bbd5d7 · inbound

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents cites this paper.

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents ToRL: Scaling Tool-Integrated RL

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.064660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T10:57:55.875707Z digest=sha256:257f4123612e57c41d681728da5bc7194630235181b6260f2ae8766f580e32b5

Observation 8a2e1419-c0c5-4db1-9992-b40160bc9eb4 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.403443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:899d114df779f551c26d1c983c8003e8187561e1e7cec4d58a89c754d1c3db12

Observation 73e0d108-c07f-4b76-95e9-cbeb0f34d7fa · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:28.554279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:28.554279Z digest=sha256:b84aacde169b14198e2c4ae6ed5f2e970fa51cdf0be64ebe5fa7bacdcdea4d00

Observation f3056329-2a84-46e7-a6ac-8d65e3d85169 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:33.209084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:d228ed49a3585562e02d2fdd1cd6051b3be2cb1f56dbc4e76a95ee8e92dc066a

Observation 206a8438-71f8-4e3a-8575-8313c113c327 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ToRL: Scaling Tool-Integrated RL

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.636889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:ac5c9b7fb0eb77ea6b396043c0508815ffe1649d69f2e63c10280f069a73ba82

Observation 23264379-35df-4bd4-bf00-2e6d31f696a2 · inbound

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It cites this paper.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It ToRL: Scaling Tool-Integrated RL

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.369980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-25T19:28:36.499352Z digest=sha256:3593170bf05be7e819b859f25bf1f3d65468c693bbcda0c71b1ce72360eee3f7

Observation baa37537-78f8-422b-b35c-ec386e6386c2 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.244054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T06:13:54.960194Z digest=sha256:403896c4d223332844698df09f9f45f38692fc1aedd6b83a0b755f9839e6c7b9

Observation 6ebedfcc-fc11-4d90-97d7-91509a9fdb01 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T07:12:34.663753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:12:34.663753Z digest=sha256:11b5639aedab2cf90bcfe13ad7fbd1bb866448f59c8a97838b7b352fb64ce6aa

Observation e3b45eb5-c77a-4597-bb91-d81577474950 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:12.937563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:12.937563Z digest=sha256:5ffc28165538135e686f62c1b6a645e110469c7b90195c13041f45fb3f3e4b41

Observation 85bdd02c-b90a-4538-b798-a0eae5c34095 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:05:44.909805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:05:44.909805Z digest=sha256:462d3e8ea7a7a5755562b8ffac1d20ff2710198e99018390362455d8af54d647

Observation bcee6077-162f-49d4-b9ec-4b8b61e6d5cd · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:36:25.292766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:36:25.292766Z digest=sha256:0d89627bf001803ccd4ef4878ea5ca3b80009f5bb23f75c19feb5d6a4556b6d5

Observation cbe8b569-abbf-4213-be6e-ee19d8d30bbf · inbound

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use cites this paper.

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use ToRL: Scaling Tool-Integrated RL

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.326575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-02T12:16:47.299349Z digest=sha256:3ecb80ee34d4de8bfbf68a228433ec6bd97c590729c8c114fc158e1542934d80

Observation b57a85d3-e8bc-44e0-b596-f4f14bc7aa84 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception ToRL: Scaling Tool-Integrated RL

Reference 134

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:d40a7f23a070330f3db7fce032b450a79ddbba44e4f70a2ffaa167217d2207b0

Observation bc658aa5-335e-456c-bbc8-569f751eb622 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b89f5819e1d10bbd604386ab6f3196f5f652419022a21c8c3ebb766d8675882f

Observation 012ef096-fe13-4c84-b804-db4c759c726e · inbound

Knowledge-Centric Agents for Workflow Generation in ComfyUI cites this paper.

Knowledge-Centric Agents for Workflow Generation in ComfyUI ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:12:53.833936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:12:53.833936Z digest=sha256:c642bacfc7eba05e4f008c56a8c4d5ea55067b7184e7605d2635e76a07f41b30

Observation e8d2fc97-d1e1-45c2-985b-edf2ec06025b · inbound

H$^2$SD: Hybrid Hindsight Self-Distillation cites this paper.

H$^2$SD: Hybrid Hindsight Self-Distillation ToRL: Scaling Tool-Integrated RL

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:49.091047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:49.091047Z digest=sha256:fd0f8a3df888ba91182d483bb388af2688e6c7bef5ddd227b816dba72d68b8d2

Observation f134a709-de0c-4f9d-9fbf-e0f22b86e1b4 · inbound

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents cites this paper.

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:56:53.271564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:56:53.271564Z digest=sha256:1e906238fa08688de2f042928308f05858430f70c9eb56ca7aa4cea32454078e

Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.645199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.645199Z digest=sha256:06e0e31bd969457103b73879134884a570140ac362bfc3d4cccb2df8a11b26eb

Observation 05877cd5-5ffe-4f1a-9b12-286a45b46ca7 · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.937398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.937398Z digest=sha256:10ca528c2fba39132474b971b94300f0dc9cc19a3434874460bba051f475a43f

Observation 927b7980-bf7f-4d01-b64b-5900562bc05e · inbound

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents cites this paper.

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents ToRL: Scaling Tool-Integrated RL

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:49:39.554840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:49:39.554840Z digest=sha256:f97e2fd40f82b7cf7898b7a54017bd628035086ca0701f0752c1b19cc0eddca6

Observation ffc6c40f-2636-40c7-a5d1-2b2f38aca1c8 · inbound

Contextual Information Policy Optimization for Search Agents cites this paper.

Contextual Information Policy Optimization for Search Agents ToRL: Scaling Tool-Integrated RL

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:44.357730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:27:44.357730Z digest=sha256:b9e7e7647610ecbd4764253a1f900cf4b6f769c20383e12b9b365a8b15e8d46e

Observation 8661dab7-fd78-44f8-99d9-6c81b7dbc4ff · inbound

Contextual Information Policy Optimization for Search Agents cites this paper.

Contextual Information Policy Optimization for Search Agents ToRL: Scaling Tool-Integrated RL

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T00:54:15.445890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:54:15.445890Z digest=sha256:2a2127094457eef02f40919a5d502f7bf54990582282b8bcd631769cfd8046dc

Observation 3c61b79c-030c-42a0-88ff-cbac19d0b4fb · inbound

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing cites this paper.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToRL: Scaling Tool-Integrated RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.633077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.633077Z digest=sha256:bf29186384dddfdf324053c5a528b440f76158f62ce55c99999230f355553c6d

Observation cfeb87f9-c46e-4edb-8262-dfd60b8b77e8 · inbound

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents cites this paper.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.120450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.120450Z digest=sha256:742a5b51e9c3e7ca8dc7aa47027a745dc42617fddf40f6914ff0e222250af45a