Pith. sign in

Paper Citation Record · LEDGER

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

As of 19 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 6 inbound Pith citation observations for arXiv:2509.06980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06980 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:07:18.380838Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:00:17.278040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:16:57.747266Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c0e1a07-587d-4e63-88ab-a40d9ff44466 · outbound

This paper cites Introducing gpt 5.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Introducing gpt 5

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.485775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:07:18.350381Z digest=sha256:aeeab1319a1f6b527508dab5155d666fe9d5e3b1acee1e391008594f8a5d396e

Observation 4e916ce5-dffe-4c26-8060-2fc22b9698e7 · outbound

This paper cites Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agent rl scaling law: Agent rl with spontaneous code execution for mathematical problem solving, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.477442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:07:18.354025Z digest=sha256:1170771f701e5cd910cf98aad81ce9c5ed536009bec241b153c710d51a87fe67

Observation 4a0e9739-c96a-49fe-9c7d-c669497d230e · outbound

This paper cites Agentic reinforced policy optimization, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reinforced policy optimization, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.357017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.357017Z digest=sha256:b5c73a77227d9901360c8b18dc63d8ad4875c272c15b7db17f2947008bd3a8f1

Observation 3357afd6-2c30-499e-b906-b4dd26856800 · outbound

This paper cites Agentic reasoning and tool integration for llms via reinforcement learning, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Agentic reasoning and tool integration for llms via reinforcement learning, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.463675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:07:18.359940Z digest=sha256:49726fae17857a2ecd868a430d93957fee23c95197527329a6f79184639982ea

Observation 536182d1-6685-4491-9134-d11dd6372bee · outbound

This paper cites Search-o1: Agentic search-enhanced large reasoning models, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-o1: Agentic search-enhanced large reasoning models, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.362922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.362922Z digest=sha256:94fee87276c125b8dc61d5d3c373a2062ee9a510d23c5f87b4dbbb2534faf807

Observation 525d9313-2543-4522-9255-460b0002a9f0 · outbound

This paper cites Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Search-r1: Training llms to reason and leverage search engines with reinforcement learning, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.450757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:07:18.365922Z digest=sha256:197aed28a7d21873ee9683d0bb1d78942dab60b734d47e4130664ff84cbca72b

Observation 8bdf368b-a736-4ff5-b227-6c47ad6c23f7 · outbound

This paper cites Mmsearch-r1: Incentivizing lmms to search, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Mmsearch-r1: Incentivizing lmms to search, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.442344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:07:18.368961Z digest=sha256:3aeb0a00ef1f439b9d6af0f89f2081ed094d744ac80d911d479da7f42515a777

Observation c4d907dc-e03e-42c8-a09e-407c739803a9 · outbound

This paper cites Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:07:18.433534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:07:18.372033Z digest=sha256:3461a275dafb8922cb941722e12bbdd19fcc25a81a9c760a48bff38e0f85a476

Observation b161d047-2c1b-404d-8a62-0ce2e56eafd8 · outbound

This paper cites Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.374868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.374868Z digest=sha256:dcc6473435d84a9a4f4fadca7f4204d678ccfc563a9d5e16c5a3792d60f5873b

Observation 8f37fafe-1ae4-47f6-9956-5a28792d3deb · outbound

This paper cites TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.377992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.377992Z digest=sha256:459438ec2b388a4d714729cc60346785318b98e16da3ac0fa95134a097ec2ab9

Observation 17e17ad6-cbbf-4845-a144-4485a5f87b7a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:07:18.380838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:07:18.380838Z digest=sha256:c4e4867142c5b9996bd62b8987d14c4fe34a53504c80e1043736993ba9bf1a92

Pith citing papers

Observation 06872784-d724-4494-9912-f89a2a61ba75 · inbound

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services cites this paper.

LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:00:17.278040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:00:17.278040Z digest=sha256:ffbe36f38d87d011f4fe1f6d38b10b481f77c2ae64e0a0251a6240ba834a005a

Observation bddbbd6a-8510-4f17-aa96-93e1f8c41bf3 · inbound

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents cites this paper.

SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:48.471379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:26:57.084201Z digest=sha256:e3b06b1b2d76203bb0e7f5952fdbdeb46e555361eeedf98ae5584c7b38cf3a96

Observation 1142874d-a039-4f24-b1a4-1b215bfacb85 · inbound

VistaHop: Benchmarking Long-Horizon Visual DeepSearch cites this paper.

VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.261024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:35:45.449334Z digest=sha256:43992089d14028b0f07684e95739c49c54629e557b91f2271fb49f861dde40f8

Observation ab8021cf-9c68-49ae-902d-b902bba21c4e · inbound

VistaHop: Benchmarking Long-Horizon Visual DeepSearch cites this paper.

VistaHop: Benchmarking Long-Horizon Visual DeepSearch RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T12:32:59.871356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:32:59.871356Z digest=sha256:70a27777f581332187766d9d40f8df2baf8d2c13ded41b8750f0eed151b0aa19

Observation aa4b8412-441a-466e-9849-001daf2e9e7e · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.748936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:8711c84d5867b1811f5cbda12b05f1de5456956577e35070b96ca0b0c41e0bc9

Observation f73945fa-ece9-4df3-9ccd-86ed0e883e3a · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:36.010374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:36.010374Z digest=sha256:60734b0763f74e718e5cabd886a9bee999a7042b39565162d3c3cd303a401898