Pith. sign in

Paper Citation Record · LEDGER

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2601.15141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.15141 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:01:51.020857Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:34:52.336080Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.265122Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4ab3d3b-027d-4553-895e-a50d9a704d3c · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.349112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.349112Z digest=sha256:f5e09d7ab92e1f3233024fff7ba415f907ff8a2175f04fd0d37cd053d7d7ff17

Observation e714dde6-3275-43e7-887e-7dc0ec4f00d8 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.654520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.654520Z digest=sha256:30590cb9ea9e827ac2a3ec2e5b70e0fe611d4527a0cbdaa4cd04f62f9908f612

Observation bbaf49e3-f50a-4ba9-8bc2-a838fd1b49c6 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.820681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.820681Z digest=sha256:d1b077018868b0afc4d99111903be7e66b3e22cebf91cd772ac6048fc0ee7c17

Observation 180c1cac-d97e-4ecd-8ddb-c91f4bfcbdc5 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.960580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.960580Z digest=sha256:2d5f60c8e246e6627e86ae3ee011075506d62f7604e4ed9c416258bdbe75c16e

Observation ed3f8c17-543b-47e7-81eb-eed34f2ded00 · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Skywork Open Reasoner 1 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.068086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.068086Z digest=sha256:2378196d68154b390d66565ed6082c1488684204d322133daa9fe7ff14eca060

Observation b5601a98-0cf0-4a9b-ab4b-e18507318bbb · outbound

This paper cites Qwen2.5-Coder Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Qwen2.5-Coder Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.175019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.175019Z digest=sha256:b43c1649f3f045993cbe2b79e7308be36d600be3d5ff98efcbe7c6a62bfaf810

Observation 11165d18-2286-4f7f-bb6d-bb09a955bfec · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.238425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.238425Z digest=sha256:bc31b80ea0755c32517a4a7d61d6c8e6c589a5140bc01e49e17c4dbe0fded9d1

Observation 142c089e-79e0-4e3d-89ff-a7f711dcfe86 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.282180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.282180Z digest=sha256:bbab4faea1b88190cb3ce143f20850e955efdec03c367890918a986e568d9349

Observation 32cf656d-bd86-4703-97a2-04814d70c89e · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.333636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.333636Z digest=sha256:ef623297e955d6101163b416c8979e986b2ad2e9e2647527e043aaf4f1b5ff2a

Observation 4a7dc707-789b-4e8c-920a-913bc6f4425d · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.390340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.390340Z digest=sha256:c80ef398e75d278f81d824467a9a94cd43bbb1a3daa856be3786b5703032a3a0

Observation a2ca8d26-6970-4f63-b54e-ed5efa763351 · outbound

This paper cites DeepSeek-V3 Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.434593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.434593Z digest=sha256:5e4253f893542b63a93e2a5a1dcad3a6ae45b0ec04a7258626732520252bd1d0

Observation 434df131-b215-494a-a712-39f06697e8c7 · outbound

This paper cites TALM: Tool Augmented Language Models.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning TALM: Tool Augmented Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.506280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.506280Z digest=sha256:12145ab07bd25b5b53e4d049757d036591c03e756634eeb2435df932483c6122

Observation ab254e2c-0c71-4078-b29c-ad48dd7389c4 · outbound

This paper cites Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.563297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.563297Z digest=sha256:c3a074331cdd4beb57693d9641a9e0d77d5dbd2dfbe2c454e2d322c0c52b78bd

Observation 7ab0f697-fc9d-4e97-8423-7290699715d7 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.606749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.606749Z digest=sha256:b8129897bcf3b919593a805af67499f648bbb7f109dfcbf9a0561a8788be1272

Observation de3eaf9c-8587-4984-a51f-5b8bc5bbad01 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.682153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.682153Z digest=sha256:0a37a4b5cc33e04bf4a8237d618d8fb311dd2b1557af2964a60abbc9034936bd

Observation 98bcfdf8-07fa-49b6-8372-7ea0a030e119 · outbound

This paper cites pocoo.org/2025/10/17/code/.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning pocoo.org/2025/10/17/code/

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.821208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.821208Z digest=sha256:acbfc10616cf208db0f36de5f59127bdd95beca5ed9ae7bce3f614936dea84e1

Observation d2bb13be-c000-42e5-a325-75a7376d07c8 · outbound

This paper cites rStar2-Agent: Agentic Reasoning Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.933722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.933722Z digest=sha256:7ef410296eedee3fd193e2513797c21e07e6e1ab1b7c5d15f7446020383325b3

Observation 5928fef3-ab27-4ab0-9044-f5aa5a6d68eb · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.047393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.047393Z digest=sha256:fb539658462300c4b23fd77df445c399aaa60b1bf148e2196483916677b43114

Observation 7323dc5c-2272-404c-9b85-c0a26531bf90 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.181248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.181248Z digest=sha256:b2dccedb5a58342bc397b75cd93a3f3f483b566f803461073d355a9a8e68afbd

Observation 7d71d9eb-b6e0-40a5-ae43-97b93d7b3733 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.527848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.527848Z digest=sha256:aecc7df83be279c785e89f1c515fe6319952515b242c8c9e12b4901fda7c52ff

Observation 3be6609d-e3ea-4139-b1a0-b2be409ba534 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.680369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.680369Z digest=sha256:bf161e265fae1bb0edd48667849e3d616e396f7453c1cb5b07b0b02b328eef99

Observation e7050150-7c8a-43fd-814b-2a4d581f2c0d · outbound

This paper cites True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.797041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.797041Z digest=sha256:5b8798444c3a3335e0639f63c9ba2d980e0fdb6cbd9376e18dd8654625f0b065

Observation bd56bd04-6138-4ab8-bd2f-20d03ae4cd79 · outbound

This paper cites DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.824574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.824574Z digest=sha256:bffa9991bce4cfe8bd8eba844a4af90b903b265503e54384437ddddbbc7d6d76

Observation 8d8949bb-f3fb-444a-a76b-3f016ae956c1 · outbound

This paper cites Qwen3 Technical Report.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.828166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.828166Z digest=sha256:409d8fea87a38a6138f4551fb28041642805b109a973a5485cc6bee48ab6fc6e

Observation e2a7845b-c40d-4353-9094-ae815c6317e5 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.876731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.876731Z digest=sha256:f4a9e46be8c0c9bb0ffbf694844d04c598e0f3e2b98d35181a77994c4d1330ab

Observation 39931dbd-5fe6-461a-8fa0-ca30c8a38a7b · outbound

This paper cites purified.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning purified

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:51.020857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:51.020857Z digest=sha256:cc244ce69c452f0772f1c438d0d99fc256e3cf914f9303a2e6c916ce3e4940c6

Observation b564bf72-8e55-4866-ab6c-5958a9e452f4 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:50.366780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:50.366780Z digest=sha256:35ad4ed5c5f67b189a8ef1e8cded64621cf76b37752e824d4936393f52fd7c31

Observation 9246615c-e139-4311-addc-cc4d0d27038d · outbound

This paper cites Agentic entropy-balanced policy optimization.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Agentic entropy-balanced policy optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.539543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.539543Z digest=sha256:4ec8e40c4856db25f16c30df759357ef9ca41111415a731d5e6548a3c0a72de0

Observation 588c30f6-4fc9-4a2e-9190-e5a041ed1f75 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.015328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.015328Z digest=sha256:c52eb40b79c79a38092ffaced3993a147bf56b257e935efbc805488d48b06f93

Observation 3d2ac0bf-b775-47b6-b766-fa0590755ef2 · outbound

This paper cites Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.241440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.241440Z digest=sha256:90678591fcc4b2ce42850f8515f3a5bde0282388af9553aaa527acfd2a2a8b79

Observation 022a5203-10bc-4fa8-970e-a6c44630f96e · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.458376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.458376Z digest=sha256:8e25b8ddfbc4026b630065b4fa967a108b7e5185219f4711bcb20d95ec00fcd6

Pith citing papers

Observation 5b4f6607-78ba-464c-acf5-87d8e6f19b8d · inbound

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents cites this paper.

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:02.718893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-22T06:10:26.185447Z digest=sha256:ff8b309fa4cf4ff6c1fcd2e6c98fbd404e0144dd26e1419424e3433aae5dff38

Observation a0d08ac2-af13-4bb9-9029-08f8a770c63f · inbound

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents cites this paper.

Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-07T02:16:02.718893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T17:34:52.336080Z digest=sha256:8a4e7a6d51bcf973ede210ff5749490cead8e718a9fc9fcca33891f7f51683e0