Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T09:01:51.020857Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 2 inbound Pith citation observations for arXiv:2601.15141.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T09:01:51.020857Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T17:34:52.336080Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T17:34:57.265122Z
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e4ab3d3b-027d-4553-895e-a50d9a704d3c · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e714dde6-3275-43e7-887e-7dc0ec4f00d8 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbaf49e3-f50a-4ba9-8bc2-a838fd1b49c6 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180c1cac-d97e-4ecd-8ddb-c91f4bfcbdc5 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3f8c17-543b-47e7-81eb-eed34f2ded00 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Skywork Open Reasoner 1 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5601a98-0cf0-4a9b-ab4b-e18507318bbb · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Qwen2.5-Coder Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11165d18-2286-4f7f-bb6d-bb09a955bfec · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 142c089e-79e0-4e3d-89ff-a7f711dcfe86 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32cf656d-bd86-4703-97a2-04814d70c89e · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7dc707-789b-4e8c-920a-913bc6f4425d · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning WebSailor: Navigating Super-human Reasoning for Web Agent
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ca8d26-6970-4f63-b54e-ed5efa763351 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DeepSeek-V3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434df131-b215-494a-a712-39f06697e8c7 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning TALM: Tool Augmented Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab254e2c-0c71-4078-b29c-ad48dd7389c4 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Defeating the training-inference mismatch via fp16.arXiv preprint arXiv:2510.26788,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab0f697-fc9d-4e97-8423-7290699715d7 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToolRL: Reward is All Tool Learning Needs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3eaf9c-8587-4984-a51f-5b8bc5bbad01 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bcfdf8-07fa-49b6-8372-7ea0a030e119 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning pocoo.org/2025/10/17/code/
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2bb13be-c000-42e5-a325-75a7376d07c8 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5928fef3-ab27-4ab0-9044-f5aa5a6d68eb · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7323dc5c-2272-404c-9b85-c0a26531bf90 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d71d9eb-b6e0-40a5-ae43-97b93d7b3733 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be6609d-e3ea-4139-b1a0-b2be409ba534 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7050150-7c8a-43fd-814b-2a4d581f2c0d · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd56bd04-6138-4ab8-bd2f-20d03ae4cd79 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8949bb-f3fb-444a-a76b-3f016ae956c1 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Qwen3 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a7845b-c40d-4353-9094-ae815c6317e5 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39931dbd-5fe6-461a-8fa0-ca30c8a38a7b · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning purified
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b564bf72-8e55-4866-ab6c-5958a9e452f4 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9246615c-e139-4311-addc-cc4d0d27038d · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Agentic entropy-balanced policy optimization
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 588c30f6-4fc9-4a2e-9190-e5a041ed1f75 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2ac0bf-b775-47b6-b766-fa0590755ef2 · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022a5203-10bc-4fa8-970e-a6c44630f96e · outbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b4f6607-78ba-464c-acf5-87d8e6f19b8d · inbound
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0d08ac2-af13-4bb9-9029-08f8a770c63f · inbound
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.