Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:39.510592Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 19 inbound Pith citation observations for arXiv:2508.20722.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:39.510592Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T06:40:25.785903Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
31 of 31 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 674f8691-2994-450f-8f27-a408c9d60ff7 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Phi-4-reasoning Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f6480d-d4f6-491a-84e2-ce22941b5d90 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Llama-nemotron: Efficient reasoning models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d392f9cb-a736-4c3f-8590-5aa401935edc · outbound
rStar2-Agent: Agentic Reasoning Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14bd8d2-9c3c-40da-a071-5805ad758f21 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Reasoning with Exploration: An Entropy Perspective
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e89626b0-8d1c-4c8a-beda-243ec2370369 · outbound
rStar2-Agent: Agentic Reasoning Technical Report The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8531d737-cf7d-4b16-ac7b-e5e06958c981 · outbound
rStar2-Agent: Agentic Reasoning Technical Report ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e6a9ace-d75a-451e-b5c1-cc2d3f741e21 · outbound
rStar2-Agent: Agentic Reasoning Technical Report rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d64c1c3-e5e3-4ed5-903a-1f9d68fb3be2 · outbound
rStar2-Agent: Agentic Reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bdb1f8e-e9e3-4861-bf3e-f7ee7c8a74e5 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff25afd-da3a-4465-8fcc-eeff55d22108 · outbound
rStar2-Agent: Agentic Reasoning Technical Report OpenAI o1 System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4a37c3-6a47-4f09-a775-3a90ffedb885 · outbound
rStar2-Agent: Agentic Reasoning Technical Report ToRL: Scaling Tool-Integrated RL
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e308ebe-28d5-4453-8387-8d91a8dea797 · outbound
rStar2-Agent: Agentic Reasoning Technical Report ToolACE: Winning the Points of LLM Function Calling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e85ad13-c4e7-40ce-8ed5-8bdbda4a1f96 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c7f279d-c669-4fc1-8ffb-568518d163cc · outbound
rStar2-Agent: Agentic Reasoning Technical Report AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435b59c5-8e68-488e-b7eb-e59cf0e6f3c0 · outbound
rStar2-Agent: Agentic Reasoning Technical Report APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d14ae9a-cb97-421e-b964-938ca4cdd7fe · outbound
rStar2-Agent: Agentic Reasoning Technical Report ToolRL: Reward is All Tool Learning Needs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6555b84-43a2-475f-b906-cd989fed6db3 · outbound
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea5839c-db55-4e8f-b609-015076d6cb7d · outbound
rStar2-Agent: Agentic Reasoning Technical Report ByteDance Seed, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, Zhiqi Lin, Mingxuan Wang, Chengyi Wang, Xiangpeng Wei, Wenyuan Xu, et al
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f49d86-4f6b-4113-b811-406ebd9d8947 · outbound
rStar2-Agent: Agentic Reasoning Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a5a147-9639-4331-819a-90690405db0f · outbound
rStar2-Agent: Agentic Reasoning Technical Report HybridFlow: A Flexible and Efficient RLHF Framework
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4bf7628-c16d-40fe-8b29-34cbe67c9916 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0527649f-084d-44be-b31d-b22c6d068207 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a0bdfc-3922-4041-8048-6598a871fdd0 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Qwen3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca34421e-6542-49e5-9a8e-4644f7166bc4 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Magicoder: Empowering Code Generation with OSS-Instruct
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f901aec4-ac7c-4ad1-bfab-7917b16802c7 · outbound
rStar2-Agent: Agentic Reasoning Technical Report MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923c90b6-2151-45b6-9b45-298703dfc78c · outbound
rStar2-Agent: Agentic Reasoning Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c478112-4f81-4129-80cc-85d947720eb5 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Promoting Efficient Reasoning with Verifiable Stepwise Reward
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ad39bf-ede6-4d10-b885-327da8382ff2 · outbound
rStar2-Agent: Agentic Reasoning Technical Report Instruction-Following Evaluation for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9b1cef-8270-4200-9f95-350fa9bf2956 · outbound
rStar2-Agent: Agentic Reasoning Technical Report ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee948300-1b9d-4fc3-912c-928028a5fbf9 · outbound
rStar2-Agent: Agentic Reasoning Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4edf230-9a3f-4cbd-8ca3-7c9b740426ab · outbound
rStar2-Agent: Agentic Reasoning Technical Report MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46961b85-16bb-4847-91a5-50b12008bd73 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey rStar2-Agent: Agentic Reasoning Technical Report
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 03de0ad8-bd7c-4deb-9eb4-a8e79f32dea3 · inbound
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving rStar2-Agent: Agentic Reasoning Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba424a7-071a-451d-853e-4bd2233f05a6 · inbound
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning rStar2-Agent: Agentic Reasoning Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2bb13be-c000-42e5-a325-75a7376d07c8 · inbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44f5df8-ac25-45ac-9c26-8ba6bfa81712 · inbound
AI Can Learn Scientific Taste rStar2-Agent: Agentic Reasoning Technical Report
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6c79d8-0511-48b1-a72b-e05d4ffab47b · inbound
SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning rStar2-Agent: Agentic Reasoning Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 52ae9bc8-3b34-4067-8e51-a32db106e2ed · inbound
Fine-Tuning Small Reasoning Models for Quantum Field Theory rStar2-Agent: Agentic Reasoning Technical Report
Reference 224
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6c362dbf-4887-4fc2-bfe5-842379e3e598 · inbound
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training rStar2-Agent: Agentic Reasoning Technical Report
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fdc7fda0-30ad-431e-8f56-c5fc5287af24 · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e1aaefe0-d289-4c88-954d-bcf99c6a9463 · inbound
Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL rStar2-Agent: Agentic Reasoning Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3e970f5-5a02-484a-a52b-31115fe346fe · inbound
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning rStar2-Agent: Agentic Reasoning Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e213d58-4445-4618-9bf7-ef103d03f21a · inbound
Teaching Language Models to Think in Code rStar2-Agent: Agentic Reasoning Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba510941-67cf-4756-b7fb-9bfc7df61b78 · inbound
Teaching Language Models to Think in Code rStar2-Agent: Agentic Reasoning Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1a0a899-ee28-4c80-a2d1-6adafbc0b040 · inbound
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning rStar2-Agent: Agentic Reasoning Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c6f55ad0-8721-44ef-ae29-20152abc4df5 · inbound
CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming rStar2-Agent: Agentic Reasoning Technical Report
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 20f9e4b4-7ab9-482f-b412-fd80cc2ac8c7 · inbound
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning rStar2-Agent: Agentic Reasoning Technical Report
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc52f01b-67d1-467f-b515-c6bb9a6670df · inbound
Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 36d350a3-81e4-46fa-91dc-29c9616c577d · inbound
Discovering Millions of Interpretable Features with Sparse Autoencoders rStar2-Agent: Agentic Reasoning Technical Report
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b32f477-f016-4790-b9ca-9fb3fd6adcbc · inbound
PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language rStar2-Agent: Agentic Reasoning Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.