Pith. sign in

Paper Citation Record · LEDGER

rStar2-Agent: Agentic Reasoning Technical Report

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 19 inbound Pith citation observations for arXiv:2508.20722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20722 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:39.510592Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:40:25.785903Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 674f8691-2994-450f-8f27-a408c9d60ff7 · outbound

This paper cites Phi-4-reasoning Technical Report.

rStar2-Agent: Agentic Reasoning Technical Report Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.369654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.369654Z digest=sha256:52cda81871d7fde83c16adcca4ee047196c21f78d62a446fe64b7ba31165e4b1

Observation d0f6480d-d4f6-491a-84e2-ce22941b5d90 · outbound

This paper cites Llama-nemotron: Efficient reasoning models.

rStar2-Agent: Agentic Reasoning Technical Report Llama-nemotron: Efficient reasoning models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.380300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.380300Z digest=sha256:94d550e293c94ddfefea7dbd9e521d2f824ba45c2a8f44339058beb6dd48180e

Observation d392f9cb-a736-4c3f-8590-5aa401935edc · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

rStar2-Agent: Agentic Reasoning Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.384805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.384805Z digest=sha256:1c56fbb92e0b57294cd35eaf7fb5a6a58a3dd098302092097d6df53f640f4165

Observation b14bd8d2-9c3c-40da-a071-5805ad758f21 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

rStar2-Agent: Agentic Reasoning Technical Report Reasoning with Exploration: An Entropy Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.390460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.390460Z digest=sha256:93ada2238371be78431cb13e5ea200ba7f49c206b04fedf558dfab41f67f8b4d

Observation e89626b0-8d1c-4c8a-beda-243ec2370369 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

rStar2-Agent: Agentic Reasoning Technical Report The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.395422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.395422Z digest=sha256:806df03d3637c3d3d021aa97c4e90adc53421eb285cfaa444d01b97363902564

Observation 8531d737-cf7d-4b16-ac7b-e5e06958c981 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

rStar2-Agent: Agentic Reasoning Technical Report ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.400085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.400085Z digest=sha256:fd24d01cc7a14c7afeff9123f74415c2e903945d05731a578be7ee1f95c465e9

Observation 1e6a9ace-d75a-451e-b5c1-cc2d3f741e21 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

rStar2-Agent: Agentic Reasoning Technical Report rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.405467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.405467Z digest=sha256:ed9e5194bc3ec6a6561f0370bd4e365f5da115af52e2636683c0891f1471be5d

Observation 3d64c1c3-e5e3-4ed5-903a-1f9d68fb3be2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

rStar2-Agent: Agentic Reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.409953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.409953Z digest=sha256:183c93838052150e4d6aa405584b895c70d82aaba344767509ffb5dc3780bf7f

Observation 6bdb1f8e-e9e3-4861-bf3e-f7ee7c8a74e5 · outbound

This paper cites Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning.

rStar2-Agent: Agentic Reasoning Technical Report Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.414554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.414554Z digest=sha256:f832cdd72d38176ede70e0b6e92fa00971b9097520a699c8cdcb04caad2f28ef

Observation eff25afd-da3a-4465-8fcc-eeff55d22108 · outbound

This paper cites OpenAI o1 System Card.

rStar2-Agent: Agentic Reasoning Technical Report OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.419049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.419049Z digest=sha256:621625d54873639dce8b994e0b2ed9f49d98be05df8406d69b655f1e2183f551

Observation 0b4a37c3-6a47-4f09-a775-3a90ffedb885 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

rStar2-Agent: Agentic Reasoning Technical Report ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.427791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.427791Z digest=sha256:209c751cf2ed57b939037c0c16b5ce2ac41d7e0252b77bdf8c2824baee70c6e0

Observation 5e308ebe-28d5-4453-8387-8d91a8dea797 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

rStar2-Agent: Agentic Reasoning Technical Report ToolACE: Winning the Points of LLM Function Calling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.437030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.437030Z digest=sha256:ff90970fa26359cf4f77e4709e50c1fb1d4ca345ec8acfa4b8e85435299b6e4b

Observation 4e85ad13-c4e7-40ce-8ed5-8bdbda4a1f96 · outbound

This paper cites Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving.

rStar2-Agent: Agentic Reasoning Technical Report Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.441530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.441530Z digest=sha256:0a260fde309d332a3abf02d43c1546a83c4007e184c0af9f6c789849553fbf4b

Observation 9c7f279d-c669-4fc1-8ffb-568518d163cc · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

rStar2-Agent: Agentic Reasoning Technical Report AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.446005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.446005Z digest=sha256:6bc743ba71e33501c0eb302eedd9cc3d037caa5618ac96c929766acdc5a3ceac

Observation 435b59c5-8e68-488e-b7eb-e59cf0e6f3c0 · outbound

This paper cites APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay.

rStar2-Agent: Agentic Reasoning Technical Report APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.450498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.450498Z digest=sha256:23e44537d51fa893359a0339925b384d026cb84b18060068e60d9e08c06e0623

Observation 2d14ae9a-cb97-421e-b964-938ca4cdd7fe · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

rStar2-Agent: Agentic Reasoning Technical Report ToolRL: Reward is All Tool Learning Needs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.455290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.455290Z digest=sha256:36d5838869f55d228efc60d8e1a9449673ee7ac22ac5d7841d02592ef9763de1

Observation a6555b84-43a2-475f-b906-cd989fed6db3 · outbound

This paper cites Magistral.

rStar2-Agent: Agentic Reasoning Technical Report Magistral

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.459947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.459947Z digest=sha256:ff06ff48c5e750fa39a97859ae90940124d121ce946e123d4020b86d71796c64

Observation 7ea5839c-db55-4e8f-b609-015076d6cb7d · outbound

This paper cites ByteDance Seed, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, Zhiqi Lin, Mingxuan Wang, Chengyi Wang, Xiangpeng Wei, Wenyuan Xu, et al.

rStar2-Agent: Agentic Reasoning Technical Report ByteDance Seed, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, Zhiqi Lin, Mingxuan Wang, Chengyi Wang, Xiangpeng Wei, Wenyuan Xu, et al

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.464672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.464672Z digest=sha256:ca11b2206d4d5144e4aab3a2fea22c9b916a1745483f6f889948e34f1da4acdb

Observation e0f49d86-4f6b-4113-b811-406ebd9d8947 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

rStar2-Agent: Agentic Reasoning Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.469144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.469144Z digest=sha256:d1e26853ba8657217ff9a955eaa98a85b4c5f8fa6a5161180139919b35e74e37

Observation 89a5a147-9639-4331-819a-90690405db0f · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

rStar2-Agent: Agentic Reasoning Technical Report HybridFlow: A Flexible and Efficient RLHF Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.473799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.473799Z digest=sha256:0264f3fcdeee745fa0b37b1a51f0414d973d34282501cb1f9d6f132e7a34edd5

Observation d4bf7628-c16d-40fe-8b29-34cbe67c9916 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

rStar2-Agent: Agentic Reasoning Technical Report Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.478441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.478441Z digest=sha256:440a168b73669363a41d5ce1bedaab4c0717a1d8738891f69b2e9f6b909d13f8

Observation 0527649f-084d-44be-b31d-b22c6d068207 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

rStar2-Agent: Agentic Reasoning Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.482938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.482938Z digest=sha256:ed7109a5dd3d0cd24e53e072b6b6b4c775c8327046227d8103003dbf4b03b89b

Observation a7a0bdfc-3922-4041-8048-6598a871fdd0 · outbound

This paper cites Qwen3 Technical Report.

rStar2-Agent: Agentic Reasoning Technical Report Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.487519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.487519Z digest=sha256:12f938cf52e90ce098cc93b38499a9acdb050fa70b95ae7abc684b9e8c76079b

Observation ca34421e-6542-49e5-9a8e-4644f7166bc4 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

rStar2-Agent: Agentic Reasoning Technical Report Magicoder: Empowering Code Generation with OSS-Instruct

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.492222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.492222Z digest=sha256:2d93da48c69075516cfd7f1df6bcb01ba298e8f3347440d73077f25d91b4c118

Observation f901aec4-ac7c-4ad1-bfab-7917b16802c7 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

rStar2-Agent: Agentic Reasoning Technical Report MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.496906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.496906Z digest=sha256:97f057202f566a95fd15db457faff64b5e86bd8017033a7b11a2e4ea72da609a

Observation 923c90b6-2151-45b6-9b45-298703dfc78c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

rStar2-Agent: Agentic Reasoning Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.501283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.501283Z digest=sha256:73d78e849c1f50e97238fe16ea13a10cbbebbd55a56083f2bbb1729b70f1c09c

Observation 5c478112-4f81-4129-80cc-85d947720eb5 · outbound

This paper cites Promoting Efficient Reasoning with Verifiable Stepwise Reward.

rStar2-Agent: Agentic Reasoning Technical Report Promoting Efficient Reasoning with Verifiable Stepwise Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.505938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.505938Z digest=sha256:15a5e16614d47f32eaab4ff88c4e1f1bed868f705365c0f1a8fc0391542d98c2

Observation 19ad39bf-ede6-4d10-b885-327da8382ff2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

rStar2-Agent: Agentic Reasoning Technical Report Instruction-Following Evaluation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.510592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.510592Z digest=sha256:d716927c48105782804514bbe855d32d08f179c5cee3095e3fafeddd9f4a93f0

Observation aa9b1cef-8270-4200-9f95-350fa9bf2956 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

rStar2-Agent: Agentic Reasoning Technical Report ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.432202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.432202Z digest=sha256:143ce6e6b3da17bfc1c75d1ba4d944bd56851ffa224556e85acea122208ec73f

Observation ee948300-1b9d-4fc3-912c-928028a5fbf9 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

rStar2-Agent: Agentic Reasoning Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.423424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.423424Z digest=sha256:0bcca7dca14185bd171b31d5f5a05d7b189767c7e39a7ef9554b26474fcb1e37

Observation d4edf230-9a3f-4cbd-8ca3-7c9b740426ab · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

rStar2-Agent: Agentic Reasoning Technical Report MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.374889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.374889Z digest=sha256:2a84d57d1cc0941c5a4272a7d3b0a89c037f63b8f69953748ee006f2380c35de

Pith citing papers

Observation 46961b85-16bb-4847-91a5-50b12008bd73 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey rStar2-Agent: Agentic Reasoning Technical Report

Reference 123

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.445880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:68700972c717f9744bc23f43ec56a5a821fcf758d22ac337c088763f1c53378e

Observation 03de0ad8-bd7c-4deb-9eb4-a8e79f32dea3 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving rStar2-Agent: Agentic Reasoning Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:25.785903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:25.785903Z digest=sha256:d8abe40000be9161907c601243fe33ee1e3c197df23c66ca01a5c32073959ee0

Observation 6ba424a7-071a-451d-853e-4bd2233f05a6 · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:41.340195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:41.340195Z digest=sha256:1dc333f8294b095bbe822891ca0364843c466a8e8c2e9f10bf18f7b8db31b797

Observation d2bb13be-c000-42e5-a325-75a7376d07c8 · inbound

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning cites this paper.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.933722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.933722Z digest=sha256:555bd57199198c02a7fa8f4d9a85b89bcb47a4cbf944fe25a8e20967c441d041

Observation f44f5df8-ac25-45ac-9c26-8ba6bfa81712 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste rStar2-Agent: Agentic Reasoning Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:54.599326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:54.599326Z digest=sha256:5287e825611cf9ff971ecd50151620a5e5910118f6d11b4620448c5d908b6d18

Observation 7d6c79d8-0511-48b1-a72b-e05d4ffab47b · inbound

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning cites this paper.

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:50.855350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:14:10.609406Z digest=sha256:37dc43882c816f648cc2ca63b6b9ef5ad4705767e72d74c54eda74314c8c1d69

Observation 52ae9bc8-3b34-4067-8e51-a32db106e2ed · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory rStar2-Agent: Agentic Reasoning Technical Report

Reference 224

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:05.839585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:cfdb83455266886c7d4b4b923b12c1de9354d4bd95595df822c799964b8442e1

Observation 6c362dbf-4887-4fc2-bfe5-842379e3e598 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training rStar2-Agent: Agentic Reasoning Technical Report

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:11.260353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:1385a548b1f690bdedc7175e9331ef5b54595473e6180e1669ee4dee9f1cbd48

Observation fdc7fda0-30ad-431e-8f56-c5fc5287af24 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.161743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:16765e3cea40a1f5c760824ea73c6d040636e21d172aaf4bf9548f76bd6adaf7

Observation e1aaefe0-d289-4c88-954d-bcf99c6a9463 · inbound

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL cites this paper.

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL rStar2-Agent: Agentic Reasoning Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.568637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:23:13.535140Z digest=sha256:4f3a403c8983005889c3ed9bb0ea9edc430bf359aaa870952e301b456bd2e1d8

Observation a3e970f5-5a02-484a-a52b-31115fe346fe · inbound

Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning cites this paper.

Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:01:10.556776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T10:30:36.330305Z digest=sha256:27b7bf454f46da6b3a5f4347d9397b4354413e9b81ef51cd92d64bcbd60b5953

Observation 9e213d58-4445-4618-9bf7-ef103d03f21a · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code rStar2-Agent: Agentic Reasoning Technical Report

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:15:56.569722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:33:02.995360Z digest=sha256:7cda82253defc1b6c373f824f56a68c4e356deffa5b1ad378b76f5cfb704438e

Observation ba510941-67cf-4756-b7fb-9bfc7df61b78 · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code rStar2-Agent: Agentic Reasoning Technical Report

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.333349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:26:11.265781Z digest=sha256:e8c1a93319b617bd9d909d0c6ac4f3459e156a6a760c75ea9a82a431e2a67387

Observation c1a0a899-ee28-4c80-a2d1-6adafbc0b040 · inbound

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning cites this paper.

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:23:14.031521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:21:39.492072Z digest=sha256:cd94583e1fdecf1696e67d859febc09ba898dbe4b62c8d652ab1360117847a25

Observation c6f55ad0-8721-44ef-ae29-20152abc4df5 · inbound

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming cites this paper.

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming rStar2-Agent: Agentic Reasoning Technical Report

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.161949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:19:22.541146Z digest=sha256:41570ecd2953cac6b4a7eff95920f5861573bd13af1845a848b6f027943845a3

Observation 20f9e4b4-7ab9-482f-b412-fd80cc2ac8c7 · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:23.734275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:72e734cf0597c6f2444edfc4e4a88d4653db334ba6fc6b6023e40d41c49ed024

Observation cc52f01b-67d1-467f-b515-c6bb9a6670df · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.248154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:44bbcf04875b344549c12ed5d803e8d94127a307e738fe225023aeb83df24de2

Observation 36d350a3-81e4-46fa-91dc-29c9616c577d · inbound

Discovering Millions of Interpretable Features with Sparse Autoencoders cites this paper.

Discovering Millions of Interpretable Features with Sparse Autoencoders rStar2-Agent: Agentic Reasoning Technical Report

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.059796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:34:00.754172Z digest=sha256:4c7037478d909130434f6c9c81d2378386552d6bb3181af61033506281c955ba

Observation 0b32f477-f016-4790-b9ca-9fb3fd6adcbc · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language rStar2-Agent: Agentic Reasoning Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:49.632692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:49.632692Z digest=sha256:af8c6c01e20c7710e4cee09c3c0af68efac2befdc9cb436cdd3e881a6add5b35