Pith. sign in

Paper Citation Record · LEDGER

rStar2-Agent: Agentic Reasoning Technical Report

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 19 inbound Pith citation observations for arXiv:2508.20722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20722 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:57:39.510592Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:40:25.785903Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 674f8691-2994-450f-8f27-a408c9d60ff7 · outbound

This paper cites Phi-4-reasoning Technical Report.

rStar2-Agent: Agentic Reasoning Technical Report Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.369654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.369654Z digest=sha256:7b2735a705fbdc7b1ae202a14a9cc7e2d00393bf57a3dcb7d04d7df22f3edce1

Observation d0f6480d-d4f6-491a-84e2-ce22941b5d90 · outbound

This paper cites Llama-nemotron: Efficient reasoning models.

rStar2-Agent: Agentic Reasoning Technical Report Llama-nemotron: Efficient reasoning models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.380300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.380300Z digest=sha256:8e8ec4f5e81921d757f80d87e2e4cc601896268fb8b33a6d1930fdfeab640401

Observation d392f9cb-a736-4c3f-8590-5aa401935edc · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

rStar2-Agent: Agentic Reasoning Technical Report MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.384805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.384805Z digest=sha256:67af6004881a92902195e087dfbfb0907e3a98ff4915955a903df84c5739ce52

Observation b14bd8d2-9c3c-40da-a071-5805ad758f21 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

rStar2-Agent: Agentic Reasoning Technical Report Reasoning with Exploration: An Entropy Perspective

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.390460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.390460Z digest=sha256:6451716d680068b02714489ecf3730e1c0aaa0cfe89d836df69f4aeab6460aba

Observation e89626b0-8d1c-4c8a-beda-243ec2370369 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

rStar2-Agent: Agentic Reasoning Technical Report The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.395422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.395422Z digest=sha256:2e570f115a83c1ab051e61704af1c124db7ee67fea85c7eb1603914039f18052

Observation 8531d737-cf7d-4b16-ac7b-e5e06958c981 · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

rStar2-Agent: Agentic Reasoning Technical Report ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.400085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.400085Z digest=sha256:394549b561bde46ced920d6a16864757fd712f36601c35383f59112313e62446

Observation 1e6a9ace-d75a-451e-b5c1-cc2d3f741e21 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

rStar2-Agent: Agentic Reasoning Technical Report rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.405467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.405467Z digest=sha256:75b12e800f34b67e7c9793c9a16fe3156fc3f595c0842d23d7a738dd50ec602b

Observation 3d64c1c3-e5e3-4ed5-903a-1f9d68fb3be2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

rStar2-Agent: Agentic Reasoning Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.409953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.409953Z digest=sha256:6a595c82633672d40ea613ce741782d7c8bc91d51c74ea14416409df2e63c5db

Observation 6bdb1f8e-e9e3-4861-bf3e-f7ee7c8a74e5 · outbound

This paper cites Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning.

rStar2-Agent: Agentic Reasoning Technical Report Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.414554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.414554Z digest=sha256:3a602b4f64efd5c47d7eb6d47b2968820751afe989d5b8d2fddf923e7122ef08

Observation eff25afd-da3a-4465-8fcc-eeff55d22108 · outbound

This paper cites OpenAI o1 System Card.

rStar2-Agent: Agentic Reasoning Technical Report OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.419049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.419049Z digest=sha256:02fdce5204320649f1f96f752eac3eef8163382f50dcf6c083515673dafd1e1a

Observation 0b4a37c3-6a47-4f09-a775-3a90ffedb885 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

rStar2-Agent: Agentic Reasoning Technical Report ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.427791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.427791Z digest=sha256:7431fe378f155ec74190af6f92d130659ebcf49ea769c5d89db70c07149ed367

Observation 5e308ebe-28d5-4453-8387-8d91a8dea797 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

rStar2-Agent: Agentic Reasoning Technical Report ToolACE: Winning the Points of LLM Function Calling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.437030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.437030Z digest=sha256:70ef0a488797e3874e0b44745b4b088d2be0d0d552de7d0c175141965fe61b56

Observation 4e85ad13-c4e7-40ce-8ed5-8bdbda4a1f96 · outbound

This paper cites Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving.

rStar2-Agent: Agentic Reasoning Technical Report Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.441530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.441530Z digest=sha256:c0af3c25fa0405489d85f957244045fbca463e6bc72003c977192d0e8a2a0f8f

Observation 9c7f279d-c669-4fc1-8ffb-568518d163cc · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

rStar2-Agent: Agentic Reasoning Technical Report AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.446005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.446005Z digest=sha256:13022c6d1fb1d8d34d7383839a7cb10442607fa8707b71c5d718ad536dfef99d

Observation 435b59c5-8e68-488e-b7eb-e59cf0e6f3c0 · outbound

This paper cites APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay.

rStar2-Agent: Agentic Reasoning Technical Report APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.450498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.450498Z digest=sha256:51913da8587bdd5b13e591acb9b3ff28f02575a6f3bce7353dd3fda59c5aa40f

Observation 2d14ae9a-cb97-421e-b964-938ca4cdd7fe · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

rStar2-Agent: Agentic Reasoning Technical Report ToolRL: Reward is All Tool Learning Needs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.455290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.455290Z digest=sha256:8c534fd1f230eaf773114b9bc76a3ce8a51f2385a3ef2ca15f5ab69eb0a0bf8f

Observation a6555b84-43a2-475f-b906-cd989fed6db3 · outbound

This paper cites Magistral.

rStar2-Agent: Agentic Reasoning Technical Report Magistral

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.459947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.459947Z digest=sha256:73b8db8b2220ead9af2d018588ef685f264cbda53f8c2fbaa243e5494188c700

Observation 7ea5839c-db55-4e8f-b609-015076d6cb7d · outbound

This paper cites ByteDance Seed, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, Zhiqi Lin, Mingxuan Wang, Chengyi Wang, Xiangpeng Wei, Wenyuan Xu, et al.

rStar2-Agent: Agentic Reasoning Technical Report ByteDance Seed, Jiaze Chen, Tiantian Fan, Xin Liu, Lingjun Liu, Zhiqi Lin, Mingxuan Wang, Chengyi Wang, Xiangpeng Wei, Wenyuan Xu, et al

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.464672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.464672Z digest=sha256:372c6e9cbe0dd591bb9fa867084d27755823130f727df54b4a4291357f72141e

Observation e0f49d86-4f6b-4113-b811-406ebd9d8947 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

rStar2-Agent: Agentic Reasoning Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.469144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.469144Z digest=sha256:b86d150d48dfcdbea16d647f983d8421beda18edd8d4700b000af7d74f8847eb

Observation 89a5a147-9639-4331-819a-90690405db0f · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

rStar2-Agent: Agentic Reasoning Technical Report HybridFlow: A Flexible and Efficient RLHF Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.473799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.473799Z digest=sha256:10072c76927ca29b7e56266f4c27564e00bd3553ddafeba6c8724af5d5e4f029

Observation d4bf7628-c16d-40fe-8b29-34cbe67c9916 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

rStar2-Agent: Agentic Reasoning Technical Report Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.478441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.478441Z digest=sha256:cf49717193b16a3567db26dcb52fed45871c98cc795e6e2da8b9453025752056

Observation 0527649f-084d-44be-b31d-b22c6d068207 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

rStar2-Agent: Agentic Reasoning Technical Report Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.482938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.482938Z digest=sha256:9948c54b5b5026128ec299c6a4fc915e588a2ea0debd23c17330c92c71c27786

Observation a7a0bdfc-3922-4041-8048-6598a871fdd0 · outbound

This paper cites Qwen3 Technical Report.

rStar2-Agent: Agentic Reasoning Technical Report Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.487519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.487519Z digest=sha256:2abd5b7059db15045aaf6cf10eb35877c464651fa6cf2b383c0bba8aade76ec2

Observation ca34421e-6542-49e5-9a8e-4644f7166bc4 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

rStar2-Agent: Agentic Reasoning Technical Report Magicoder: Empowering Code Generation with OSS-Instruct

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.492222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.492222Z digest=sha256:b6964d59e62cda99fb1464a726420a1de3a0a2bef91a54a2e699fe133646dec5

Observation f901aec4-ac7c-4ad1-bfab-7917b16802c7 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

rStar2-Agent: Agentic Reasoning Technical Report MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.496906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.496906Z digest=sha256:80e43592b8f3718422354fe5955416210ca49f2b6933cdfcb9061bea582b619e

Observation 923c90b6-2151-45b6-9b45-298703dfc78c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

rStar2-Agent: Agentic Reasoning Technical Report DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.501283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.501283Z digest=sha256:2d4026c388b6d78d009028152dc155e5d0335b42bcae9f31d47eb3dc37c6e895

Observation 5c478112-4f81-4129-80cc-85d947720eb5 · outbound

This paper cites Promoting Efficient Reasoning with Verifiable Stepwise Reward.

rStar2-Agent: Agentic Reasoning Technical Report Promoting Efficient Reasoning with Verifiable Stepwise Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.505938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.505938Z digest=sha256:38b265cbc7ea92bf01c58c1ea70764586625ac53c0b1e7115586baa8c518041a

Observation 19ad39bf-ede6-4d10-b885-327da8382ff2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

rStar2-Agent: Agentic Reasoning Technical Report Instruction-Following Evaluation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.510592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.510592Z digest=sha256:9d6c7ed0a095474f1e3b7d0c7c545e3beb1aa7b0e456b13c6f7182bb6fbdabb1

Observation aa9b1cef-8270-4200-9f95-350fa9bf2956 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

rStar2-Agent: Agentic Reasoning Technical Report ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.432202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.432202Z digest=sha256:d71f658d751cdb3bbe1398fa99cb13faa014cffe2c9315b8183ce9e7f053a768

Observation ee948300-1b9d-4fc3-912c-928028a5fbf9 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

rStar2-Agent: Agentic Reasoning Technical Report From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.423424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.423424Z digest=sha256:a40f665fdd71012fa471780be2ad0132a5e11e460a9fb157a6df83800ee60581

Observation d4edf230-9a3f-4cbd-8ca3-7c9b740426ab · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

rStar2-Agent: Agentic Reasoning Technical Report MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.374889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.374889Z digest=sha256:0c69577d101a8d4f8ed0d747b329754cbba82053453f8b7354d6ac69a6b84da4

Pith citing papers

Observation 46961b85-16bb-4847-91a5-50b12008bd73 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey rStar2-Agent: Agentic Reasoning Technical Report

Reference 123

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.445880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:20618d49428502de4a461924b56381aedb032e521d6b146c086229c5091ed42f

Observation 03de0ad8-bd7c-4deb-9eb4-a8e79f32dea3 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving rStar2-Agent: Agentic Reasoning Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:25.785903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:25.785903Z digest=sha256:098a4d0a986202bc2d5f0391038626fd62104f46e1b41a6f0431363149abeba9

Observation 6ba424a7-071a-451d-853e-4bd2233f05a6 · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:41.340195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:41.340195Z digest=sha256:a38c5f18f7116fd7ff2591f2b25f10c2ba79a266915cea2cc05e11cb22f3012c

Observation d2bb13be-c000-42e5-a325-75a7376d07c8 · inbound

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning cites this paper.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:49.933722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:49.933722Z digest=sha256:892b644c7695be49594ef8becad2ff45df73510eb1910ad952eb26d1e4ef50f0

Observation f44f5df8-ac25-45ac-9c26-8ba6bfa81712 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste rStar2-Agent: Agentic Reasoning Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:54.599326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:54.599326Z digest=sha256:2ae0c0763d722ff25fc1a2c707e34697bbfc81b036af93c36dc0fe221cb1a493

Observation 7d6c79d8-0511-48b1-a72b-e05d4ffab47b · inbound

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning cites this paper.

SHAPE: Stage-aware Hierarchical Advantage via Potential Estimation for LLM Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:20:50.855350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:14:10.609406Z digest=sha256:3f1180d92c741d0c33a4617caedf771b6b3899cd180f10ddf56792bb80c54ff6

Observation 52ae9bc8-3b34-4067-8e51-a32db106e2ed · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory rStar2-Agent: Agentic Reasoning Technical Report

Reference 224

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:05.839585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:58b70c75b9cfac8fd833cdc686ef5f5e001327bb4930bf5dbec0eebd6a8cee95

Observation 6c362dbf-4887-4fc2-bfe5-842379e3e598 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training rStar2-Agent: Agentic Reasoning Technical Report

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:11.260353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:eac53389d641e24125707ec91178e598b574c733ee1a2da5aa67380caaba5691

Observation fdc7fda0-30ad-431e-8f56-c5fc5287af24 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.161743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:56752d4c492f09e0554a2ff17a1f444824abc5e236ed616120f05b2233e2e5a4

Observation e1aaefe0-d289-4c88-954d-bcf99c6a9463 · inbound

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL cites this paper.

Every Step Counts: Step-Level Credit Assignment for Tool-Integrated Text-to-SQL rStar2-Agent: Agentic Reasoning Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:16:10.568637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T16:23:13.535140Z digest=sha256:de65179a64aadbd98c780d2152445e82397fc04eb27902c4ade445539c71d2d6

Observation a3e970f5-5a02-484a-a52b-31115fe346fe · inbound

Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning cites this paper.

Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:01:10.556776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T10:30:36.330305Z digest=sha256:c04b38525103f3af41a48e22b96f8fd48012259b7eaed0a49f25d8d5008001b3

Observation 9e213d58-4445-4618-9bf7-ef103d03f21a · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code rStar2-Agent: Agentic Reasoning Technical Report

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:15:56.569722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T02:33:02.995360Z digest=sha256:73882899930eb568e7dab7082930141d7c83ed70e622845a88adcc72bb5b7950

Observation ba510941-67cf-4756-b7fb-9bfc7df61b78 · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code rStar2-Agent: Agentic Reasoning Technical Report

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.333349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T04:26:11.265781Z digest=sha256:33d3b9b692dfef9cf5029c2ca636850ac47b49714dd544f99772c1a4a36d42cf

Observation c1a0a899-ee28-4c80-a2d1-6adafbc0b040 · inbound

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning cites this paper.

Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:23:14.031521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:21:39.492072Z digest=sha256:15c76f30127bacb15f5a51b99e254cb4b9e0e721c1d2f8816a923e63007ebc1d

Observation c6f55ad0-8721-44ef-ae29-20152abc4df5 · inbound

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming cites this paper.

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming rStar2-Agent: Agentic Reasoning Technical Report

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.161949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T13:19:22.541146Z digest=sha256:9a722b4c56e6cee996082d0de87159c1b84835e4eaa4c1d73f9719a3680b3eeb

Observation 20f9e4b4-7ab9-482f-b412-fd80cc2ac8c7 · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning rStar2-Agent: Agentic Reasoning Technical Report

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:23.734275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:6904ee4759f3ffda4ba47e3422c3ad379d41810d4ab25b810f924c957cc2cf63

Observation cc52f01b-67d1-467f-b515-c6bb9a6670df · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning rStar2-Agent: Agentic Reasoning Technical Report

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.248154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:d420fe93e02474fd044ebc60205032143f8c0d1066091be62c3e94acf36b59a1

Observation 36d350a3-81e4-46fa-91dc-29c9616c577d · inbound

Discovering Millions of Interpretable Features with Sparse Autoencoders cites this paper.

Discovering Millions of Interpretable Features with Sparse Autoencoders rStar2-Agent: Agentic Reasoning Technical Report

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.059796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T05:34:00.754172Z digest=sha256:9aa06729a8a86676bc71e0977ba4222e6d95490042979cf8af4d3f15866329f1

Observation 0b32f477-f016-4790-b9ca-9fb3fd6adcbc · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language rStar2-Agent: Agentic Reasoning Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:49.632692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:49.632692Z digest=sha256:62667754555fadacf9cd3811cf5c97200e728d40526b6e904d2d4b1ea7f2a610