Pith. sign in

Paper Citation Record · LEDGER

Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2504.04736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.151443Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:26:23.083412Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 941ac46f-d5da-45aa-9232-a13a9bfc8ad8 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.259463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:f5687d8bedf1596c1b863169a5d993206b165fb2a4208668d2c1defeeb2bafc8

Observation a6bd3552-31f0-4518-b16d-d01d8f677910 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.151443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.151443Z digest=sha256:9e6f132747fd9393094f09626e85a681aaa98749103c38b12d4142d1f7d1eb8d

Observation 71362ded-996a-4e10-8c80-66dfb4b9f1e6 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:23.355764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:23.355764Z digest=sha256:255b788f3a641e20caa2f1f92c8b0bfc0084c8092ab898da108e29ffad8aa1c8

Observation c96330a2-5f5e-4403-b8c0-954be3c8e936 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:28.578751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:28.578751Z digest=sha256:62c24a6aa4c77fefa46e25f46adc6e6d081f8606c206c9e64b1628ee8e07bf9a

Observation a8369b8b-1fca-444f-92c3-4f022cc3f304 · inbound

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models cites this paper.

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:41.804315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:41.804315Z digest=sha256:aa97f204fb34f7c3e6053d85114d315d90b364210d1a62259ac4d46981c616c0

Observation 3a734589-8f92-48a9-b25d-56e36e6999b1 · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:39.115854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:39.115854Z digest=sha256:4e73ff2f0711104426f4ffc31bd95441dbae3f1e1e629b4272d996ebc1ae2cc1

Observation 288dda77-2013-491c-8b05-cff357d26bd1 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.609002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.609002Z digest=sha256:c715bb95ab869e868b1af59e9d24b2a8926f00b0536dce757f504162843487a0

Observation a150e92a-810d-4da8-abbe-a176a4b9352d · inbound

Distilling Tool Knowledge into Language Models via Back-Translated Traces cites this paper.

Distilling Tool Knowledge into Language Models via Back-Translated Traces Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:58.302886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:58.302886Z digest=sha256:f2ce3dba059e8c7e5b188a86a8dd62b30079e1e43b3b79659ec13e5bd7475d29

Observation 4b64dbb8-5014-4415-9e30-908677aa8367 · inbound

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning cites this paper.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:08.936392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:08.936392Z digest=sha256:5389663d1582994d4d72f4c2ce3b84b3bfdb96fe10a64c840ea4d5dfee3b5026

Observation bd220da0-e9ad-43e9-b6fd-e1996fb0e4de · inbound

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs cites this paper.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.523176Z digest=sha256:3d8b867974dfe64da7de8c88498003fe0f5a5aac9b1d3fc079fa3d6c71d4b8c8

Observation cf99c4e5-03f3-4284-9d33-bb448d99d66f · inbound

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs cites this paper.

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:16:54.103748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:16:54.103748Z digest=sha256:e6ec12de8b7aa5d4e3396b1db59ba28f50f792ba24485705d7475e08f0f4a610

Observation 70168ea2-0162-4c4b-90f1-c425f7d5460e · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:21:42.397702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:97f09f757fc8cf37a6bd1ccaded27ce6e2346be751f12378050fa7724c235518

Observation 1eb84988-09be-4dc8-9337-a1b723f9e5a5 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.383625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.383625Z digest=sha256:554ad88759820057585ae9409ceeff20ff4b02e28de7839db830a413a91c54e0

Observation 9989590c-175f-49ea-992e-207d1eca1144 · inbound

Agent Learning via Early Experience cites this paper.

Agent Learning via Early Experience Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.616326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.616326Z digest=sha256:f1cf5202d2366c03a6834a766fe2896d1789a43d893906a1c629930e4f729dfd

Observation 5bc1d60a-e7fa-42b7-abf2-e4b4564a36b4 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.728609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:a3366887c15d69deb0a055769c7c2552d0ad428d45d2b1e9d09d02ff27b7aba4

Observation 3f1600b7-245c-486b-aaa5-00cad53aeb2b · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:06.043412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:ae44cd16246afae14531da170c2c811b813757dcd63f87bfdcd7bf90d809349d

Observation c9324866-5083-4271-b0cc-967a7d707468 · inbound

Process Supervision of Confidence Margin for Calibrated LLM Reasoning cites this paper.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.225348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:a293c4994bbf0a0f7ef182deee9ed1bff9cb26e53eceb58e3143b2a087602f05

Observation 015f74c6-388b-49af-814b-0079ff9a3add · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.952986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:08:53.731480Z digest=sha256:83501ba874a9e9af42a1feff8f695f9487c932a5438dad921d7362e873ba4b71

Observation 4058fa5e-7634-4bd6-8b2c-85de24e3e2f9 · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:55:12.125531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:48:54.797750Z digest=sha256:007d6baa314d08acfd363b5d391fe314ba835354ac5591a3982be17f0de415c1

Observation 3d28ceff-6fc2-47a2-b680-5f48a6933e2f · inbound

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning cites this paper.

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:33.632647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:45:06.199636Z digest=sha256:93341db44350616c8872ad1aed7fec18c971d8f72a03c9db23e24028d415d1cf

Observation 3556ff3f-d5ed-4060-81a0-293f833c2f98 · inbound

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning cites this paper.

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:23:48.443509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:19:49.016156Z digest=sha256:6f3e6aacdae83ffa313f3115b06206a72cf3169ad19bbeda6c11f1dff3dd6d71

Observation a62075b5-d03b-49d2-a0e8-d4329cca1248 · inbound

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents cites this paper.

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:26:23.085582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:18:45.821224Z digest=sha256:1a714586194f4b55e6bb94abf681b165d1565d0a187bf0dc7299c15ec08918da

Observation ff9b56ae-1d57-4c40-9193-11888ee39645 · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.527752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:39881b6ec2afbc7ff9f6cf55f8e3b9356526caec3cb2e8e43f103a8154286265

Observation 218db873-a18b-4dc8-818e-34e6eed4657c · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:11.165028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:11.165028Z digest=sha256:4224f90f7c1fda3a3a4d989d596bb8c4d1196f7c043edc31e55e7ce47596eee2

Observation f7674e92-eec6-4e79-b192-36a5a2754d15 · inbound

Chained Recursive Language Models for Multi-Iteration Reasoning cites this paper.

Chained Recursive Language Models for Multi-Iteration Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:17.289396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:17.289396Z digest=sha256:0ac6a32bc541770fe8e101179793560d238d9b96f78ae33647751f30ccf839aa