Pith. sign in

Paper Citation Record · LEDGER

Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2504.04736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.151443Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:26:23.083412Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 941ac46f-d5da-45aa-9232-a13a9bfc8ad8 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.259463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:f971bd7ffad20de00a6210d05d23b38e3648a11fb6123589e82c6379c3c2f21d

Observation a6bd3552-31f0-4518-b16d-d01d8f677910 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.151443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:26.151443Z digest=sha256:9e6f132747fd9393094f09626e85a681aaa98749103c38b12d4142d1f7d1eb8d

Observation 71362ded-996a-4e10-8c80-66dfb4b9f1e6 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:23.355764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:23.355764Z digest=sha256:255b788f3a641e20caa2f1f92c8b0bfc0084c8092ab898da108e29ffad8aa1c8

Observation c96330a2-5f5e-4403-b8c0-954be3c8e936 · inbound

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review cites this paper.

CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:28.578751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:28.578751Z digest=sha256:dd3b67857364100ece145683b5e05f7ce14a4a2b382da7c3da7026215f6cd6ba

Observation a8369b8b-1fca-444f-92c3-4f022cc3f304 · inbound

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models cites this paper.

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:41.804315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:41.804315Z digest=sha256:aa97f204fb34f7c3e6053d85114d315d90b364210d1a62259ac4d46981c616c0

Observation 3a734589-8f92-48a9-b25d-56e36e6999b1 · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:39.115854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:39.115854Z digest=sha256:31104e6fb45aa28f4b64f552b4083f357f9e32eceef8a212bbab9b5a3fcf6db4

Observation 288dda77-2013-491c-8b05-cff357d26bd1 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.609002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.609002Z digest=sha256:c715bb95ab869e868b1af59e9d24b2a8926f00b0536dce757f504162843487a0

Observation a150e92a-810d-4da8-abbe-a176a4b9352d · inbound

Distilling Tool Knowledge into Language Models via Back-Translated Traces cites this paper.

Distilling Tool Knowledge into Language Models via Back-Translated Traces Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:58.302886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:58.302886Z digest=sha256:f2ce3dba059e8c7e5b188a86a8dd62b30079e1e43b3b79659ec13e5bd7475d29

Observation 4b64dbb8-5014-4415-9e30-908677aa8367 · inbound

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning cites this paper.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:08.936392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:08.936392Z digest=sha256:74a93edecec5f84701cf0bf0038a6a59858742ecc6c33e29004ee33c54c54080

Observation bd220da0-e9ad-43e9-b6fd-e1996fb0e4de · inbound

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs cites this paper.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:58.523176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:58.523176Z digest=sha256:3d8b867974dfe64da7de8c88498003fe0f5a5aac9b1d3fc079fa3d6c71d4b8c8

Observation cf99c4e5-03f3-4284-9d33-bb448d99d66f · inbound

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs cites this paper.

Step-wise Policy for Rare-tool Knowledge (SPaRK): Offline RL that Drives Diverse Tool Use in LLMs Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:16:54.103748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:16:54.103748Z digest=sha256:e6ec12de8b7aa5d4e3396b1db59ba28f50f792ba24485705d7475e08f0f4a610

Observation 70168ea2-0162-4c4b-90f1-c425f7d5460e · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:21:42.397702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:a6b9f23c0c7dc02849e8c882f0de3515bec97f13941b6b631ce367cbbe0391b6

Observation 1eb84988-09be-4dc8-9337-a1b723f9e5a5 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.383625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.383625Z digest=sha256:554ad88759820057585ae9409ceeff20ff4b02e28de7839db830a413a91c54e0

Observation 9989590c-175f-49ea-992e-207d1eca1144 · inbound

Agent Learning via Early Experience cites this paper.

Agent Learning via Early Experience Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:58.616326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:47:58.616326Z digest=sha256:f1cf5202d2366c03a6834a766fe2896d1789a43d893906a1c629930e4f729dfd

Observation 5bc1d60a-e7fa-42b7-abf2-e4b4564a36b4 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:28:05.728609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:89b18c3be3a85af7cf147d9b6e68b085804092f651e0ec5dd0d3a638d2dfd467

Observation 3f1600b7-245c-486b-aaa5-00cad53aeb2b · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:06.043412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:4b1bf4a9747d0775e0e8fee71d29c6ca4aadf9ecd479cec7af2d17b76cae76e1

Observation c9324866-5083-4271-b0cc-967a7d707468 · inbound

Process Supervision of Confidence Margin for Calibrated LLM Reasoning cites this paper.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.225348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:21077723a163926c12bee91e675f58dd0220f76275553d66c49bdb2f48227c7b

Observation 015f74c6-388b-49af-814b-0079ff9a3add · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:46:06.952986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:08:53.731480Z digest=sha256:274146920717734a0c59ca39f1581133def3447881dd4e39351d82bad8e61381

Observation 4058fa5e-7634-4bd6-8b2c-85de24e3e2f9 · inbound

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data cites this paper.

$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:55:12.125531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T00:48:54.797750Z digest=sha256:bde0897e45dc68dfda3ba0931eae780df0ccbcd6c1acfca9b5407b16df6e67e6

Observation 3d28ceff-6fc2-47a2-b680-5f48a6933e2f · inbound

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning cites this paper.

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:33.632647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:45:06.199636Z digest=sha256:5aec99c73ed281e4e22858cfbcb7135449ecf7c3aa4c0f5bd8fb305bf7a15439

Observation 3556ff3f-d5ed-4060-81a0-293f833c2f98 · inbound

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning cites this paper.

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:23:48.443509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:19:49.016156Z digest=sha256:aa62cb1ca9925d5caf1c578ef5d58dc0aa43d12a11a8e73bbe823d53a8cf9e7c

Observation a62075b5-d03b-49d2-a0e8-d4329cca1248 · inbound

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents cites this paper.

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:26:23.085582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:18:45.821224Z digest=sha256:3cf99662d403f86c01e1e8763e22737387753ade384797a1dcff6d995ab644ef

Observation ff9b56ae-1d57-4c40-9193-11888ee39645 · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.527752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:f02a2abe514fba752561320b4db4f7676beccb539154c7d805a0a756e0663ffe

Observation 218db873-a18b-4dc8-818e-34e6eed4657c · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:11.165028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:11.165028Z digest=sha256:4224f90f7c1fda3a3a4d989d596bb8c4d1196f7c043edc31e55e7ce47596eee2

Observation f7674e92-eec6-4e79-b192-36a5a2754d15 · inbound

Chained Recursive Language Models for Multi-Iteration Reasoning cites this paper.

Chained Recursive Language Models for Multi-Iteration Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:17.289396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:17.289396Z digest=sha256:f350aa3afaa55ec16bad97257e21f4a5a06fdd93d21434162539fe6856c8b5fa