Pith. sign in

Paper Citation Record · LEDGER

Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.00658.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.00658 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:21:05.597420Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:42:26.911183Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5709a708-bbc7-4505-815a-6730cf786dff · inbound

Preference Optimization for Reasoning with Pseudo Feedback cites this paper.

Preference Optimization for Reasoning with Pseudo Feedback Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:05.597420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:05.597420Z digest=sha256:64a3f3575ebf148c1e9b8cdb9b439147e6970b8446f73f916090338212345278

Observation da187040-5038-43de-a5f9-d2fffd4447d7 · inbound

Mars-PO: Multi-Agent Reasoning System Preference Optimization cites this paper.

Mars-PO: Multi-Agent Reasoning System Preference Optimization Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:39:36.879146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:39:36.879146Z digest=sha256:82c12a86d7b0d47f664f874932e4bf02c02fd77466a49fdb3ac67987ca834800

Observation 7fd0cef1-7967-4e70-b4d2-e75dd1d7d11b · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.597847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.597847Z digest=sha256:2928f9c9c8c60f84272bff61aaffef2a9d553aceb1fb674800e46a4f9fd2daf6

Observation 7749b063-e0fb-4649-b3c8-43bf035ed78e · inbound

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models cites this paper.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:42:27.076360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:42:22.042716Z digest=sha256:00c4aa82e8a8037857cd7f9bf47f7df21373e2e6a83a610daf32d71f0e2744fe

Observation 0c4735c7-4597-42d8-a28f-021c161a5a18 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:57.753578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:57.753578Z digest=sha256:cd365da0aaa8547f3fdf05b9d7d93def9657a3f645e2db6566294db34d5ea419

Observation 990160fb-6307-4daa-9858-00d3fdbbe47c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:d0320418ca941cb86b470fa488fee4d50b294394c98674618066954f2b4d15ef

Observation 594a8fd3-7461-4573-91c8-b2fdff369f9e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.861888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.861888Z digest=sha256:99c28d2d48ac28f361da5e009b957c16621bd1db066690f25509c12b9a23483a