Pith. sign in

Paper Citation Record · LEDGER

ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2406.03816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.03816 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:40.771963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.622776Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d718454-7d82-4fec-81f0-200f2633fc22 · inbound

Improve Mathematical Reasoning in Language Models by Automated Process Supervision cites this paper.

Improve Mathematical Reasoning in Language Models by Automated Process Supervision ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:45.987128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:53:45.878221Z digest=sha256:c784aa81e5955b27cf8d89d65e526327652c0e4f260f55d6c7df20ad7d967a37

Observation fdc81c39-977f-4985-87d5-629e578dcf47 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.421827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:f0f69699d5cfda5d4c7a9e0c151a43ef2b1c90af55981e10b066eb5e38e32f9d

Observation 42aaf597-5211-412e-b112-53e20c5aebc7 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.771963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.771963Z digest=sha256:e43534426d41aaf2469cc55a0133a13dc0f04a8c2fffb2501e54fd03e76284ca

Observation 1ec4b393-0e04-4406-a39f-d5adf40228bc · inbound

Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models cites this paper.

Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T00:27:17.680587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:27:17.680587Z digest=sha256:8689b2752a5911f2480557313849d0e58079bf47819d15fa4c626c5c96862133

Observation 19cb2dbc-c86d-4fa6-aa64-4abbd030d2d6 · inbound

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking cites this paper.

Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T21:40:26.993871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:40:26.993871Z digest=sha256:868fe4dce6131d5f8d0756c74c183b73403af9995f8e3247b62294b5fa05e926

Observation 2c915cf3-30ac-4f5a-a29b-5b1223555aed · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.446974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.446974Z digest=sha256:c02728ccca468c7f2b0b72b31a4f4f883eb2f1339fa7fbacaf155efede8d1006

Observation e4199f7f-f4b0-424c-8252-8677147d40fd · inbound

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition cites this paper.

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:53.635679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:25:53.635679Z digest=sha256:fec0dced225f08db1a880c586c1273d5e41c5051dbb91e49dfa00c6c5b43c6df

Observation d9fb5c6d-7ef9-4cc8-9766-d999c3b3bc77 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 172

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.208820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:89645eb49fc1cff052683b0ffa1fb70b452012460bd65fc4232da324777a0b22

Observation db777363-2f3f-434a-b00f-222e6d3c1778 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.788835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:39125c4c5770cffd7214509d61bcc7b50e8c33a753d92ecc9a582320d4c98d93

Observation 5ab26c4c-3faf-40a7-a2b6-f62068ae0116 · inbound

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning cites this paper.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.099971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.099971Z digest=sha256:72b918157601099856d0f89af90608c5a0ced8a7565b967a001028cdc8759f62

Observation 657dc109-a0cd-412b-860d-200919d6f8d1 · inbound

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering cites this paper.

O$^2$-Searcher: A Searching-based Agent Model for Open-Domain Open-Ended Question Answering ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:53.041020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:01:53.041020Z digest=sha256:2824165b9122667162ef115c36ec3ac9e3a30a6ba7118191d3925f4f8b43482e

Observation 8f62cfd8-7bdb-4bc1-b40d-e42e352344dc · inbound

Fostering Video Reasoning via Next-Event Prediction cites this paper.

Fostering Video Reasoning via Next-Event Prediction ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:45.635718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:45.635718Z digest=sha256:31d935e1e8a19bc7fe42de2408d93d42339b1e86f06449e8d969350bf3317c6a

Observation b468fc3c-b25a-4c5b-9425-e2a8afe5b8c3 · inbound

Structured Pruning for Diverse Best-of-N Reasoning Optimization cites this paper.

Structured Pruning for Diverse Best-of-N Reasoning Optimization ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:23.724916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:23.724916Z digest=sha256:389885146588027974b3180241d63bd82d1a68b4aa77501ee1ed6b01f35b230d

Observation cc91fd04-9f02-4dc1-80dc-a1c9320020f8 · inbound

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning cites this paper.

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:38.545811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:38.545811Z digest=sha256:c883511a05200dcdd1579b8e66bb79cdfb21307d1406b6332e5b594bf67ea8a1

Observation f4a3a9e7-5c69-4fc8-bffa-56f9115de103 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:23.240182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:23.240182Z digest=sha256:c7e6fbcb533d82a36a5d2e70c2d5561e85098fd42b485db27a1c0306789a2409

Observation 835d2202-c34e-47ce-8dd7-cbb7280c706c · inbound

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search cites this paper.

TreeRL: LLM Reinforcement Learning with On-Policy Tree Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:25.265647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:25.265647Z digest=sha256:67d1bbd775359bce853ca0a250ee054551007e8dbc72aa05134ae729e516b31a

Observation e9bd3714-6d04-4bae-b32b-264eec0dba3e · inbound

Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs cites this paper.

Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:54.490103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:03:54.490103Z digest=sha256:c358ffbb521b6202ec15550d8d08940d2697c20cd7ff5195f96e1303eccba552

Observation 5be8b77b-eb10-4589-9779-57bf1ac7d3d4 · inbound

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments cites this paper.

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:16.296100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:16.296100Z digest=sha256:fc3edb60248157553515b3c8691dbe3b354400f6736938e400aacfed4a4ce237

Observation 60d23114-b4cc-4d04-a3e8-356b4f0a3d6f · inbound

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique cites this paper.

EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:04:55.263407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:04:55.263407Z digest=sha256:961f92aef7d5bb18c8dcc2acfebb6a437a1168820c57cde11ca93ed6ff0d5499

Observation 4404c5fd-2fa1-423a-a3b1-419b3d1c9fc1 · inbound

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL cites this paper.

Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T04:42:06.478842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:42:06.478842Z digest=sha256:574b12597db0ae3a09a7a82c6a210aa85cedddd71644dd72fd8677e6e29a155f

Observation 55537287-1e32-4ea5-97b8-a6533ee96552 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:27.771393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:27.771393Z digest=sha256:69002e9ea36674183358d048dadfd7b9bf0ee590f1b9abf3b28630a5cc4673e3

Observation 96cc5f10-7a1c-42e7-95eb-358b94a1a8c9 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:02.754586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:7279f091e8aa5b3d1a3b741324cab1f68a4af02d95ca15c5dd6b7a4575a47ce1

Observation 6856d430-8a02-4ba5-a41e-b32e1038547c · inbound

PARM: Pipeline-Adapted Reward Model cites this paper.

PARM: Pipeline-Adapted Reward Model ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.512713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:15:26.015817Z digest=sha256:ba454fee7aff828cbf122c27afe18d9eebae33ce1eb1e0ec4f580008571ef1a4

Observation 77e634fa-d469-4914-96f5-b82c1ad7af7e · inbound

Self-Improvement for Fast, High-Quality Plan Generation cites this paper.

Self-Improvement for Fast, High-Quality Plan Generation ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:30.662509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T04:13:17.612428Z digest=sha256:74bb3b8ffb1e9935c5c0a304ace80a20b15130e950c569d92fecdaa53bdcca8b

Observation 96f5121e-c87d-416a-abe8-776e950c0f9c · inbound

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning cites this paper.

Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:06.893478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:16:00.499779Z digest=sha256:1f555f6fc5ea99116b9da41ed8a35f0e7635cb97f70efd764bb9287d6cbd3d90

Observation 3141565e-11ba-4c49-99e7-97baa43cf0d9 · inbound

Efficient Test-time Inference for Generative Planning Models with OCL Search cites this paper.

Efficient Test-time Inference for Generative Planning Models with OCL Search ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:35.134640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T18:53:08.784890Z digest=sha256:c17757560e617d42aa2ade3ed761ade418839efb21bc8f2e876a2c17ff7ccfca

Observation c8b26145-bd85-472d-ac7c-7dd1757adcc6 · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.915321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:123f4779732dd7ee90c99ba526f1bb88772e9759c36779af5492c39168aad73a

Observation fc6e7769-7fc5-47cd-a782-be6c4bdcb884 · inbound

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States cites this paper.

Theoria: Rewrite-Acceptability Verification over Informal Reasoning States ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.432836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T12:12:20.879377Z digest=sha256:1162e30bfbc5686bd50abb3e7cf42496917b253f877e06f0a5b563cd7d108974

Observation d5ea40c1-a14f-41a6-8463-a32c16913dd9 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.624038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:22eab0e50f2a95b26d6e07f474a98ef6df34127131678189cdadea511c5c5299

Observation 98df491d-4780-4f9e-9cc5-a0fea5e01a3a · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:35f8e4594412845d4327c71294846f5835e1843923722a7c669545cd1aa94b47

Observation 02923fb9-c267-4ac0-a44d-b3571225cc57 · inbound

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents cites this paper.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.726132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.726132Z digest=sha256:e34d32fb1dbd29df0a73ab50d083cafb95216a4ac0198506bc0fb179f23d5933

Observation 76aac002-9dd0-447e-aa6e-1327ef9e72ab · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:24.768710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:24.768710Z digest=sha256:34f9f6a8989fbf746b77a0bd309939cd7e2d2810c5818ca9fd147ce575108ba7