Pith. sign in

Paper Citation Record · LEDGER

ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.10003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:16.864213Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d5d98b6-93e2-4658-8d54-4003accc3b97 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:57:38.252446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:3ec81e37f1dc943f6468f74edba53067dddc890578fd51bd29c1d0d01f7e22d1

Observation e146ce08-8442-47ea-baba-77c1d7e78c41 · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:16.864213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:16.864213Z digest=sha256:289b2f2514b2f9eecc05d84a97096ebe570ec418e9b2dfc223519443af7b6408

Observation 7bf92f9b-25fa-45a2-949f-e52db5e4775e · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:28.337466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:28.337466Z digest=sha256:ddb3e82e35ed7cf06cf5cf63afbbb7b15690b964e70d7e522060108f3a6b4624

Observation af7de01d-ebad-4268-b54c-87434c6f3bcd · inbound

Estimating the Empowerment of Language Model Agents cites this paper.

Estimating the Empowerment of Language Model Agents ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:53:14.599607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:53:14.599607Z digest=sha256:6d073ff8e40adfceb5f91dd87679a0d2f0f60967525fbfef7f19173d5f7ed72d

Observation 70ba945b-1d86-4212-81f1-419d5feffc55 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:15.679804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:e61abb989ca99cfd8c5aec2f65f221d5168ae1d5a3f63c5009db5f2041bc19e0

Observation cb05160e-8cee-452a-8503-e66f5b5c1c96 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.988543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:68dff209e60935906a74d040142676e081b3e83cd46bccd3097b4dab153819aa

Observation c3be1648-17fc-4441-ab3d-2bf62885c8f5 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:32.684964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:32.684964Z digest=sha256:d66758dc34ca9a1d3119209442017cf29947442244f06e8285fec27e9a74b469

Observation 037ae162-fc22-492f-bde3-b91cfe59372d · inbound

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents cites this paper.

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:18:01.300645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T17:16:17.464937Z digest=sha256:1427533d2916a7b21c3dc8d2b7022a6f130c687d1a23eeddaec5736fd526e245

Observation 7258b9e7-2cdf-4f52-9335-e577337d4d82 · inbound

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition cites this paper.

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.922872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T10:30:48.957568Z digest=sha256:a0dbe8c9fcf2993208e3a450ba1fb0c2cfb4b1dea49328b1c12dcdf12ebbb362

Observation 3352d076-4318-4d6b-b64a-b506e727b548 · inbound

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation cites this paper.

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.146779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:29:44.078911Z digest=sha256:826221cedd8f08d4cc31623404cebebeb92fd0b0ff2e0c24f234e077d8514484

Observation 4c09d938-fcb1-48b2-865a-423e2a18bd28 · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.710917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:dd4566730ab2624dd21a22e3bfc852ab36fa19fe0dd8ed99abf1b84d18658c2d

Observation d379452a-44fc-4c5a-bf92-c4ce6d53b990 · inbound

Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts cites this paper.

Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.037482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T16:39:36.945808Z digest=sha256:36d28861242dbd5b1523d8829b3cdb1427ef97fbafd057590fd4710c5835597c

Observation 079675ac-308f-40a4-a1e8-8918767a667b · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:16.726094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:a3bd37767fc5f14143d04393aa656d0fdb387f29e38a9982cfc158519aad2c01

Observation b5f02c49-29fc-4e89-8f0d-210781c0d1e7 · inbound

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting cites this paper.

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:22.277629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T20:40:11.059493Z digest=sha256:9614e06af9fa759a9e4f932c6078c42c37ffc6fc5f0065cbf27c04e20ae54362

Observation 8ceacf65-bebc-4793-b2c4-7fc4ebf54b23 · inbound

LLMs for Agentic Home Energy Management cites this paper.

LLMs for Agentic Home Energy Management ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T17:08:27.234032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:08:27.234032Z digest=sha256:1e6b92c344140eaf5271f92b5fbf672b2d26e265260c6f499710e0469504f6b7

Observation ee669dc4-56fa-4ab4-9f3a-51b5978ba771 · inbound

LLMs for Agentic Home Energy Management cites this paper.

LLMs for Agentic Home Energy Management ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.501969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:33.501969Z digest=sha256:32fed751d329ac4da381ac666cf549dfedadbd7caa7b348bf017e6226e352fcd

Observation 7ffba785-912c-490c-9786-776a4b314105 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:8f74a9f268fdebe08ec8713159f11efa223bacb69e835fc17d31663cf4adac53

Observation 6f322826-12d6-45ff-ba02-ecde46fa59be · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.326974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.326974Z digest=sha256:1ba6741ded7fd1852efc1a7d4c387dd1a35b019a795ac7325100173e33da5aab

Observation a194edc7-af73-4ddd-8eb2-ca258a572d92 · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:28.624570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:28.624570Z digest=sha256:af8482ad11794133bb3bfd2b35b5777a1d3cb6cec693567e76937265050eaf8f

Observation 63ef0a9b-0eb2-4062-acbb-75db8b113e97 · inbound

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks cites this paper.

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:52:44.777206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:52:44.777206Z digest=sha256:b4b35224c3e555c63fe0d66d1e8766b065ed56f1eed662d43f9f4ae0c4175e3f