Pith. sign in

Paper Citation Record · LEDGER

ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.10003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:16.864213Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d5d98b6-93e2-4658-8d54-4003accc3b97 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:57:38.252446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:455603b4477d69fe737d3c2bade21d099ffef654b306e443a58a2ab65e83ead1

Observation e146ce08-8442-47ea-baba-77c1d7e78c41 · inbound

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning cites this paper.

Truly Self-Improving Agents Require Intrinsic Metacognitive Learning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:16.864213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:28:16.864213Z digest=sha256:289b2f2514b2f9eecc05d84a97096ebe570ec418e9b2dfc223519443af7b6408

Observation 7bf92f9b-25fa-45a2-949f-e52db5e4775e · inbound

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives cites this paper.

AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T19:34:28.337466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:34:28.337466Z digest=sha256:ddb3e82e35ed7cf06cf5cf63afbbb7b15690b964e70d7e522060108f3a6b4624

Observation af7de01d-ebad-4268-b54c-87434c6f3bcd · inbound

Estimating the Empowerment of Language Model Agents cites this paper.

Estimating the Empowerment of Language Model Agents ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:53:14.599607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:53:14.599607Z digest=sha256:6d073ff8e40adfceb5f91dd87679a0d2f0f60967525fbfef7f19173d5f7ed72d

Observation 70ba945b-1d86-4212-81f1-419d5feffc55 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 234

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:15.679804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:52824107afaf560f52a0d068f5eb138c7a402c92b3c663b0e330218bfee7a263

Observation cb05160e-8cee-452a-8503-e66f5b5c1c96 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.988543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:fd57242bdffa5b66781d0c5450997c03eb50ea688f9d0d055476b3e75a7bcc6a

Observation c3be1648-17fc-4441-ab3d-2bf62885c8f5 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:32.684964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:32.684964Z digest=sha256:d66758dc34ca9a1d3119209442017cf29947442244f06e8285fec27e9a74b469

Observation 037ae162-fc22-492f-bde3-b91cfe59372d · inbound

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents cites this paper.

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:18:01.300645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T17:16:17.464937Z digest=sha256:eed39d13a55d54addd116d6556a9a22cdf545bdad788ade1439be626ee5549bb

Observation 7258b9e7-2cdf-4f52-9335-e577337d4d82 · inbound

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition cites this paper.

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.922872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T10:30:48.957568Z digest=sha256:735ec437946154ef4c6be7b3c26bf25e02cf77bd0e59f9348a01a6c5bfee17a0

Observation 3352d076-4318-4d6b-b64a-b506e727b548 · inbound

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation cites this paper.

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.146779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:29:44.078911Z digest=sha256:851ec47c0eeedd4338ef6479abfd4a61833f73b123092dd181a9c047b2bc0af1

Observation 4c09d938-fcb1-48b2-865a-423e2a18bd28 · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.710917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:ba827d86c159aa71eaf545ae8215fcf90ab226b3a8276d8c46e2e61b35b4d49d

Observation d379452a-44fc-4c5a-bf92-c4ce6d53b990 · inbound

Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts cites this paper.

Graph2Idea:Retrieval-Augmented Scientific Idea Generation with Graph-Structured Contexts ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.037482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:39:36.945808Z digest=sha256:8f8b9487329951acc6ca576000a556ac1a293c527f7db98c962f7a46f4faa3a2

Observation 079675ac-308f-40a4-a1e8-8918767a667b · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:16.726094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:67da7d856bd345d10888f87b47eeefa636374ac4b4cf297bb3afb17374840c66

Observation b5f02c49-29fc-4e89-8f0d-210781c0d1e7 · inbound

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting cites this paper.

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:22.277629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T20:40:11.059493Z digest=sha256:0292fc9c77cf13f4384b07b80ffade541b23ddf0c1f1c0fee604b079d77fb3c3

Observation 8ceacf65-bebc-4793-b2c4-7fc4ebf54b23 · inbound

LLMs for Agentic Home Energy Management cites this paper.

LLMs for Agentic Home Energy Management ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T17:08:27.234032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:08:27.234032Z digest=sha256:1e6b92c344140eaf5271f92b5fbf672b2d26e265260c6f499710e0469504f6b7

Observation ee669dc4-56fa-4ab4-9f3a-51b5978ba771 · inbound

LLMs for Agentic Home Energy Management cites this paper.

LLMs for Agentic Home Energy Management ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:33.501969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:33.501969Z digest=sha256:32fed751d329ac4da381ac666cf549dfedadbd7caa7b348bf017e6226e352fcd

Observation 7ffba785-912c-490c-9786-776a4b314105 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:8f74a9f268fdebe08ec8713159f11efa223bacb69e835fc17d31663cf4adac53

Observation 6f322826-12d6-45ff-ba02-ecde46fa59be · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:40.326974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:40.326974Z digest=sha256:1ba6741ded7fd1852efc1a7d4c387dd1a35b019a795ac7325100173e33da5aab

Observation a194edc7-af73-4ddd-8eb2-ca258a572d92 · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:28.624570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:28.624570Z digest=sha256:af8482ad11794133bb3bfd2b35b5777a1d3cb6cec693567e76937265050eaf8f

Observation 63ef0a9b-0eb2-4062-acbb-75db8b113e97 · inbound

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks cites this paper.

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:52:44.777206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:52:44.777206Z digest=sha256:bcb7c936a1486c0e93544a802065f1b7ffe6cc4945c07dbf60068c5b831d2b81