Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2310.08118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.08118 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:11:28.174002Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T11:05:42.217787Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0eaa7b4b-6f74-45f9-a1be-e544b5295761 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 232

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.544552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:d7ec8d259d57401e488c9d1cef83a17e43c452a1904f67703e78841463236ea8

Observation 3cc974f2-a2c4-4171-8590-00b868f03bd0 · inbound

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? cites this paper.

Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:28.174002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:28.174002Z digest=sha256:33bb8d72774efa17fa90585158bb3ab4ca0fa6ab213861cdba773229df88bb3c

Observation 1d07930e-325d-4ba0-a1bd-68fc220040b9 · inbound

S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency cites this paper.

S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T21:33:37.640825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:33:37.640825Z digest=sha256:9d32f4c72b365b7b407c9a7f5b269351cd892fec4ab0611b8818038a161b4c0e

Observation e4360466-1e79-4357-9803-c21c70600514 · inbound

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models cites this paper.

Exchange of Perspective Prompting Enhances Reasoning in Large Language Models Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:00.034064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:00.034064Z digest=sha256:66ae49fdcacf69b48596906db15815626d23c4a2cccd10206e8b9f0a9e7fb51f

Observation 1839f655-5ea1-40b2-87e4-919933da9d28 · inbound

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms cites this paper.

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.395654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.395654Z digest=sha256:f1734ee7b00c351d220c7e0f18db807c034a09b1cf6d3e234ac802fa69600b51

Observation fc1db1c5-5c14-49db-8e7c-7a23c02e4b59 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.287151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:0b32d71e137935204ffe8d10d5a07c13298561f86740c0472b48f77bc3eedf72

Observation 62d2848f-8c45-4134-a8c1-ca9a5060a426 · inbound

End-to-end PDDL Planning with Hardcoded and Dynamic Agents cites this paper.

End-to-end PDDL Planning with Hardcoded and Dynamic Agents Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:38:41.621289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T23:34:20.411940Z digest=sha256:9096a58e3617dc710bd211422ac94ea82871ec2811472df73f3048c59c351009

Observation 5218f3c7-8cc0-4bda-93fa-e2492323df5e · inbound

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding cites this paper.

Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T04:06:11.908439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:06:11.908439Z digest=sha256:fe493ca35dd750e68c620289b4274832bfe41c46d51b166d89608cf996450ad0

Observation 5258a308-80a1-47b3-9b9d-f41dc493e3a2 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.655422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:a6cd7ee26543f903bf85ca93c0946841e928867725761009f35b190a9eb57486

Observation 8ae7bdbd-a4bb-46d7-9beb-345465abd922 · inbound

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning cites this paper.

Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.650977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T20:46:31.368628Z digest=sha256:4d3fe4ead5f9cd26093b61d85688c84bf439d3edf77ebb07e9cd0f627d91d075

Observation 67c87912-15ca-415d-864e-66c22e0171fe · inbound

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers cites this paper.

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:42:46.127304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T20:41:48.063871Z digest=sha256:324f7502bed04eb4175f7c4ef1e9aa22950932939abf8bd7ce2883c4e73ee9ec

Observation b6a978cc-5345-4543-bba1-919efe96c1e3 · inbound

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models cites this paper.

Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:05:42.219249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T04:44:56.520156Z digest=sha256:dc32b4e6c1a6257c8fabdcb9c26da77e1728b00b35f4982fbef3c72f12f8a907