Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:16:43.153818Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2608.00017.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:16:43.153818Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fd057453-96b2-4df1-a1ab-376a65485ac9 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Sutton and Andrew G
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 447acec5-3a1d-4f99-958d-76c0f90fe96f · outbound
Memory Reward Inflation in Self-Improving LLM Agents Reflexion: Language agents with verbal reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f4abcf-f05e-4d6c-bce7-1f0790d49458 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Self-refine: Iterative refinement with self-feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eae2184-c78d-4153-b98b-ae7922e7dc3b · outbound
Memory Reward Inflation in Self-Improving LLM Agents Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7930bc81-c268-43a0-9e0b-f5d687120db2 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Christiano, Jan Leike, Tom B
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5786c519-f5b3-41e4-842d-710903a365b2 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c6b4a9-8649-4919-83f0-a8c63062c1aa · outbound
Memory Reward Inflation in Self-Improving LLM Agents Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e747de1-a639-4923-9c64-bc7ae891cec7 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Spontaneous Reward Hacking in Iterative Self-Refinement
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b6a77f-a2ce-47d1-bcad-4aa26abeec89 · outbound
Memory Reward Inflation in Self-Improving LLM Agents ReAct: Synergizing reasoning and acting in language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efce4652-528d-4fbf-b651-7b5ec098f964 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Le, and Denny Zhou
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f74aae-e796-4c1a-a602-ca878731518f · outbound
Memory Reward Inflation in Self-Improving LLM Agents Retrieval-augmented generation for knowledge-intensive NLP tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6f0ec12-58c4-472a-acd2-36a4f878530d · outbound
Memory Reward Inflation in Self-Improving LLM Agents O’Brien, Carrie J
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f525c53-dfb7-4ca9-bea7-94f9b8ef7b97 · outbound
Memory Reward Inflation in Self-Improving LLM Agents ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3908e104-b9e0-4dac-8be4-f7b14aeb47bd · outbound
Memory Reward Inflation in Self-Improving LLM Agents MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ce45e0-ef31-4735-8a5e-1e9c439c1224 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ca1f9d-4d80-4d02-9ffe-3b8d22a1f76f · outbound
Memory Reward Inflation in Self-Improving LLM Agents Mem-{\alpha}: Learning Memory Construction via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c472d61b-1dc4-45c2-841a-9b957d0c94ef · outbound
Memory Reward Inflation in Self-Improving LLM Agents Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38068a7-e067-42f2-8516-ed08db3d8f06 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53f58fce-b58a-4ba1-b605-2a8acaf8246a · outbound
Memory Reward Inflation in Self-Improving LLM Agents Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82440356-8ab7-48d8-9169-ed5612ca0019 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Process reward models that think.arXiv preprint arXiv:2504.16828, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4080ab-d542-4623-b6fc-f8b558591d70 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Self-Preference Bias in LLM-as-a-Judge
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4352c5a-0e5b-440b-af5c-2a8cb2ba9922 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Beyond the Surface: Measuring Self-Preference in LLM Judgments
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5bacfb-663a-450b-a497-04c25024b9a9 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93271f3e-459c-4d8a-a4a1-31539340c8a6 · outbound
Memory Reward Inflation in Self-Improving LLM Agents QuickCheck: A lightweight tool for random testing of Haskell programs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af03401f-98e5-498a-a261-110f4ce576aa · outbound
Memory Reward Inflation in Self-Improving LLM Agents Finding and understanding bugs in C compilers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ef34c9-c955-4cc7-b7c5-511c74ad8216 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Metamorphic testing: A new approach for generating next test cases
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 007d0614-1a62-46dc-a84f-6dcd08ccd8f5 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Sanchez, and Antonio Ruiz-Cortés
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627f7534-db2a-4faf-84b4-54f9f4c36b80 · outbound
Memory Reward Inflation in Self-Improving LLM Agents CodeT: Code Generation with Generated Tests
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789a2fc0-204f-4784-9c26-0b7f2f84402e · outbound
Memory Reward Inflation in Self-Improving LLM Agents SEDM: Scalable self-evolving distributed memory for agents.arXiv preprint arXiv:2509.09498, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92c542c-80b1-4a3c-93e5-ed051c3adf30 · outbound
Memory Reward Inflation in Self-Improving LLM Agents A-MemGuard: A proactive defense framework for LLM-based agent memory.arXiv preprint arXiv:2510.02373, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6697af4e-3474-45b7-880a-d8c1346fca88 · outbound
Memory Reward Inflation in Self-Improving LLM Agents MemMA: Coordinating the memory cycle through multi-agent reasoning and in-situ self-evolution.arXiv preprint arXiv:2603.18718, 2026
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513546c6-3ba1-4884-aca6-dd24b0455df1 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Useful Memories Become Faulty When Continuously Updated by LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982f9154-3941-4ad3-9d34-9ddeabfc1b34 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Large Language Models Cannot Self-Correct Reasoning Yet
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ca1904-22a6-4d15-8504-7dec3c645ce7 · outbound
Memory Reward Inflation in Self-Improving LLM Agents Estimating the Accuracies of Multiple Classifiers Without Labeled Data
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db240d47-19ca-4da0-8977-24a1072d9624 · outbound
Memory Reward Inflation in Self-Improving LLM Agents The logic of NTQR evaluations of noisy AI agents: Complete postulates and logically consistent error correlations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972e9bf5-95b7-4051-8b8c-e970ef6b557d · outbound
Memory Reward Inflation in Self-Improving LLM Agents Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f5eede-59ea-405a-adce-2633ee73e701 · outbound
Memory Reward Inflation in Self-Improving LLM Agents SimCSE: Simple contrastive learning of sentence embeddings
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.