Pith. sign in

Paper Citation Record · LEDGER

Self-Improving Large Language Models via Progressive Experience Evolution

As of 7 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.02139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02139 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:53.840638Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a86376f1-91ec-4e3f-a42a-ab7681d51475 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Improving Large Language Models via Progressive Experience Evolution Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.628042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.628042Z digest=sha256:febf1627e426a77e2a8632ad6e63aa98bedb4e75e9470f5bf77e52cd7747a598

Observation 8ae55e69-9a49-4504-b1b4-43e0f52c452d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.132451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.132451Z digest=sha256:d1794518be43ed99df8ca1f5da479ceba81c222b33ce3c9e30deb4a25cee2e9c

Observation 7f9e970c-930a-4876-a105-f50ebe8d9dfc · outbound

This paper cites Let's Verify Step by Step.

Self-Improving Large Language Models via Progressive Experience Evolution Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.433047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.433047Z digest=sha256:df260af61d3aab83d50a28c0193db2c6f9a00b1eca75e839bebcbef68e56f6cc

Observation fd8b029b-c614-420d-963a-c6901beb1782 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.493528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.493528Z digest=sha256:d4ddb235c209c20fdeb1f6c038df0f5ae5dad66b511e3b28c0f1db2374e65fc2

Observation 8ab77ad3-d340-48c2-8ca2-09ad79301ced · outbound

This paper cites Mexico City, Mexico: Association for Compu- tational Linguistics.

Self-Improving Large Language Models via Progressive Experience Evolution Mexico City, Mexico: Association for Compu- tational Linguistics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:54.527323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:52.584832Z digest=sha256:4dcb6658d92c34ee19a2b5c5d99175cb3a142af463bbe698bc5c71b88fb872c4

Observation 17bae05d-6cd2-4496-a84d-ee2afd62df82 · outbound

This paper cites Privileged Information Distillation for Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Privileged Information Distillation for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.667860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.667860Z digest=sha256:99322f15f88f660c45f1522ce83c62ec07e847f87de9626d6eb0a5bb92a2e477

Observation f931f582-dfb6-4887-89ad-1bd468f449cd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.737140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.737140Z digest=sha256:ee7fd5048d748bf4124212dd7d6543e81224dd7e520fd02d03bd1b3e45f7aac2

Observation bf57b854-9cf4-4a50-9db3-79684a2b7804 · outbound

This paper cites Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning.

Self-Improving Large Language Models via Progressive Experience Evolution Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:54.057384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:52.852963Z digest=sha256:5022e3ac82403f0c5b8f8f8f3a4c038509cb53bd2edc44df4866efbe7771a625

Observation cb45d8de-3697-4783-9d55-96d7e867809c · outbound

This paper cites Learning to summarize from human feedback.

Self-Improving Large Language Models via Progressive Experience Evolution Learning to summarize from human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.945625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.945625Z digest=sha256:5d496221154df7bdaac3e18c661bab99e6bc83a083c33c358f950379b44c5ebe

Observation 6fc808de-7b5f-48e0-8ca0-15ed7ed9a7d0 · outbound

This paper cites an unresolved cited work.

Self-Improving Large Language Models via Progressive Experience Evolution Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:14:54.399597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:53.053392Z digest=sha256:97154bfde62629e2285c425bd09940eb6834eb1bbf23291e8b830a429c36ee05

Observation 7ee965ff-f280-4811-a330-c9a58e959cee · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Self-Improving Large Language Models via Progressive Experience Evolution Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.167075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.167075Z digest=sha256:78809bfdfdc8e9e5ac491665b8b2c0a9cf0b12f5acf7e7b1ebc7fd9c89805415

Observation 917f942a-4e44-4f62-9862-d96d96fa09b2 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.270921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.270921Z digest=sha256:5e69be371391573f7f20f34d47df6a3ceced2df86d2cbfe97d2e6b7fe0197e7d

Observation ae2fa355-1d36-42e6-b2c9-856a172610a1 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.385502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.385502Z digest=sha256:a1ffc8e23c1bfd60b5bcff303e567e8ce26f3ea37ed85ede2191b18cfa4cf7d0

Observation 90a4fbce-1e03-4d14-8fca-69c32febfd6f · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.476750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.476750Z digest=sha256:db311b503f12387a90a11ff672819a67c89aee273b8aaaf430d9973970f848a7

Observation 552e9929-b0b7-4a2f-a4b3-f196aa85df3c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution ReAct: Synergizing Reasoning and Acting in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.569449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.569449Z digest=sha256:5fab93051e9de557ba95818d1deb923fa1a3ea34a7b2ee5ac69a1dbc1f54fc98

Observation 69d5ac64-ecc7-4360-ad19-e41cc7ad756c · outbound

This paper cites Expel:Llmagentsareexperientiallearners.

Self-Improving Large Language Models via Progressive Experience Evolution Expel:Llmagentsareexperientiallearners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:54.242192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:14:53.661199Z digest=sha256:4c83f795ada083004fcf1b2d6a3877f6accf5c1b4cea7c6f596f50259f30d5ca

Observation 6884dd0e-b56d-48ac-b1f8-b0a7828cbbea · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.742184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.742184Z digest=sha256:bcf005793a325615b44b5e8c0275b70bee512522362bcd37bf39cab1ab6c2687

Observation f922c2e8-bb6f-4c45-8aeb-fe438a28dee5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Self-Improving Large Language Models via Progressive Experience Evolution Fine-Tuning Language Models from Human Preferences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.840638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.840638Z digest=sha256:8d64386e2405bb41b05df75404c80b21f094af1d6d0d025e5d255759bc4617d0

Observation 3d7fd242-6134-4c8c-87ed-4cc0e74efd5a · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Self-Improving Large Language Models via Progressive Experience Evolution Distilling the Knowledge in a Neural Network

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.018406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.018406Z digest=sha256:e9cd8d7e87f4aceb3a9161c8ff3a3e32767f1cf81355f1bda1fed68fd784d7f3

Observation 936ce180-9f81-4839-9535-6323bcc33898 · outbound

This paper cites Scaling Laws for Neural Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Scaling Laws for Neural Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.247773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.247773Z digest=sha256:7386233e47c88b189d397b5ab8b889caa7175527b9754f2211ded2a893d637ed

Observation 4058ba69-5813-41c8-83de-99ec992faf60 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.339631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.339631Z digest=sha256:561bd91dc882ba6cb83b3d872891f8ad7947326ec9a2f44d045503bc41eda9a8

Observation 1261451e-4e58-4460-b123-6dab728803a0 · outbound

This paper cites GPT-4 Technical Report.

Self-Improving Large Language Models via Progressive Experience Evolution GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.435355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.435355Z digest=sha256:e0137ac30fad24cad2b121882807d83f908e2c43a95a1a99bd01ce884abefe69

Observation 562d20f0-28e4-49f3-a75b-25b9f7dc0ce7 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Self-Improving Large Language Models via Progressive Experience Evolution On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.525171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.525171Z digest=sha256:a0693ee19b2f154e2b9be994e9fdb447319990dbe893001fd5aaef871bdb3157

Observation e3dc3a32-dfd4-493b-91d6-fa61925135e4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.898510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.898510Z digest=sha256:6a66de4c4fa543d256ea4c19320f48719f81a3871ef076ba87bf766e1918c31f

Observation 4c0bdd01-606b-4211-8a2b-71f6b2bf37e1 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution MiniLLM: On-Policy Distillation of Large Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.747842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.747842Z digest=sha256:7074f50bde6c5ec7edca7c633d62756eddf86121852c41cf8bdea7afcd2158df

Pith citing papers

No inbound Pith citation observations are available.