Pith. sign in

Paper Citation Record · LEDGER

Self-Improving Large Language Models via Progressive Experience Evolution

As of 13 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.02139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02139 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:53.840638Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a86376f1-91ec-4e3f-a42a-ab7681d51475 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Improving Large Language Models via Progressive Experience Evolution Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.628042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.628042Z digest=sha256:b33c296dc5985a276963329b435e817389663501c125e3a959bd8f9f4c0c0dbd

Observation 8ae55e69-9a49-4504-b1b4-43e0f52c452d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.132451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.132451Z digest=sha256:2f96f341c328ee5cb5316885d2a046c17eba1c6f4fbd489503896b0fdf1254c7

Observation 7f9e970c-930a-4876-a105-f50ebe8d9dfc · outbound

This paper cites Let's Verify Step by Step.

Self-Improving Large Language Models via Progressive Experience Evolution Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.433047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.433047Z digest=sha256:97025dbe74b067324664bf542cefbef5baeca2107906d9bed8253c41843f93db

Observation fd8b029b-c614-420d-963a-c6901beb1782 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.493528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.493528Z digest=sha256:2facdbe416c5d719587dc04c8ec73bf5f30db44cffa1ba171d7e088b94c5f892

Observation 8ab77ad3-d340-48c2-8ca2-09ad79301ced · outbound

This paper cites Mexico City, Mexico: Association for Compu- tational Linguistics.

Self-Improving Large Language Models via Progressive Experience Evolution Mexico City, Mexico: Association for Compu- tational Linguistics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:54.527323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T00:14:52.584832Z digest=sha256:b17fc8caefcf55384ef177fb867de31891eced77fcd8c19f9412719af11b56b4

Observation 17bae05d-6cd2-4496-a84d-ee2afd62df82 · outbound

This paper cites Privileged Information Distillation for Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Privileged Information Distillation for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.667860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.667860Z digest=sha256:e0956c4306bcf1b98538b0b2434380fb22e5d9b0fb6cf213c2b43cce350b81fa

Observation f931f582-dfb6-4887-89ad-1bd468f449cd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.737140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.737140Z digest=sha256:1b2604809663401feb2ccac4a779bd20ae289ea11f6b2188c4f6f378fc83421d

Observation bf57b854-9cf4-4a50-9db3-79684a2b7804 · outbound

This paper cites Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning.

Self-Improving Large Language Models via Progressive Experience Evolution Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:54.057384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T00:14:52.852963Z digest=sha256:bc8653b73a379d798e8c3d0a1dd0353d3e4cbfd24835ccd7eb40da71a3c0e6c5

Observation cb45d8de-3697-4783-9d55-96d7e867809c · outbound

This paper cites Learning to summarize from human feedback.

Self-Improving Large Language Models via Progressive Experience Evolution Learning to summarize from human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.945625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.945625Z digest=sha256:14736b0b4594410e9f31f4cf19e66d34b75d9e731ede6dc52ae65bc7063106f5

Observation 6fc808de-7b5f-48e0-8ca0-15ed7ed9a7d0 · outbound

This paper cites an unresolved cited work.

Self-Improving Large Language Models via Progressive Experience Evolution Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:14:54.399597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T00:14:53.053392Z digest=sha256:51a545e96e7ea68b8adf7e32de2f215ed70038d749e540b34a261bba06ab1e06

Observation 7ee965ff-f280-4811-a330-c9a58e959cee · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Self-Improving Large Language Models via Progressive Experience Evolution Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.167075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.167075Z digest=sha256:eed770dcfce59035ecab044c0e2269ef044d13ebba0c747142cee9f9f27156bf

Observation 917f942a-4e44-4f62-9862-d96d96fa09b2 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.270921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.270921Z digest=sha256:0063200464afffb2ea5f3226fbc71bd98306f1053bb22ec6ffef109dcfc97457

Observation ae2fa355-1d36-42e6-b2c9-856a172610a1 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.385502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.385502Z digest=sha256:4ca77f491941ab10045a9c068448ce55af50b9686234aa02ba4873ff5637a73b

Observation 90a4fbce-1e03-4d14-8fca-69c32febfd6f · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.476750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.476750Z digest=sha256:aaeb4d3a82f6113cb55bbdf7129de7925387ceeea80c321d8f1ab48676c6e053

Observation 552e9929-b0b7-4a2f-a4b3-f196aa85df3c · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution ReAct: Synergizing Reasoning and Acting in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.569449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.569449Z digest=sha256:af1785fc6746991ffe88a9137ab2c3d9dd4ffdd15040d4d0fd2e8ce755287bb7

Observation 69d5ac64-ecc7-4360-ad19-e41cc7ad756c · outbound

This paper cites Expel:Llmagentsareexperientiallearners.

Self-Improving Large Language Models via Progressive Experience Evolution Expel:Llmagentsareexperientiallearners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:14:54.242192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T00:14:53.661199Z digest=sha256:a72bd301aeb78494729adcb2700f34ea3571fb050d5c237a5cff48b41e906592

Observation 6884dd0e-b56d-48ac-b1f8-b0a7828cbbea · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.742184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.742184Z digest=sha256:2f933241c2257a16df66eff0026c42076269fb700cc3c67866904e0048c6107e

Observation f922c2e8-bb6f-4c45-8aeb-fe438a28dee5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Self-Improving Large Language Models via Progressive Experience Evolution Fine-Tuning Language Models from Human Preferences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:53.840638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:53.840638Z digest=sha256:41ee3ae64652db4ee107dc3742957aa5f774abfab558c990a7a6dda92a300182

Observation 3d7fd242-6134-4c8c-87ed-4cc0e74efd5a · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Self-Improving Large Language Models via Progressive Experience Evolution Distilling the Knowledge in a Neural Network

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.018406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.018406Z digest=sha256:e9b4c995f2f10ee4794f85c102c84ba323cc012fa9fa6a3f988f9183e54a1e27

Observation 936ce180-9f81-4839-9535-6323bcc33898 · outbound

This paper cites Scaling Laws for Neural Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution Scaling Laws for Neural Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.247773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.247773Z digest=sha256:5f54e056fc790e68d0ae233bb43072e564d20e027c64b93752fe7af71feb50f5

Observation 4058ba69-5813-41c8-83de-99ec992faf60 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:52.339631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:52.339631Z digest=sha256:ffba4e457ed8368d45227c7b63343052d344436c531620c35e589f5207ee946c

Observation 1261451e-4e58-4460-b123-6dab728803a0 · outbound

This paper cites GPT-4 Technical Report.

Self-Improving Large Language Models via Progressive Experience Evolution GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.435355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.435355Z digest=sha256:bbf484110fad0dda6a6781267a67aa1a8681e9aa390e66c550a8770cdd6d8647

Observation 562d20f0-28e4-49f3-a75b-25b9f7dc0ce7 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Self-Improving Large Language Models via Progressive Experience Evolution On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.525171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.525171Z digest=sha256:41dbd8cfecb81412a3bf14d54ca4170fcb077cb1812a754a88a6b0359329940e

Observation e3dc3a32-dfd4-493b-91d6-fa61925135e4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Self-Improving Large Language Models via Progressive Experience Evolution DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.898510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.898510Z digest=sha256:5dead459c5b0bfd2717933fd352fdc8fbba9264fb0a5d62c3456c8ef5e1425b6

Observation 4c0bdd01-606b-4211-8a2b-71f6b2bf37e1 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Self-Improving Large Language Models via Progressive Experience Evolution MiniLLM: On-Policy Distillation of Large Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:51.747842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:51.747842Z digest=sha256:3a238cb326b4b7775ec82cb6cb2ebd6ab5fbd88abc1b9bb4c15f7bf2f53202af

Pith citing papers

No inbound Pith citation observations are available.