Pith. sign in

Paper Citation Record · LEDGER

Language Model Self-improvement by Reinforcement Learning Contemplation

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2305.14483.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14483 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:28:48.464964Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.461679Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28bd4512-e841-440e-9f47-4f4374a6446d · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.809974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:b17c7150a5aa7109779d6f903e56f6f0d28f740b79c822f59a6bbcb8b88c7236

Observation 0606a24e-2b8a-4b71-8ce6-82a7fb95be5d · inbound

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering cites this paper.

Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T18:28:48.464964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:28:48.464964Z digest=sha256:1d2f493f20e5ba18dda9df0f72fccd5671ca3e440d48c2f4466392e9395a5268

Observation e0a80c5d-efc9-4ee1-ac42-e7fcbf1dbff1 · inbound

Self-Improvement in Language Models: The Sharpening Mechanism cites this paper.

Self-Improvement in Language Models: The Sharpening Mechanism Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.748447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.748447Z digest=sha256:a4a5dab55fd7b00807ddf38dd5300dbd4dd612caa81ae2d75d6e9b5df26244f4

Observation 772a2b83-e15b-477e-8172-fcd57b69dc48 · inbound

ALoFTRAG: Automatic Local Fine Tuning for Retrieval Augmented Generation cites this paper.

ALoFTRAG: Automatic Local Fine Tuning for Retrieval Augmented Generation Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:47:46.146757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:47:46.146757Z digest=sha256:7fffb280d22ebac4fb235a3821987fa941b4495c4e9f00a71932dd398c898b30

Observation b070e946-c315-49d8-9c69-342c55e82370 · inbound

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression cites this paper.

Training Dialogue Systems by AI Feedback for Improving Overall Dialogue Impression Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:52.007090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:57:52.007090Z digest=sha256:1d2941ec811da91190ad44233ba44fb5c5e3768e2f060d0991dd30248b0d67fd

Observation db0cbb67-4e25-44f8-8137-86ad2df344de · inbound

ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection cites this paper.

ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:40.606818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:04:40.606818Z digest=sha256:29014fe5b3b5c389a9f2fc04d0266855aeb200dfb0bfe78275989c8c47f12e75

Observation 146c8736-2eec-4e14-a68a-1608ba13afdb · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.233251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T04:40:23.895430Z digest=sha256:01f3a961019efb6fb13bee7ca7c801c80bee81966d4a94a094db02727e0cf117

Observation 935ec621-b806-41bb-8664-74fac4b71b8c · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.282952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:48:49.322744Z digest=sha256:942db5d51dfcea416481175b27ebaede199cd4bd98e4ec083cedfa40885c48d2

Observation bb69859f-194a-438b-a684-327ef2e2384c · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:24:08.664516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:95c772d5eec1c976675ca90d4df82dd3a827c7522a6b0df7c41de5880b17ac7b

Observation 5b85f5bf-8ad4-41b0-8a03-48eef99ed6c6 · inbound

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models cites this paper.

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:14.199980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T07:36:19.786202Z digest=sha256:4486562c685f5a5c245cb58f7d62bac0ce026674a2f72d39a0ee77181535b27e

Observation 6911b7cb-e2a2-4332-af9f-7fed2b06b107 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.463123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:63586fb116ff9a8c29d79f2d9edf45e46af6373eee0307144e02526ebcf9d1e8

Observation a1838133-bd51-4c1f-bf27-58ee53ce2e5c · inbound

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR cites this paper.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.462207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.462207Z digest=sha256:53a0712f08e0c3fd190ccdf9df7b49edd072f25385c7ccbeba96711117e230e6